Devices and methods for in-network learning for dynamic wireless networks

The relay device with a relay machine learning model enhances In-Network Learning's robustness in dynamic wireless networks by aggregating encoded data and adapting to topology changes, addressing the limitations of conventional frameworks.

WO2026076576A1PCT designated stage Publication Date: 2026-04-16HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/123501
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Conventional In-Network Learning frameworks struggle to handle dynamic wireless networks with variable numbers of edge devices, fluctuating network topologies, and changing channel conditions, necessitating robustness improvements.

Method used

Implementing a relay device with a relay machine learning model that aggregates encoded input data from edge devices and a fusion center, trained through a two-stage In-Network Learning process, allowing topology-agnostic operation and reducing the need for full re-training.

Benefits of technology

Enhances the robustness of In-Network Learning against topology changes, conserving communication bandwidth, reducing energy consumption, and minimizing training/configuration time in dynamic wireless networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123501_16042026_PF_FP_ABST
    Figure CN2024123501_16042026_PF_FP_ABST
Patent Text Reader

Abstract

Devices and methods are disclosed for in-network learning for dynamic wireless networks. For instance, a relay device (120a, b) for a wireless network (100) is disclosed, wherein the wireless network (100) comprises a plurality of edge devices (110a-n), each edge device (110a-n) configured to collect input data of one or more modalities of a plurality of modalities, one or more relay devices (120a, b) and a fusion center device (130). The relay device (120a, b) is configured to implement a relay machine learning, ML, model (121a-c) configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of the specific modality, wherein the encoded input data of the specific modality is provided by one or more of the plurality of edge devices (110a-n) and wherein each of the one or more edge devices (110a-n) implements an encoder ML model (111a-n) configured to generate the encoded input data of the specific modality. Moreover, the relay device (120a, b) is configured to provide the aggregated encoded input data of the specific modality to the fusion center device (130) implementing a decoder ML model (131) configured to generate output data by decoding the aggregated encoded input data of the specific modality provided by the relay device (120a, b). The plurality of encoder ML models (111a-n) and the decoder ML model (131) are based on a first In-Network Learning stage of a first virtual wireless network, wherein the first virtual wireless network comprises a plurality of virtual edge devices implementing the plurality of encoder ML models (111a-n) for collecting input data of different modality and a virtual fusion center device implementing the decoder ML model (131). The relay ML model (121a-c) is trained with a second In-Network Learning stage of a second virtual wireless network comprising the plurality of virtual edge devices for collecting input data of the specific modality, each virtual edge device implementing the trained encoder ML model (111a-n), a virtual relay device implementing the relay ML model (121a, b), and a virtual fusion center device implementing a decoder ML model for input data of the specific modality.
Need to check novelty before this filing date? Find Prior Art

Description

DEVICES AND METHODS FOR IN-NETWORK LEARNING FOR DYNAMIC WIRELESS NETWORKSTECHNICAL FIELD

[0001] The present invention relates to data processing. More specifically, the present invention relates to devices and methods for in-network learning of wireless networks, in particular dynamic networks, such as mobile networks, for instance 3GPP mobile networks.BACKGROUND

[0002] Future communication systems, including 3GPP 6G, are envisioned to rely on cooperating intelligent agents whose combined processing capabilities will be on par with the capabilities of state-of-the-art centralized large language models (LLMs) such as GPT-5. These distributed agents will be able to jointly process multimodal data (e.g., images, videos or sensing data collected by each agent) , adapt to changing, i.e. dynamic environments, solve various complex tasks, and share obtained knowledge with other devices.

[0003] In the above vision, all relevant input data are distributed across multiple agents while the joint data interpretation and decision making should be done by a remote fusion center. This poses a significant technical challenge, because limited communication bandwidth and privacy constraints usually prevent transmitting raw input data of all agents to the fusion server. In order to mitigate this challenge, a possible and increasingly popular solution has become In-Network Learning (INL) (disclosed in I. Aguerri and A. Zaidi, "Distributed Variational Representation Learning" in IEEE Transactions on Pattern Analysis &Machine Intelligence, vol. 43, no. 01, pp. 120-138, 2021, doi: 10.1109 / TPAMI. 2019.2928806, and in M. Moldoveanu and A. Zaidi, “In-Network Learning: Distributed Training and Inference in Networks” , arXiv preprint arXiv: 2107.03433) which allows to train or evaluate machine learning models in a distributed manner without a need to transmit all local data at a single place. In In-network Learning, each agent keeps a piece of a large model and uses it to process local input data. The result of the local processing are latent descriptions which gather only most relevant information of the local input data and thus are much easier to transmit to the fusion center.

[0004] However, the basic framework of In-Network Learning (plain INL) assumes that the number of agents, the network topology, and the distribution of input data remains identical during training and inference phases. Unfortunately, these requirements are difficult to meet in highly dynamic mobile networks with a variable number of edge devices, dynamic network topologies, and fluctuating channel conditions. For this reason, there is an urgent need to propose novel techniques that increase the robustness of In-network Learning against changes of a network topology, i.e. for dynamic wireless networks.SUMMARY

[0005] It is an object of the invention to provide improved devices and methods for in-network learning of wireless networks, in particular highly dynamic networks, such as mobile networks, for instance 3GPP mobile networks.

[0006] The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.

[0007] According to a first aspect a relay device (herein also referred to as relay node) for a wireless network is provided. The wireless network comprises a plurality of edge devices (or agents) , wherein each edge device is configured to collect input data of one or more modalities of a plurality of modalities, one or more relay devices, including the relay device according to the first aspect, and a fusion center device.

[0008] The relay device according to the first aspect is configured to implement a relay machine learning, ML, model configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of the specific modality, wherein the encoded input data of the specific modality is provided by one or more of the plurality of edge devices, wherein each of the one or more edge devices implements an encoder ML model  configured to generate the encoded input data of the specific modality. Moreover, the relay device according to the first aspect is configured to provide the aggregated encoded input data of the specific modality to a further relay device or to the fusion center device implementing a decoder ML model configured to generate output data by decoding the aggregated encoded input data of the specific modality provided by the relay device. The plurality of encoder ML models implemented by the plurality of edge devices and the decoder ML model implemented by the fusion center device are based on a first In-Network Learning, i.e. training stage of a first virtual wireless network, wherein the first virtual wireless network comprises a plurality of virtual edge devices implementing the plurality of encoder ML models for collecting input data of a different modality and a virtual fusion center device implementing the decoder ML model. The relay ML model implemented by the relay device is trained with a second In-Network Learning stage of a second virtual wireless network comprising a plurality of virtual edge devices for collecting input data of the specific modality, each edge device implementing the trained, i.e. fixed or frozen encoder ML model, a virtual relay device implementing the relay ML model, and a virtual fusion center device implementing a decoder ML model for input data of the specific modality. As will be appreciated, the relay device according to the first aspect allows making In-Network Learning more robust against topology changes of the wireless network during inference phase without the need to fully re-train all distributed machine learning models whenever the topology of the wireless network changes. This may, for instance, save communication bandwidth, reduce energy consumption and reduce network training / configuration time in a wireless network implementing In-Network Learning.

[0009] As used herein, “In-Network Learning (INL) ” allows to train or evaluate machine learning models in a distributed manner without a need to transmit all local data at a single place. In In-network Learning, each agent, such as an edge device or a relay device, keeps a piece of a large model and uses it to process local input data. The result of the local processing are latent descriptions which gather only most relevant information of the local input data and thus are much easier to transmit to the fusion center device.

[0010] In a further possible implementation form, the relay ML model comprises a cross-view attention module.

[0011] In a further possible implementation form, each encoder ML model comprises one or more feed-forward neural networks.

[0012] In a further possible implementation form, the relay device is configured to inform the plurality of edge devices and / or the fusion center device about a change of the topology of the wireless network for triggering a re-training of the wireless network with In-Network Learning.

[0013] In a further possible implementation form, the relay ML model implemented by the relay device is configured to be further trained, i.e. fine-tuned with a third In-Network Learning, i.e. fine-tuning stage of the wireless network comprising the plurality of edge devices for collecting input data of the specific modality, each edge device implementing the encoder ML model trained in the first In-Network Learning stage, the relay device implementing the relay ML model trained in the second In-Network Learning stage, and the fusion center device implementing the decoder ML model trained in the first In-Network Learning stage.

[0014] According to a second aspect a wireless network, such as a 3GPP mobile network, is provided, comprising:

[0015] a plurality of edge devices, each edge device edge device implementing an encoder ML model;

[0016] one or more relay devices according to the first aspect; and

[0017] a fusion center device implementing a decoder ML model.

[0018] According to a third aspect a method is provided for operating a relay device, i.e. a relay node of a wireless network, wherein the wireless network comprises a plurality of edge devices, each edge device configured to collect input data of one or more modalities of a plurality of modalities, one or more relay devices and a fusion center device. The method according to the third aspect comprises the following steps executed by the relay device:

[0019] implementing a relay machine learning, ML, model configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of the specific modality, wherein the encoded input data of the specific modality is provided by one or more of the plurality of edge devices, wherein each of the one or more edge devices implements an encoder ML model configured to generate the encoded input data of the specific modality; and

[0020] providing the aggregated encoded input data of the specific modality to a further relay device or to a fusion center device implementing a decoder ML model configured to generate output data by decoding the aggregated encoded input data of the specific modality provided by the relay device, wherein the plurality of encoder ML models and the decoder ML model are based on In-Network Learning of a virtual wireless network, wherein the virtual wireless network comprises a plurality of virtual edge devices implementing the plurality of encoder ML models for collecting input data of different modality and a virtual fusion center device implementing the decoder ML model, and wherein the relay ML model is trained with a second In-Network Learning stage of a second virtual wireless network comprising the plurality of virtual edge devices for collecting input data of the specific modality, each virtual edge device implementing the trained encoder ML model, a virtual relay device implementing the relay ML model, and a virtual fusion center device implementing a decoder ML model for input data of the specific modality.

[0021] The method according to the third aspect can be performed by the relay device according to the first aspect. Thus, further features of the method according to the third aspect result directly from the functionality of the relay device according to the first aspect and its different implementation forms described above and below.

[0022] According to a fourth aspect a computer program or a computer program product is provided, comprising a computer-readable storage medium carrying program code which causes a computer or a processor to perform the method according to the third aspect, when the program code is executed by the computer or the processor.

[0023] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In the following embodiments of the invention are described in more detail with reference to the attached figures and drawings, in which:

[0025] Fig. 1 is a schematic diagram illustrating a conventional wireless network;

[0026] Fig. 2 is a schematic diagram illustrating different scenarios for dynamic changes of a conventional wireless network;

[0027] Fig. 3 is a schematic diagram illustrating the architecture of a first virtual wireless network, including a plurality of virtual edge devices and a virtual fusion center device, used for a first stage of In-Network Learning for a wireless network according to an embodiment;

[0028] Fig. 4a is a schematic diagram illustrating the architecture of a second virtual wireless network, including a plurality of virtual edge devices, a virtual relay device and a fusion center device, used for a second stage of In-Network Learning for a wireless network according to an embodiment;

[0029] Fig. 4b is a schematic diagram illustrating a concatenation of ML models implemented by relay devices of a wireless network according to an embodiment;

[0030] Fig. 4c is a schematic diagram illustrating In-Network Learning of ML models implemented by different devices of a wireless network, including a plurality of edge devices, a plurality of relay devices, and a fusion center device according to an embodiment;

[0031] Fig. 4d is a schematic diagram illustrating In-Network Learning of a wireless network, including a plurality of edge devices, a plurality of relay devices, and a fusion center device according to an embodiment;

[0032] Fig. 4e is a schematic diagram illustrating in more detail two of the plurality of edge devices of the wireless network of figure 4d;

[0033] Fig. 5a is a signaling diagram illustrating messaging between the devices of a wireless network, including a plurality of edge devices, a plurality of relay devices, and a fusion center device according to an embodiment for re-training of the wireless network;

[0034] Fig. 5b is a diagram illustrating a message exchanges between the devices of a wireless network, including a plurality of edge devices, a plurality of relay devices, and a fusion center device according to an embodiment for re-training of the wireless network;

[0035] Fig. 6a is a diagram illustrating a fusion of latent descriptions implemented by a relay device according to an embodiment;

[0036] Fig. 6b is a diagram illustrating a cross-view attention mechanism used for the fusion of latent descriptions implemented by a relay device according to an embodiment; and

[0037] Fig. 7 is a flow diagram illustrating a method according to an embodiment for operating a relay device of a wireless network for In-Network Learning.

[0038] In the following identical reference signs refer to identical or at least functionally equivalent features.

[0039] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the invention or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the invention may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0041] For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps) , even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units) , even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.

[0042] Figure 1 is a schematic diagram illustrating a conventional wireless network 10 implemented for distributed machine learning, ML, in the form of In-Network Learning (more specifically, figure 1 illustrates the distributed inference using In-Network Learning) . The conventional wireless network 10 illustrated in figure 1 comprises a plurality of edge devices 11a-c (referred to as nodes 1 to 3 in figure 1) , a relay device 12a (referred to as node 4 in figure 1) , and a fusion center device 13, for instance, a server 13 (referred to as node N in figure 1) . As indicated in figure 1, each of these nodes may implement ML models and communicate with each other via communication channels having different communication capacities Ci, j.

[0043] The plurality of edge devices 11a-c may have access to local data (possibly multimodal, e.g., images, audio, text, etc. ) and the fusion center device 13 may be configured to make a joint decision based on the relevant information jointly possessed by all the edge devices 11a-c. To this end, as already mentioned above, respective ML models may be implemented on the different devices 11a-c, 12, and 13 and trained using In-Network Learning. Briefly summarized, In-Network Learning allows to train or evaluate machine learning models in a distributed manner without a need to transmit all local data at a single place. In In-network Learning, each agent, such as the edge devices 11a-c or the relay device 12, keeps a piece of a large model and uses it to process local input data. The result of the local processing are latent descriptions which gather only most relevant information of the local input data and thus are much easier to transmit to the fusion center device 13. Further details of In-Network Learning are disclosed in I. Aguerri and A. Zaidi, "Distributed Variational Representation Learning" in IEEE Transactions on Pattern Analysis &Machine Intelligence, vol. 43, no. 01, pp. 120-138, 2021, doi: 10.1109 / TPAMI. 2019.2928806, and in M. Moldoveanu and A. Zaidi, “In-Network Learning: Distributed Training and Inference in Networks” , arXiv preprint arXiv: 2107.03433 as well as in the published patent applications “Robust in-network learning: system and method for learning and inference in presence of missing  data” , WO2023 / 147857 and “Joint relevance-channel aware wireless device scheduling for in-network learning” , WO2023 / 151802, which are fully incorporated herein by reference.

[0044] Conventional In-Network Learning, as disclosed, for instance, in the documents mentioned above, requires that the topology of the wireless network 10 and the channel capacities remain the same during both training and inference phases. Thus, conventional In-Network Learning is not capable of handling dynamic scenarios, such as the ones illustrated in figure 2. As will be described in more detail in the following, embodiments disclosed herein allow to make In-Network Learning robust against topology changes of the wireless network during inference phase without the need to fully re-train all distributed machine learning models whenever the topology changes. More specifically, embodiments disclosed herein allow for topology-agnostic In-Network Learning with a variable number of edge devices 110a-n processing K input modalities, arbitrarily-connected intermediate relay devices 120a, b, and a fusion center device 130.

[0045] Figure 4d is a schematic diagram illustrating a wireless network 100 with a plurality of devices trained with In-Network Learning. More specifically, the wireless network 100 comprises a plurality of edge devices 110a-n, a plurality of relay devices 120a, b according to an embodiment, and a fusion center device 130. The plurality of edge devices 110a-n are configured to collect input data of one or more modalities of a plurality of modalities, such as images, audio, text, and the like, and the fusion center device 130 may be configured to make a joint decision based on the relevant information jointly possessed by all the edge devices 110a-n. To this end, as will be described in more detail in the following, each of the devices of the wireless network may implement one or more suitably configured and trained ML models. In an embodiment, the wireless network 100 may be a mobile network 100, in particular a 3GPP mobile network 100, such as a 5G or 6G network. In an embodiment, the edge devices 110a-n may be, for instance, UEs 110a-n, the plurality of relay devices 120a, b may be base stations 120a, b or other Radio Access Network or Core Network devices 120a, b, and the fusion center device 130 may be an application server.

[0046] Each of the edge devices 110a-n may comprise processing circuitry, such as one or more processors or cores, for processing data and a memory for storing and retrieving data and implementing one or more ML models. Furthermore, each of the edge devices 110a-n may comprise a communication interface for exchanging data with the other devices of the wireless network 100 via a wireless connection. The processing circuitry of each edge device 110a-n for operating one or more ML models may be implemented in hardware and / or software. The hardware may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs) , field-programmable gate arrays (FPGAs) , digital signal processors (DSPs) , or general-purpose processors. The memory of each edge device 110a-n may store executable program code which, when executed by the processing circuitry causes the edge device 110a-n to perform the functions and methods described herein.

[0047] Likewise, each of the relay devices 120a, b may comprise processing circuitry, such as one or more processors or cores, for processing data and a memory for storing and retrieving data and implementing one or more ML models. Furthermore, each of the relay devices 120a, b may comprise a communication interface for exchanging data with the other devices of the wireless network 100 via a wireless or wired connection. The processing circuitry of each relay device 120a, b for operating one or more ML models may be implemented in hardware and / or software. The hardware may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs) , field-programmable gate arrays (FPGAs) , digital signal processors (DSPs) , or general-purpose processors. The memory of each relay device 120a, b may store executable program code which, when executed by the processing circuitry causes the relay device 120a, b to perform the functions and methods described herein.

[0048] Likewise, the fusion center device 130, e.g. server 130 may comprise processing circuitry, such as one or more processors or cores, for processing data and a memory for storing and retrieving data and implementing one or more ML models. Furthermore, the fusion center device 130 may comprise a communication interface for exchanging data with the other devices of the wireless network 100 via a wireless or wired connection. The processing circuitry of the fusion center device 130 for operating one or more ML models may be implemented in hardware and / or software.  The hardware may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs) , field-programmable gate arrays (FPGAs) , digital signal processors (DSPs) , or general-purpose processors. The memory of the fusion center device 130 may store executable program code which, when executed by the processing circuitry causes the fusion center device 130 to perform the functions and methods described herein.

[0049] As will be described in more detail in the following under further reference to figures 3, 4a, 4b, 4c, and 4e, each relay device, such as the exemplary relay device 120a, is configured to implement one or more relay machine learning, ML, models 121a-c, wherein each relay ML model 121a-c is configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of the specific modality. The encoded input data of a respective specific modality is provided by one or more of the plurality of edge devices 110a-n, wherein each of the plurality of edge devices 110a-n implements one or more encoder ML models 111a-n, wherein each encoder ML model 111a-n is configured to generate the encoded input data of the respective specific modality. For instance, in the exemplary embodiment shown in figure 4d, encoded input data of three different modalities are provided by the edge device 110a to the relay device 120a and encoded input data of two different modalities are provided by the edge device 110b to the relay device 120a. Thus, in the embodiment of figure 4d, the edge device 110a implements the encoder ML models 111a, 111b and 111c for generating encoded input data for three different modalities, while the edge device 110b implements the encoder ML models 111a and 111c for generating encoded input data for two different modalities. This is illustrated in more detail in figure 4e for the further edge devices 110m and 110n. As can be taken from figure 4e, the edge device 110m implements the encoder ML model 111c for generating encoded input data for one specific modality, while the edge device 110n implements the encoder ML models 111a, 111b and 111c for generating encoded input data for three different modalities. As further illustrated in figures 4d and 4e, each edge device 110a-n may implement a demultiplexer 112 for demultiplexing the multimodal input data for the one or more encoder ML models 111a-c and a multiplexer 113 for generating the multimodal output data.

[0050] In the embodiment shown in figure 4d, the relay device 120a is further configured to provide the aggregated encoded input data of the specific modalities provided by the edge devices 110a, 110b to the further relay device 120b. The further relay device 120b, in turn, is configured to aggregate the encoded input data of the specific modalities from the relay device 120a and the edge device 110m and to provide the aggregated encoded input data of the specific modalities to the fusion center device 130. The fusion center device 130 implements a decoder ML model 131, 133 configured to generate output data by decoding the aggregated encoded input data of the specific modalities provided by the further relay device 120b and the edge device 110n.

[0051] As can be taken from figure 3, the plurality of encoder ML models 111a-n implemented by the plurality of edge devices 110a-n and the decoder ML model 131, 133 implemented by the fusion data center are based on a first In-Network Learning stage of a first virtual wireless network 300, wherein the first virtual wireless network 300 comprises a plurality of virtual edge devices implementing the plurality of encoder ML models 111a-n for collecting input data of different modality and a virtual fusion center device implementing the decoder ML model 131. As illustrated in figure 4a, the relay ML model 121a, b implemented by the respective relay device 120a, b is trained with a second In-Network Learning stage of a second virtual wireless network 400 comprising the plurality of virtual edge devices for collecting input data of the specific modality, each virtual edge device implementing the trained encoder ML model 111a-n, a virtual relay device implementing the relay ML model 121a, b, and a virtual fusion center device implementing a decoder ML model 131’ for input data of the specific modality.

[0052] In an embodiment, the wireless network 100 may be implemented in three phases, namely (a) off-line training of the encoder ML models 111a-c, the decoder ML model 131, and the relay ML models 121a, b, (b) ML model deployment on the distributed devices of the wireless network 100, and (c) on-line fine-tuning of the ML models implemented by the devices of the wireless network 100.

[0053] As already described above in the context of figure 3, in a first stage of the offline multi-modal training phase of the ML models a plurality of encoder ML models 111a-n, each given a different input modality, and the joint multimodal decoder ML model 131 are trained. In an embodiment, the encoder ML models 111a-n and / or the decoder  ML model 131 may be feed-forward neural networks, with the only requirement that the dimensionality of the decoder’s input should be a sum of all encoders’ output dimensionalities. The training data may have an identical modality and similar distribution as planned for deployment (e.g., images captured by an onboard camera of an autonomous vehicle, patients’ health records kept by a hospital, and the like) . Once the training is finished, the trained encoder ML models 111a-n and the decoder ML model 131 are retrieved and their weights are fixed.

[0054] As already described above in the context of figure 4a, once the plurality of encoder ML models 111a-n and the joint decode ML model 131 have been trained, a plurality of relay ML models 121a-n are trained (i.e., one for each modality) . In an embodiment, each relay ML model 121a-n may be a neural network equipped with a cross-view attention module, as will be described in more detail further below. The role of each relay ML model 121a-n is to fuse multiple latent descriptions of the same modality into one combined latent description of the same modality. The combined latent description gathers relevant information from all input latent descriptions, and keeps the same dimensionality as the inputs. Due to this latter property, it is possible to concatenate relay models of the same type, as illustrated in figure 4b.

[0055] As already described above, training of the relay ML models 121a-n may be done offline, separately for each of the plurality, i.e. K modalities and by emulating the multi-access topology 400 shown in figure 4a. Assuming some selected modality with a corresponding encoder ML model 111a-n has been obtained in the way described above, the emulated distributed network 400 consists of multiple copies of the same encoder ML model, one relay model to be trained, and a single-modality decoder to be discarded after training. During training, the training inputs of all encoder ML models have the same modality and are correlated (e.g., images of the same area taken from different angles, complementary text descriptions, and the like) . In this way, the relay model learns to summarize latent descriptions and output a single latent description will all relevant information from all encoders.

[0056] Thus, as will be appreciated, the virtual wireless network 400 shown in figure 4a contains several copies of one encoder model (by way of example the encoder ML model 111a) of the plurality of the encoder ML models 111a-n of figure 3. More specifically, performing the first training phase results in the K encoder ML models 111a-n of figure 3 (one per modality) . The weights of these encoder ML models 111a-n are fixed and the encoder ML models 111a-n with the fixed weights are used in the second training phase. The second training phase for training a respective relay ML model, such as the relay ML model 121a illustrated in figure 4a, is repeated independently K times (one training for each modality) . In other words, the virtual wireless network topology 400 shown in figure 4a is used K times, i.e. once for each modality. The result of the second training phase are relay ML models 121a-n (one for each modality) , such as the relay ML model 121a illustrated in figure 4a. One repetition of the second training phase is done in the following way. For a modality k (for instance the modality image data) with 0 < k <= K a corresponding encoder ML model from the first training phase is chosen (in this example, the encoder ML model 111a trained in the first phase for processing image data) and duplicated several times (for instance, three times) to obtain the virtual wireless network topology 400 of figure 4a. As already mentioned above, during the second training phase, the weights of the encoder ML models 111a-n remain fixed, i.e. the encoder ML models 111a-n do not change anymore in the second training phase.

[0057] As will be appreciated, the trained K encoder ML models 111a-n (one per modality) , the K relay ML models 121a, b (one per modality) , and the joint decoder ML model 131 may be used as building blocks to design arbitrary tree-type distributed topologies, such as the exemplary wireless network topology illustrated in figure 4c. The encoder ML models 111a-n of the same modalities may be duplicated multiple times, and the latent descriptions output by each encoder ML model 111a-n may be fused to a single stream using a respective relay ML model 121a, b. Moreover, as already described above in the context of figure 4b, the relay ML models 121a, b of the same modality may be concatenated. In this way, it is possible to merge multiple descriptions from multiple sources, and to obtain one unified stream of data per modality. In the last step, the obtained K single-modal streams are fused and processed by the joint decoder ML model 131.

[0058] As already described above, figure 4d illustrates an example of model deployment on a wireless network 100 with multiple edge devices 110a-n, arbitrarily connected intermediate relay devices 120a, b, and the fusion center device 130. The input of an edge device 110a-n could be a combination of K multimodal data. In that case, each modality may  be processed separately by a dedicated encoder ML model 111a-n. As already described above, the procedure for deploying the plurality of ML models on the different devices of the wireless network 100 illustrated in figure 4d is as follows. First, the one or more encoder ML models 111a-n are loaded to the edge devices 110a-n depending on the devices’ input modalities. Secondly, relay ML models 121a-n are loaded for each modality to the intermediate relay devices 120a, b and the fusion center device 130. Thirdly, the multi-modal joint decoder ML model 131 is loaded to the fusion center device 130. In this way, it can be ensured that different modalities are processed separately until the decoder.

[0059] As already mentioned above, after the actual deployment the relay ML models 121a-n may be fine-tuned end-to-end over the network 100 using In-Network Learning using small training datasets. The goal is to fine-tune the relay ML models 121a-n to better fuse latent descriptions by taking into account particularities of the network topology (e.g., positioning and number of edge devices, correlations of input data, and the like) . In this fine-tuning stage the weights of the encoder ML models 111a-n and of the decoder ML model 131 are preferably fixed.

[0060] For adapting to changes of the network topology, in an embodiment, each relay device 120a, b (as well as the other devices of the wireless network 100) may be configured to inform neighboring device, whenever the relay device 120a, b detects a change in the topology. In a negotiation phase neighboring devices of the wireless network may re-asses new channel links and their capacities. To this end, messages 510 may be exchanged between the different devices as illustrated in figure 5a. If necessary, some relay models 121a-n may be fine-tuned again to consider a different number of edge devices, or number of processing steps. Figure 5b illustrates an exemplary structure of the message 510 exchanged between the devices of the wireless network 100. The message 510 may contain an ID of the sender, a counter (or a timestamp) for synchronization, and a set of latent descriptions for each modality.

[0061] For allowing a relay ML model 121a-n to fuse multiple data streams of the same modality into one stream, the respective relay ML model 121a-n may use a cross-view attention mechanism, as illustrated in figures 6a and 6b. The cross-view attention mechanism implemented by a relay ML model 121a-n according to an embodiment exploits attention to find relations and semantic similarities among multiple descriptions. These detected similarities may be exploited to appropriately combine latent descriptions into one description, such that the useful information from all inputs is preserved. The procedure for performing cross-view attention implemented by a relay ML model 121a-n may be as follows (single head of a single layer) :

[0062] For each latent description Uj, compute a query Qj and a key Kj by using learned matrices WQ and WK: Qj=WQUj, Kj=WKUj

[0063] Compute cross-attention (scores) between all queries and keys using the following formula:

[0064] (c) Combine received latent descriptions Uj using computed attention scores:

[0065] (d) Repeat the above steps several times if multi-layer, multi-head cross-view attention is used (each time using dedicated matrices WQ and WK) .

[0066] Figure 7 shows a flow diagram illustrating a method 700 of operating, for instance, the relay device 120a of the wireless network 100, such as the wireless network 100 of figure 4d, including the plurality of edge devices 110a-n, the further relay device 120b and the fusion center device 130. The method 700 comprises the following steps executed by the relay device 120a. In a step 701 the relay device 120a implements, i.e. operates a relay machine learning, ML, model 121a configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of a specific modality, wherein the encoded input data of the specific modality is provided by the edge devices 110a, b of the plurality of edge devices 110a-n, wherein each of the plurality of edge devices 110a-n implements an encoder ML model 111a-n configured to generate the encoded input data of the specific modality. In a step 703 the relay device 120a provides the aggregated encoded input data of the specific modality to the further relay device 120b. In a further embodiment, the aggregated encoded input data of the specific modality may be directly provided to the fusion center device 130 implementing the decoder ML model 131 configured to generate  output data by decoding the aggregated encoded input data of the specific modality provided by the relay devices 120a, b and the edge devices 110a-n. As already described in detail above, the plurality of encoder ML models 111a-n and the decoder ML model 131 are based on In-Network Learning of the first virtual wireless network 300 illustrated in figure 3, wherein the first virtual wireless network 300 comprises a plurality of virtual edge devices implementing the plurality of encoder ML models 111a-n for collecting input data of different modality and a virtual fusion center device implementing the decoder ML model 131. As further described in detail above, the relay ML model 121a is trained with a second In-Network Learning stage of the second virtual wireless network 400 illustrated in figure 4a, wherein the second virtual wireless network 400 comprises a plurality of virtual edge devices for collecting input data of the specific modality, each virtual edge device implementing the trained encoder ML model 111a-n, a virtual relay device implementing the relay ML model 121a, and a virtual fusion center device implementing a decoder ML model 131’ for input data of the specific modality.

[0067] The person skilled in the art will understand that the "blocks" ( "units" ) of the various figures (method and apparatus) represent or describe functionalities of embodiments of the invention (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step) .

[0068] In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiment is merely exemplary. For example, the unit division is merely logical function division and may be other division in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.

[0069] The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.

[0070] In addition, functional units in the embodiments of the invention may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

Claims

1.A relay device (120a, b) for a wireless network (100) , the wireless network (100) comprising a plurality of edge devices (110a-n) , each edge device (110a-n) configured to collect input data of one or more modalities of a plurality of modalities, one or more relay devices (120a, b) and a fusion center device (130) , wherein the relay device (120a, b) is configured to:implement a relay machine learning, ML, model (121a-c) configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of the specific modality, wherein the encoded input data of the specific modality is provided by one or more of the plurality of edge devices (110a-n) , wherein each of the plurality of edge devices (110a-n) implements an encoder ML model (111a-n) configured to generate the encoded input data of the specific modality; andprovide the aggregated encoded input data of the specific modality to a further relay device (120a, b) or to the fusion center device (130) implementing a decoder ML model (131) configured to generate output data by decoding the aggregated encoded input data of the specific modality provided by the relay device (120a, b) ,wherein the plurality of encoder ML models (111a-n) and the decoder ML model (131) are based on a first In-Network Learning stage of a first virtual wireless network (300) , wherein the first virtual wireless network (300) comprises a plurality of virtual edge devices implementing the plurality of encoder ML models (111a-n) for collecting input data of different modality and a virtual fusion center device implementing the decoder ML model (131) ; andwherein the relay ML model (121a-c) is trained with a second In-Network Learning stage of a second virtual wireless network (400) comprising the plurality of virtual edge devices for collecting input data of the specific modality, each virtual edge device implementing the trained encoder ML model (111a-n) , a virtual relay device implementing the relay ML model (121a-c) , and a virtual fusion center device implementing a decoder ML model (131’) for input data of the specific modality.2.The relay device (120a, b) of claim 1, wherein the relay ML model (121a-c) comprises a cross-view attention module (121a-c) .3.The relay device (120a, b) of claim 1 or 2, wherein each encoder ML model (111a-n) comprises one or more feed-forward neural networks (111a-n) .4.The relay device (120a, b) of any one of the preceding claims, wherein the relay device (120a, b) is configured to inform one or more of the plurality of edge devices (110a-n) and / or the fusion center device (130) about a change of the topology of the wireless network (100) for triggering a re-training of the wireless network (100) with In-Network Learning.5.The relay device (120a, b) of any one of the preceding claims, wherein the relay ML model (121a-c) is configured to be further trained with a third In-Network Learning stage of the wireless network (100) comprising the plurality of edge devices (110a-n) for collecting input data of the specific modality, each edge device (110a-n) implementing the encoder ML model (111a-n) trained in the first In-Network Learning stage, the relay device (120a, b) implementing the relay ML model trained (121a, b) in the second In-Network Learning stage, and the fusion center device (130) implementing the decoder ML model (131) trained in the first In-Network Learning stage.6.A wireless network (100) , comprising:a plurality of edge devices (110a-n) ;one or more relay devices (120a, b) according to any one of the preceding claims; anda fusion center device (130) .7.A method (700) of operating a relay device (120a, b) of a wireless network (100) , the wireless network (100) comprising a plurality of edge devices (110a-n) , each edge device (110a-n) configured to collect input data of one or  more modalities of a plurality of modalities, one or more relay devices (120a, b) and a fusion center device (130) , wherein the method (700) comprises the following steps executed by the relay device (120a, b) :implementing (701) a relay machine learning, ML, model (121a-c) configured to aggregate encoded input data of a specific modality of the plurality of modalities into aggregated encoded input data of the specific modality, wherein the encoded input data of the specific modality is provided by one or more of the plurality of edge devices (110a-n) , wherein each of the plurality of edge devices (110a-n) implements an encoder ML model (111a-n) configured to generate the encoded input data of the specific modality; andproviding (703) the aggregated encoded input data of the specific modality to a further relay device (120a, b) or to the fusion center device (130) implementing a decoder ML model (131) configured to generate output data by decoding the aggregated encoded input data of the specific modality provided by the relay device (120a, b) , wherein the plurality of encoder ML models (111a-n) and the decoder ML model (131) are based on In-Network Learning of a first virtual wireless network (300) , wherein the virtual wireless network (300) comprises a plurality of virtual edge devices implementing the plurality of encoder ML models (111a-n) for collecting input data of different modality and a virtual fusion center device implementing the decoder ML model (131) , and wherein the relay ML model (121a-c) is trained with a second In-Network Learning stage of a second virtual wireless network (400) comprising the plurality of virtual edge devices for collecting input data of the specific modality, each virtual edge device implementing the trained encoder ML model (111a-n) , a virtual relay device implementing the relay ML model (121a-c) , and a virtual fusion center device implementing a decoder ML model (131’) for input data of the specific modality.8.A computer program product comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method (700) of claim 7 when the program code is executed by the computer or the processor.

Citation Information

Patent Citations

  • Methods and apparatus for multi-modal prediction using a trained statistical model

    CA3100065A1

  • Multi-task multi-modal machine learning model

    CN110574049A

  • Multimodal domain embeddings via contrastive learning

    US20230067528A1

  • System energy efficiency in a wireless network

    US20230188233A1