System and methods for fine-tuning of foundational models in communication networks
PEFT methods in network-based distributed learning address privacy and cost challenges by selectively updating and transferring relevant model parameters, achieving efficient and cost-effective fine-tuning of large models.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-10-23
- Publication Date
- 2026-04-30
AI Technical Summary
The challenge of fine-tuning large and complex foundational models like GPT series and diffusion models is hindered by privacy concerns and prohibitive computational and communication costs, with existing techniques failing to achieve the necessary efficiency for large-scale applications.
Implementing parameter-efficient fine-tuning (PEFT) methods in network-based distributed learning environments, utilizing adapters such as LoRA, DyLoRA, and prompt tuning, to selectively update and transfer relevant model parameters, reducing communication costs and computational overhead while preserving performance.
PEFT methods enable efficient fine-tuning of foundational models across multiple devices, reducing communication costs by up to 1000-fold, improving latency, and maintaining user data privacy.
Smart Images

Figure CN2024126781_30042026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHODS FOR FINE-TUNING OF FOUNDATIONAL MODELS IN COMMUNICATION NETWORKSTECHNICAL FIELD
[0001] The present disclosure pertains to the field of machine learning, and in particular to systems and methods for fine-tuning of foundational models in communication networks based on integrating parameter-efficient fine-tuning (PEFT) methods with network-based distributed learning.BACKGROUND
[0002] Foundation models, such as the generative pre-trained transformer (GPT) series and diffusion models, have shown significant potential across various fields, including, but not limited to, natural language processing and computer vision. However, fine-tuning these models for specific tasks, such as, for example, time-series forecasting, presents challenges due to the inaccessibility of user data driven by privacy concerns.
[0003] In addition, as models grow larger and more complex, the computational and communication costs associated with fine-tuning become increasingly prohibitive. Transferring large models or their parameters across multiple devices requires significant resources, with costs scaling based on a variety of variables which may include, but are not limited to, the number of parameters, users, and iterations involved.
[0004] While various techniques have been proposed to reduce these costs, such as model compression and parameter quantization, these techniques often lead to diminished performance or provide insufficient reductions. Current approaches fail to achieve the necessary efficiency for large-scale fine-tuning, highlighting the need for improved methods that balance performance, privacy, and cost.
[0005] Therefore, there is a need for systems and methods for fine-tuning of foundational models in communication networks that obviates or mitigates one or more limitations of the prior art.
[0006] This background information is provided to reveal information believed by the applicant to be of possible relevance to the present application. No admission is necessarily intended, nor should be construed, that any of the preceding information constitutes prior art against the present application.SUMMARY
[0007] Embodiments of the present application provides systems and methods for fine-tuning of foundational models in communication networks. According to embodiments, the method includes receiving a request including configuration instructions, and at least one module. The configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model. The method further includes sending at least one of the at least one updated module and the updated machine learning model.
[0008] In some embodiments, the sending includes sending, by the network node, at least one of the at least one updated module and the updated machine learning model to at least one other network node, wherein the request includes an indication of the at least one other network node.
[0009] In some embodiments, the configuration instructions indicate at least a portion of the machine learning model to which the at least one module is to be applied. In some embodiments, the configuration instructions are indicated in either a header or a payload of one or more packets, wherein the request includes the one or more packets.
[0010] In some embodiments, the one or more operations include one or more of training the machine learning model, evaluating the machine learning model and combining the at least one module with one or more other modules based on one or more of: a type of combination and one or more hyperparameters. The training of the machine learning model is based on one or more of: training data, a number of epochs, a learning rate, and a loss function. The evaluating of the machine learning model is based on one or more of: evaluation data and one or more metrics for evaluation.
[0011] In some embodiments, the at least one module is based on one of: Low-Rank Adaptation (LoRA) , dynamic Low-Rank Adaptation (DyLoRA) , prompt tuning, and prefix tuning.
[0012] In some embodiments, the request further indicates a data signal representing the at least one module and data related to the at least one module, wherein the one or more operations include combining, by the network node, the at least one module with a second module based on a similarity metric, wherein the similarity metric is determined using the data signal, and wherein the second module is either a local module or a module received from a second network node.
[0013] In some embodiments, the method further includes receiving from a plurality of other network nodes, a plurality of data signals, each data signal of the plurality of data signals corresponding to a network node of the plurality of other network nodes. The method further includes organizing the plurality of other network nodes into one or more groups, wherein the organizing is based on similarity metrics associated with the plurality of data signals and selecting an aggregator for each group of the one or more groups based on the similarity metrics.
[0014] In some embodiments, the method further includes sending to each network node of the plurality of other network nodes, a notification indicating a corresponding group of the one or more groups to which each network node belongs and a corresponding aggregator of the corresponding group to send at least one module of said each network node.
[0015] In some embodiments, the method further includes receiving, from the plurality of other network nodes, a plurality of modules. Each module of the plurality of modules corresponds to a network node of the plurality of other network nodes, wherein each data signal of the plurality of data signals represents at least one of a corresponding module of the plurality of modules and data related to the corresponding module.
[0016] In some embodiments, the organizing includes organizing the plurality of other network nodes into one or more groups based on similarity metrics of the plurality of modules, and the selecting includes selecting the aggregator for each group based on the similarity metrics of the plurality of modules. In some embodiments, the method further includes sending to each aggregator of each group of the one or more groups, one or more modules and corresponding one or more data signals of one or more network nodes belonging to each group.
[0017] According to embodiments, the method includes sending to a second network node, a request including configuration instructions and at least one module, wherein the configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model. The method further includes receiving from the second network node, at least one of the at least one updated module and the updated machine learning model.
[0018] In some embodiments, the method further includes receiving from a controller, the configuration instructions. In some embodiments, the configuration instructions further indicate the machine learning model to serve as a base model for fine-tuning and instructions for sending at least one of the at least one updated module and the updated machine learning model to the second network node.
[0019] According to another aspect, a (e.g. non-transitory) computer readable medium, computer program, or computer program product, comprising stored thereon statements and instructions which, when executed by a computer processor perform one or more methods described herein.
[0020] According to another aspect, an apparatus or system is provided, where the apparatus includes modules configured to perform one or more methods described herein. According to another aspect, another apparatus or system is provided that includes computing electronics and is configured to perform the methods described herein. According to another aspect, another apparatus is provided that includes processing and wireless communication electronics and is configured to operate as described herein.
[0021] According to another aspect, a method is provided for execution by processing and wireless communication electronics. The method includes performing operations as described herein. In some embodiments a computer program product is provided. The computer program product includes a non-transitory computer readable medium having recorded thereon statements and instructions which, when executed by a computer, cause the computer to perform one or more methods described herein.
[0022] According to another aspect, a chip is provided, where the chip includes a processor and a data interface, and the processor reads, by using the data interface, an instruction stored in a memory, to perform the different aspects described herein.
[0023] Other aspects of the application provide for apparatus, and systems configured to implement the methods according to the different aspects disclosed herein. For example, wireless stations and access points can be configured with machine readable memory containing instructions, which when executed by the processors of these devices, configures the device to perform the methods disclosed herein.
[0024] Embodiments have been described above in conjunction with aspects of the present application upon which they can be implemented. Those skilled in the art will appreciate that embodiments may be implemented in conjunction with the aspect with which they are described but may also be implemented with other embodiments of that aspect. When embodiments are mutually exclusive, or are incompatible with each other, it will be apparent to those skilled in the art. Some embodiments may be described in relation to one aspect, but may also be applicable to other aspects, as will be apparent to those of skill in the art.BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Further features and advantages of the present application will become apparent from the following detailed description, taken in combination with the appended drawings, in which:
[0026] FIG. 1 illustrates the use of a target_module information element in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0027] FIG. 2 illustrates the use of an applying_module (s) information element in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0028] FIG. 3 illustrates encoding options for information elements, according to an embodiment of the present disclosure.
[0029] FIG. 4A illustrates a combination operation in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0030] FIG. 4B illustrates another combination operation in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0031] FIG. 4C illustrates another combination operation in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0032] FIG. 4D illustrates a clustering process, according to an embodiment of the present disclosure.
[0033] FIG. 4E illustrates another clustering process, according to an embodiment of the present disclosure.
[0034] FIG. 5A illustrates the use of a base_model information element in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0035] FIG. 5B illustrates the application of a base_model information element in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0036] FIG. 6 illustrates the use of a call_back information element in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0037] FIG. 7 illustrates a receiving procedure at a network node from an application layer perspective, according to an embodiment of the present disclosure.
[0038] FIG. 8 illustrates a transmitting procedure at a network node from an application layer perspective, according to an embodiment of the present disclosure.
[0039] FIG. 9 illustrates a network architecture for enabling a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0040] FIG. 10 illustrates a procedure for implementing PEFT in a network-based distributed learning environment, according to an embodiment of the present disclosure.
[0041] FIG. 11A illustrates a method for fine-tuning a machine learning model in a distributed learning environment, according to an embodiment of the present disclosure.
[0042] FIG. 11B illustrates another method for fine-tuning a machine learning model in a distributed learning environment, according to an embodiment of the present disclosure.
[0043] FIG. 12A illustrates an example communication system, according to an embodiment of the present disclosure.
[0044] FIG. 12B illustrates another example communication system, according to an embodiment of the present disclosure.
[0045] FIG. 13 illustrates an apparatus communicating with another apparatus, according to an embodiment of the present disclosure. FIG. 14A illustrates an example apparatus, according to an embodiment of the present disclosure.
[0046] FIG. 14B illustrates another example apparatus, according to an embodiment of the present disclosure.
[0047] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION
[0048] Embodiments of the present application provides systems and methods for fine-tuning of foundational models in communication networks. According to an aspect, a method is provided for that includes receiving a request including configuration instructions, and at least one module. The configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model. The method further includes sending at least one of the at least one updated module and the updated machine learning model.
[0049] According to an aspect, another method is provided that includes sending to a second network node, a request including configuration instructions and at least one module, wherein the configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model. The method further includes receiving from the second network node, at least one of the at least one updated module and the updated machine learning model.
[0050] Federated learning (FL) refers to a decentralized machine learning framework that learns global knowledge from a network of clients without sharing their private data. Federated averaging (FedAvg) is a widely adopted FL approach that shares knowledge through the federation of devices by averaging the parameters of their models. FedAvg collects models from all participating devices, averages the parameters, and creates a global model. This global model is then communicated back to the participating clients.
[0051] In recent years, foundation models (e.g., GPT series, diffusion models) have demonstrated their effectiveness across a diverse set of tasks in natural language understanding, computer vision, and other fields.
[0052] To enhance performance on specific downstream tasks, foundation models require fine-tuning. For instance, a foundation model used for time-series forecasting should be fine-tuned with network traffic data. However, due to privacy concerns, direct access to user data and specific downstream tasks is not feasible. Consequently, a solution is to collaboratively train models by sharing only model parameters, rather than raw data, using distributed learning methods.
[0053] Distributed learning involves frequent transfer of model parameters among users, which can result in significant communication costs. The total communication cost of decentralized learning scales with variables that can include, but are not limited to: (a) the number of model parameters, (b) the number of participating users, and (c) the number of communication rounds. As the number of model parameters and active users increases, these costs tend to rise. For example, the total communication cost for 10 clients sending Llama2-7B to a server over 100 rounds of FL, amounts to 28 TB, which is quite large and could benefit from a reduction by a factor of at least 1000 to be practical.
[0054] To address this, exploring communication-efficient methods for fine-tuning foundational models in distributed learning may be beneficial for various reasons. Some benefits may include being able to use the capabilities of foundational models and reduce communication costs. Further benefits may include accommodating the transfer of large models and enabling participation from a greater number of users. The benefits may further include improving latency and learning time across users and preserving user data privacy.
[0055] The communication cost of distributed learning scales with nparameters*nbits / param*nusers *nrounds. Various techniques have been implemented to reduce the cost associated with each component. These techniques include decreasing the model size to reduce the number of parameters, nparameters. To reduce the number of bits representing each parameter, nbits / param, techniques such as quantization and mixed precision training have been employed. Some techniques include reducing the number of participants, nusers, by using methods such as partial participation and scheduling. To reduce the number of communication rounds, nrounds, fast-convergence methods such as FedYogi have been employed.
[0056] Existing techniques used to address communication costs in distributed learning each have their own drawbacks. For reducing the number of parameters (nparameters) , techniques involving using smaller models can lead to decreased model performance. Further, methods aimed at reducing the number of bits per parameter (nbits / param) can achieve up to a 4-fold reduction but may still fall short of the necessary cost reduction. Techniques used to address the number of participating users (nusers) can achieve up to 10 times reduction. Nevertheless, these techniques require additional information and may not fully meet the required reduction goals. Similarly, methods used to reduce the number of communication rounds (nrounds) can achieve up to a 10-fold reduction. However, these methods introduce additional computational overhead on clients, which can offset some of the benefits. While each method provides some level of reduction, none come close to the 1000-fold reduction which may be needed for practical implementation.
[0057] According to some embodiments of the present disclosure, parameter-efficient fine-tuning (PEFT) methods are employed in network-based distributed learning to train a subset of parameters and transmit the relevant portion of model parameters. For example, the relevant portions of the models may be one or more: of a subset of model parameters, add-on modules and configuration instructions. PEFT methods include, but are not limited to, adapters (e.g., LoRA, DyLoRA) , prefix tuning, prompt tuning, partial fine-tuning, and other appropriate methods. Some embodiments may reduce communication costs and lower the computation and memory requirements for clients. Some embodiments may achieve performance levels comparable to full fine-tuning.
[0058] Some embodiments may combine PEFT with network-based distributed learning. These embodiments may provide for protocols, information elements and training procedures to accommodate PEFT in network settings, which may enable integration of PEFT with network-based decentralized learning.
[0059] Some embodiments may provide for the use of PEFT in distributed learning environments with model heterogeneity. Some embodiments may offer guidance on how clients should manage adapters received from the network.
[0060] Some embodiments may provide protocols and information elements to enable PEFT in network-based distributed learning environments. According to an embodiment, an information element indicative of target_module (s) , is provided to indicate one or more targets, parts, or portions of a base model (e.g., a machine learning model) to which a PEFT module is to be applied. For example, the target_module (s) information element may indicate one or more of the following: fully connected (FC) layers, convolutional layers, attention layers, embeddings, etc.
[0061] According to an embodiment, an information element indicative of replacement_module (s) or applying_module (s) , is provided to indicate one or more PEFT modules to be applied for fine-tuning the base model. For example, the applying_module (s) information element may indicate one or more of following: LoRA, prompt tuning, prefix tuning, or other PEFT-based modules.
[0062] According to an embodiment, an information element indicative of procedure (s) , is provided to indicate one or more operations involving the PEFT module to be performed. In some embodiments, the information element indicative of procedure (s) , indicates the type of operation (s) to perform, such as training, evaluating, or combining. In some embodiments, for training, the procedure (s) information element indicates details such as which part of the model to train, the data to be used for training, the number of epochs, the learning rate, and / or the loss function. In some embodiments, when evaluating the model, the procedure (s) information element indicates the data used for evaluation and the metrics by which the model is assessed.
[0063] In some embodiments, in the case of combination operations, the procedure (s) information element includes the type of combination method, such as averaging, clustering, ties merging, DARE merging, learnable weighted averaging, etc. In some embodiments, the procedure (s) information element further includes hyperparameters needed for the combination method. In some embodiments, the procedure (s) information element, indicates one or more operations for applying the PEFT module to obtain an updated module.
[0064] According to an embodiment, an information element indicative of base_model, is provided to indicate the base model that should be used as foundational model for fine-tuning. For example, the base_model information element may specify a particular machine learning model, such as Llama2 or Llama3, that is requested by the client to serve as the foundation for fine-tuning. In some embodiments, base_model information element is used to request a base model to fine-tune (heterogenous) models.
[0065] According to an embodiment, an information element indicative of call_back, is provided to indicate one or more operations or actions to be performed after performing the one or more operations indicated by procedures. For example, the call_back information element may instruct that the aggregated model be sent back to two specific users in the network.
[0066] The one or more information elements described herein may facilitate using PEFT in network-based distributed learning.
[0067] Some embodiments may provide a learning framework where clients (e.g., network nodes) transfer PEFT modules along with auxiliary information elements to allow distributed PEFT. In this framework, each participating node may perform various operations to facilitate the learning process. For instance, a node may perform a transmission operation such as sending, receiving, requesting, or routing PEFT modules or base models using information elements such as applying_module (s) information element (s) , target_module (s) information element (s) , call_back information element (s) , and / or base_model information element (s) .
[0068] In some embodiments, a node combines multiple PEFT modules according to instructions in the procedure (s) information elements. In some embodiments, a node fine-tunes a base models using one or more PEFT modules according to instructions in procedure (s) information element.
[0069] In some embodiments, a node evaluates base models and PEFT modules according to instructions in the procedure (s) information elements. In some embodiments, a node saves or caches a base model or PEFT module to save communication costs according to base_model information elements.
[0070] FIG. 1 illustrates the use of target_module information element in a network-based distributed learning environment, according to an embodiment. In FIG. 1, a base model 102, such as a transformer, is depicted with various components, including the linear layer 104, feed-forward layer 106, masked multi-head attention layer 108, and embedding portion. In some embodiment, a PEFT module 110, such as LoRA, can be applied to one or more targets or parts of the base model. In some embodiments, a network node or a node, which could be a server 120, is capable of transmitting this PEFT module 110 to client 130, and indicating one or more targets or portions of the base model (e.g., linear, feed-forward, and multi-head attention layers) to which the PEFT module should be applied using the target_module (s) information element.
[0071] The target_module (s) information element may enable selective fine-tuning of certain components of the base model which may reduce the computational and communication overhead. For instance, the target_module information element allows the node 120 to specify that the PEFT module is relevant only to the linear, feed forward, and multi-head attention layers. This information is communicated between nodes (such as between a client and server) to ensure that the PEFT module is applied efficiently to the intended portions of the base model, streamlining the distributed fine-tuning process.
[0072] FIG. 2 illustrates the use of applying_module (s) information element in a network-based distributed learning environment, according to an embodiment. As described herein the applying_module (s) information element is used to specify which PEFT module (s) is being applied according to the procedure (s) information element. As illustrated, network nodes (e.g., a client 202 and a server 204) , can communicate with each other, using applying_module (s) information element, one or more PEFT modules, such as prefix tuning 206, prompt tuning 208, and LoRA 210 among others.
[0073] In some embodiments, the applying_module (s) information element includes or is accompanied by additional information elements necessary for the proper application of the PEFT module according to the procedure (s) information element (e.g., to the base model) . For example, for prefix tuning 206, the additional information elements may include one or more of: the base model, the number of tokens (N_tokens) , the token dimensions (Token_dim) , and the embedding. In the case of prompt tuning 208, additional elements may include one or more of: the base model, N_tokens, token dimensions, initial text, and the embedding. For LoRA, the additional information elements may include one or more of: the base model, rank, scale, matrices A and B, and sigma. These information elements may ensure that the PEFT modules are applied properly to the appropriate base model, with each element serving a role in the fine-tuning process. For instance, the rank and scale values in LoRA allow for control over how much influence the PEFT module has on the base model, while in prompt tuning, initial text and embedding define the context within which the model operates. The information exchanged between the nodes may ensure that these PEFT modules are appropriately configured and applied in network-based distributed learning environments.
[0074] FIG. 3 illustrates encoding options for information elements, according to an embodiment. FIG. 3 presents a packet format 300 including a header 302 and a payload 304. As may be appreciated, the packet format 300 is an example and may refer to other packet types and formats used within the network.
[0075] In some embodiments, one of both of the target_module (s) information element and applying_module information elements can be encoded into one or more fields of the packet header 302. Encoding in the header may allow for quick interpretation and use of these information elements as the packet is processed by network nodes.
[0076] In some embodiments, one of both of the target_module (s) information element and applying_module (s) information element can be encoded within the payload 304 of the packet, which may allow for more flexibility in terms of the data size and structure.
[0077] As may be appreciated, both encoding strategies (either in the header or the payload) may provide flexibility depending on the network’s needs. For faster processing and lower latency, encoding in the header may be preferred, as it enables network nodes to quickly interpret the information and apply the PEFT modules. Alternatively, encoding in the payload can support more detailed information exchange, offering a broader scope for data and control in larger-scale networks or more complex operations.
[0078] In some embodiments, the procedure (s) information element includes data signal information that reflects or represents the data of a user or a node for data-aware decision making (e.g., clustering, weighted combination etc. ) . In some embodiments, the data signal information represents a PEFT modules or characteristics of the PEFT module of the node, which may have been updated based on the data of the node. In some embodiments, data signal information may provide information about extracted features.
[0079] In some embodiments, the data signal information provides information about similarity metrics (e.g., centered kernel alignment (CKA) ) to determine how a PEFT module should be combined with one or more other PEFT modules (e.g., a local PEFT module) . In some embodiments, the data signal information is used in adjusting or weighting the combination of PEFT modules. In some embodiments, the adjusting or weighting the combination of PEFT modules is based on a similarity metric between data from different nodes, which may improve fine-tuning by guiding how much influence a received PEFT module may have during a combination process.
[0080] In some embodiments, a similarity metric, based on the data signal information, is used to assess how closely a received PEFT module aligns with the local data, allowing for data-aware decision-making and adjusting the influence of the received PEFT module during combination.
[0081] FIG. 4A illustrates a combination operation in a network-based distributed learning environment, according to an embodiment. The combination operation 400 combines a received PEFT module 402 with a local module 404 to obtain an updated PEFT module 406 in a distributed learning environment. According to an embodiment, a network node 412 sends a PEFT module 402 (e.g., LoRA) and a data signal information 410, denoted as function f, to network node 414. In some embodiments, the data signal information is sent as part of the procedure (s) information element or independently.
[0082] In some embodiments, the data signal information f 410 represents information about the data of network node 412. In some embodiments, the data signal information f 410 represents characteristics of the PEFT module itself 402.
[0083] In some embodiments, after receiving the PEFT module 402 and the data signal information f 410, node 414 uses the received PEFT module 402 with its own locally stored PEFT module 404 to perform a combination operation 400.
[0084] The combination operation 400, illustrated FIG. 4A as a weighted sum, allows the two PEFT modules to be merged, resulting in an updated local PEFT module 406. The data signal information f 410 may help node 414 determine the relevance and applicability of the received PEFT module. For example, the data signal information f 410 can be understood as a weight that adjusts how much influence the received PEFT module should have during the combination process. In some embodiments, similarity metrics, such as CKA, are used to compute an alpha value (α) that modulates the combination. In some embodiments, the alpha value is calculated by comparing the extracted features from both nodes based on the similarity metrics, and the data signal information f 410. In some embodiments, the alpha value determines how the two PEFT modules are weighted when combined.
[0085] The use of similarity metrics may ensure that the combination is based on how similar the data or features between the two nodes are, which may improve the integration of the PEFT modules. For example, if the similarity between the data at node 412 and node 414 is high, the received PEFT module from node 412 will be weighted more heavily in the combination. Based on the similarity, a weight (α) is calculated that determines how much the local PEFT module 404 should be combined with a received model PEFT module 402. The combination may then result in a new PEFT module 406 that balances the local and received information, enhancing the model’s adaptation based on data similarity. This approach may allow for adaptive fine-tuning across distributed nodes, taking into account both local and remote data signals information to improve the performance of the combined model.
[0086] FIG. 4B illustrates another combination operation in a network-based distributed learning environment, according to an embodiment. In this scenario, network node 412 transmits its PEFT module 402 along with data signal information f 410 to another network node 414, where a combination operation 420 is performed. In some embodiments, the combination operation 420 is conditional based on satisfying a similarity threshold or condition 421. For example, to determine whether combination operation 420 will proceed, in some embodiments, network node 414 checks the similarity between its local PEFT module 404 and the received PEFT module 402 by comparing the data signal information f 410 to determine whether condition 421 is satisfied. If network node 414 determines that the condition 421 is satisfied (e.g., the similarity threshold is met) , then the network node 414 performs the combination operation 420 involving the received PEFT module 402 and the local PEFT modules 404 to obtain the updated PEFT module 406.
[0087] Using a similarity condition or threshold may ensure that PEFT modules are combined only when they are sufficiently similar, preventing the mixing of incompatible or irrelevant models. This approach may improve the distributed learning process by ensuring that only compatible modules contribute to the fine-tuning of the local model.
[0088] FIG. 4C illustrates another combination operation in a network-based distributed learning environment, according to an embodiment. According to an embodiment, a receiver node 432 receives multiple PEFT modules 434, 436, and 438 from corresponding sender nodes 435, 437, and 439. In some embodiments, the receiver node 432 uses training data 440 (which may be local) to train an initial set of weights W0, which may serve as the starting point for a base model’s parameters. For example, in some instances W0 can be associated with a pre-trained model currently existing on the device or in some instances W0 can be received from the user.
[0089] In some embodiments, training involves applying combination weights (or scaling factor) which may be referred to as alphas (αs) to each of the received PEFT modules. In some embodiments, the alphas (αs) may be calculated or determined locally or may be received from other nodes associated with the training. Each of the received PEFT modules may be assigned a corresponding alpha (e.g., α1, α2, α3) to control how much influence that module has in the combination process, with a corresponding weight (α) for each PEFT module as illustrated. In some embodiments, the receiver node uses these alphas to combine the contributions from the received PEFT modules with its local data, thereby learning and obtaining an updated set of weights W1.
[0090] As may be appreciated, the purpose of the initial weights W0 is to provide a foundation or a starting point for the base model, which is then adjusted by the contributions from the PEFT modules. These weights typically come from the base model being used for training, before the PEFT modules are applied. Accordingly, the weights W0 and W1 are part of the model parameters that define how the model makes predictions. The initial weights W0 are the starting parameters, while W1 represents the model after incorporating new knowledge or adjustments from the received PEFT modules. The alphas help to balance the contribution or impact of each module to ensure that the model update reflects the relevant data.
[0091] In some embodiments, after the training process, the updated weight W1 can be evaluated to assess the effectiveness of the combination operation. In some embodiments, the updated weights W1 are transmitted back to the sender nodes 435, 437, and 439. This combination and evaluation process may allow for efficient and adaptive learning across multiple network nodes.
[0092] FIG. 4D illustrates a clustering process, according to an embodiment. In some cases, data signal information, f, is used for clustering or grouping purposes, where the receiving node uses the data signal information to group nodes based on similarities in their data or PEFT modules. After clustering, the combination operation can be performed, allowing nodes within the same cluster to combine their PEFT modules more effectively. Thus, the data signal information plays a versatile role in the procedure (s) information element, guiding both clustering and combination operations in the distributed learning environment.
[0093] In some embodiments, a centralized controller 455 (which is a network node such as a 5G or 6G base station) uses data signal information f and PEFT modules gathered from multiple network nodes 451, 452, 453, and 454 to cluster the nodes into groups based on similarity metrics. According to an embodiment, the clustering process 450 includes at step 461, the centralized controller 455 receiving the data signal information f and PEFT modules from the one or more network nodes, such as nodes 451, 452, 453, and 454. In some embodiments, at step 462, the centralized controller 462 clusters or groups the nodes into clusters or groups based on similarity metrics using the data signal information f and the PEFT modules. For example, the centralized controller forms two clusters or groups: group 457, which includes nodes 451 and 453, and group 458, which includes nodes 452 and 454. The centralized controller then selects an aggregator for each group, with node 451 serving as the aggregator for group 457, and node 452 acting as the aggregator for group 458. It is to be understood that the selection of an aggregator for each group may be based on or associated with signal information received by the centralized controller.
[0094] According to an embodiment, once the groups and aggregators are established, the centralized controller 455, at step 463 sends the PEFT modules and data signal information f of each group member to the selected aggregator. For example, for group 457, the centralized controller 455 sends, to node 451 (the aggregator for group 457) , the PEFT modules and data signal information f of node 453 and node 451. Similarly, for group 458, the centralized controller sends to node 452 the PEFT modules and data signal information f of node 454 and node 452.
[0095] According to an embodiment, after receiving the PEFT modules and data signal information f, each aggregator at step 464 performs the combination operation based on the received PEFT modules and data signal information. This grouping process may allow the network to efficiently manage and combine PEFT modules in a way that improves the distributed learning by grouping nodes with similar data and model characteristics.
[0096] FIG. 4E illustrates another clustering process, according to an embodiment. In some embodiments, a centralized controller 455 is responsible for managing the clustering and decentralized routing between network nodes. According to an embodiment, the clustering process 470 includes at step 471, the centralized controller 455 receiving data signal information f from multiple nodes, such as nodes 451, 452, 453, and 454. In some embodiments, the process 470 further includes at step 462, the centralized controller 455 using similarity metrics applied to the data signal information f to cluster the nodes into groups. For example, the centralized controller 455 may form two groups: group 477 including nodes 451 and 453, and group 478 including nodes 452 and 454. For each group, the centralized controller 455 selects an aggregator. For example, node 451 is selected as the aggregator for group 477, and node 454 is selected as the aggregator for group 478.
[0097] After clustering the nodes, process 470 includes at step 473, the controller informing each node of its group and the selected aggregator for the group. For example, node 453 is informed that it belongs to group 477 with node 451 as the aggregator, while node 452 is informed that it is part of group 478 with node 454 as the aggregator. In some embodiments, node 451 may be notified that it belongs to group 477 as the aggregator. Similarly, node 454 may be notified that it belongs to group 478 as the aggregator.
[0098] Following this, in some embodiments, at step 473, each non-aggregator node sends its PEFT module to the aggregator of its group. For instance, node 453 sends its PEFT module to node 451, and node 452 sends its PEFT module to node 454. Once the PEFT modules are received, at step 474, the aggregators perform the combination operation based on the received modules. This decentralized approach may allow for efficient routing and clustering, ensuring that nodes with similar characteristics are grouped together, and the combination of PEFT modules is improved within each group.
[0099] FIG. 5A illustrates the use of the base_model information element in a network-based distributed learning environment, according to an embodiment. The figure shows a setup 500 where a base station 502 receives messages from sender nodes 510 and 520, the messages including a base_model information element 514 and 524, respectively. The base station 502 holds one or more base models, represented here as base models 504 and 506. According to an embodiment, the base_model information element 514 specifies that base model 504 is to be used for as the base model or foundation model for fine-tuning, while base_model information element 524 specifies that base model 506 should be used as the base model or foundational model for fine-tuning.
[0100] The use of the base_model information element allows the base station 502 to determine which base model to apply a received PEFT modules to, ensuring proper fine-tuning based on the sender node’s requirements. This embodiment provides flexibility and precision in the selection of base models for different sender nodes.
[0101] FIG. 5B illustrates the application of the base_model information element in a network-based distributed learning environment, according to an embodiment. FIG. 5B shows the base station 502 receiving PEFT modules 516 and 518 from sender node 510, and PEFT module 526 from sender node 520. These PEFT modules may have been updated based on the local data 512 from sender node 510 and local data 522 from sender node 520, respectively.
[0102] According to an embodiment, based on the base_model information elements 514 and 524 (as shown in FIG. 5A) , the base station 502 applies PEFT modules 516 and 518 to base model 504, and applies PEFT module 526 to base model 506. The server data 508 accessible by the base station 502 may be used in applying the received PEFT modules to the corresponding base models. For example, the received PEFT modules may be applied to the corresponding base models and trained using server data 508.
[0103] This process may allow for customized fine-tuning of each base model based on the sender node's local data and transmitted PEFT modules. This embodiment highlights the dynamic nature of distributed learning by allowing multiple nodes to contribute updates to different base models held by the base station.
[0104] FIG. 6 illustrates the use of the call_back information element in a network-based distributed learning environment, according to an embodiment. The figure shows an aggregator node 602 receiving call_back information elements 611, 613, and 615 from sender nodes 610, 612, and 614, respectively. The call_back information element indicates one or more actions to be taken by the aggregator after performing operations 604, which may be based on the procedure (s) information element.
[0105] For instance, each call_back information element may instruct the aggregator 602 to return the updated PEFT module 606 to the sender nodes after performing the operations 604. Once the operations are completed, the updated PEFT module 606 is transmitted back to the sender nodes 610, 612, and 614. This embodiment may ensure that after fine-tuning or other operations, the results are communicated back to the contributing nodes, allowing for continuous updates in the distributed learning process.
[0106] FIG. 7 illustrates a receiving procedure at a network node from an application layer perspective, according to an embodiment. In some embodiments, the procedure 700 is carried out by one or more modules within the node, including a communication module 720, a storage module 730, and a processing module 740.
[0107] According to an embodiment, procedure 700 includes the communication module 720 receiving a message 701 including one or more information elements such as base_model and / or applying_module (s) . In some embodiments, the received base model and / or PEFT modules are buffered 702 by the communication module and then stored 703 in the storage module 730 for future use. In some embodiments, the communication module 720 receives information elements 704, such as procedure (s) , target_module (s) , and / or call_back, which are buffered 705 by the communication module 720 and stored 706 in the storage module.
[0108] In some embodiments, the processing module 740 requests and retrieves 707 these information elements (base_model, applying_module (s) , target_module (s) , and / or procedure (s) ) from the storage module to perform one or more operations. In some embodiments, the processing module 740 performs a combination operation 708, which may be based on the retrieved applying_module (s) and procedure (s) . In some embodiments, the processing module 740 requests and retrieves local data 709 from the storage module to perform training 710 and / or evaluation 711 based on the retrieved information elements. In some embodiments, once these operations are completed, the results are stored 712 in the storage module 730. The results may include an updated PEFT module for example. In some embodiments, the storage module may then send a message 713 to the communication module, including the updated PEFT module and call_back information. In some embodiments, the communication module 720 then transmits 714 the updated PEFT module according to the instructions in the call_back information element.
[0109] It should be noted that the operations and order of operations within procedure 700 may vary from the illustrated option, which may provide flexibility in how the node processes information and performs its functions.
[0110] FIG. 8 illustrates a transmitting procedure at a network node from an application layer perspective, according to an embodiment. In some embodiments, the procedure 800 is carried out by one or more modules within the node, including a communication module 720 and a storage module 730.
[0111] According to an embodiment, procedure 800 includes the communication module 720 sending a request message 801 for one or more information elements such as base_model information element and / or applying_module (s) information element. In some embodiments, the communication module 720 sends a request message 802 for other information elements, such as target_module (s) information element and call_back information element. In some embodiments, following the transmission of these requests, the communication module 720 receives 803 one or more base models, PEFT modules and configuration instructions or information, which may be based on the requested information elements.
[0112] Once received, in some embodiments, the communication module buffers 804 the received base models and PEFT modules and stores 805 them in the storage module 730 for future use. It should be noted that, like FIG. 7, the operations within procedure 800 do not have to follow this specific order as different sequences of operations can occur, allowing for a flexible and adaptable approach to handling network transmissions in a distributed learning environment.
[0113] Embodiment including procedure 700 and 800 may allow for transmission and reception of PEFT modules and information elements, enabling communication and data flow within the distributed learning system.
[0114] FIG. 9 illustrates a network architecture for enabling a network-based distributed learning environment, according to an embodiment. The network architecture 900 can be implemented in future generation communication networks such as, for example, 5G or 6G networks, enabling efficient machine learning (ML) while preserving user privacy. According to an embodiment, the architecture 900 includes a control plane 910 and a data plane 920.
[0115] In some embodiments, in the control plane 910, a network controller 912 manages network operations and coordinates the distributed learning process. The data plane 920 may include one or more data plane functions (DPFs) 921, which may be responsible for managing data processing tasks, routing data traffic, and facilitating communication between network nodes in the distributed learning environment. In some embodiments, the network architecture 900 includes a processing service controller (PSC) 906, which manages, controls and influences the behavior of different types of processing service functions (PSF) . In some embodiments, the PSC interfaces with the network controller 912 of the control plane 910 and the PSFs (type-1 PSF 902 and type 2 PSF 904) .
[0116] According to an embodiment, the network architecture 900 supports different types of PSFs including type-1 PSF 902 and type-2 PSF 904. In some embodiments, type-1 PSF 902 includes both an AI model and a local dataset, for example, a wireless terminal device such as a user equipment (UE) may be a type-1 PSF. In some embodiments, type-1 PSF interfaces with the control plane 910, the network controller 912, the data plane 920, and the PSC 906. In some embodiments, type-2 PSF 904 does not contain an AI model but can train an AI model. For example, a wireless terminal device such as a server or other backend system may be a type-2 PSF. In some embodiments, the type-2 PSF 902 interfaces with one or more of: the data plane 920 and the PSC 906 as shown.
[0117] The network architecture 900 may enable PEFT in network-based distributed learning environment. The network architecture 900 may reduce communication costs in decentralized learning. According to an embodiment, the architecture 900 is mapped to a 6G network architecture, e.g., the 6G X-centric network architecture. In some embodiments, the network controller 912 is mapped to a mobile control function (MCF) of the 6G network architecture. In some embodiments, the PSC 906 is mapped to a task control function (TCF) in a NET4AI module in the 6G network architecture. It is to be understood that a TCF is at least in part responsible for optimizing NET4AI decisions. In some embodiments, the type-1 PSF 902 is mapped to a UE, an application server (AS) or a network function (NF) in the 6G network architecture. In some embodiments, the type-2 PSF 904 is mapped to a processing service function (PSF) in the Net4AI module of the 6G network architecture.
[0118] Embodiments may be implemented within the 6G X-Centric network architecture, which is a next-generation framework designed to provide a user-centric and context-aware communication experience. This architecture is based on real-time adaptability, dynamically adjusting network resources and services based on user behavior, environmental conditions, and specific application needs. The X-Centric network architecture integrates artificial intelligence (AI) and machine learning (ML) for optimized decision-making and automation across network operations, while edge computing reduces latency by processing data closer to the user. Furthermore, the architecture supports a variety of communication technologies, such as Terahertz (THz) and visible light communication (VLC) , to enable ultra-high-speed, high-capacity, and low-latency connectivity for use cases like extended reality (XR) and smart cities.
[0119] A component of this X-Centric network architecture is the NET4AI module, which introduces network intelligence through AI-driven management of resources, traffic, and performance. This module may allow for end-to-end AI integration, facilitating real-time analysis and decision-making to optimize network functions autonomously. The NET4AI module enhances the network's ability to self-manage and respond dynamically to user demands, reducing the need for manual intervention. In this context, the embodiments described herein may be applied to the 6G X-Centric architecture, leveraging the NET4AI module to achieve efficient and intelligent network operations.
[0120] According to an embodiment, the network architecture 900 is mapped to a 5G network architecture. In some embodiments, the network controller 912 is mapped to a session management function (SMF) in the 5G network architecture. In some embodiments, the PSC 906 is mapped to an Application Function (AF) of the 5G network architecture. In some embodiments, the type-1 PSF 902 is mapped to a UE of the 5G network architecture. In some embodiments, the type-2 PSF 904 is mapped to an AS of the 5G network architectures.
[0121] FIG. 10 illustrates a procedure for implementing PEFT in a network-based distributed learning environment. The procedure is based on the network architecture 900, where various components work together to facilitate learning and communication between nodes.
[0122] According to an embodiment, procedure 1000 includes performing operations 1002 to request for influencing traffic routing, such as selecting an aggregator node. In some embodiments, operations 1002 include selecting type-1 PSFs 903 and 905, as well as type-2 PSF 904. The selection criteria may be based on factors such as location, task, or application. Location can refer to the physical proximity of the PSFs, which may reduce communication delays, while task or application refers to the specific operations each node is best suited to handle.
[0123] In some embodiments, the PSC 906 sends a control command 1004 to type-2 PSF 904. In some embodiments, the control command 1004 identifies the type-1 PSFs 903 and 905 based on assigned identifiers (IDs) . In some embodiments, the identification of the type-1 PSFs 903 and 905 may be used to facilitate or enable the receipt of information 1012 from these nodes.
[0124] In some embodiments, the control command 1004 includes PEFT configurations which includes one or more information elements for distributed learning. The one or more information elements include target_module (s) , applying_module (s) , procedures, and base_model. In some embodiments, the PEFT configurations is used for one or more processing operations 1014.
[0125] In some embodiments, the control command 1004 includes error-handling instructions for addressing issues such as transmission problems issues, training issues, and received model issues. These error-handling instructions may be applicable when receiving information 1012 from the type-1 PSFs and when sending notification 1016 to, for example, report status or errors.
[0126] In some embodiments, the PSC 906 configures, by sending type-2 PSF configuration instructions 1006 to, each type-1 PSFs 903 and 905. The configuration instructions 1006 include identification information indicating assigned IDs for the type-1 PSFs 903 and 905. In some embodiments, the assigned IDs are used to facilitate or enable the sending of information 1012.
[0127] In some embodiments, configuration instructions 1006 include PEFT configurations which include one or more information elements, such as target_module (s) information element, applying_module (s) information element, procedures information element, and base_model information element. In some embodiments, the PEFT configurations in configuration instruction 1006 is similar to or sometimes the same as the PEFT configurations provided in the control command 1004. In some embodiments, the PEFT configurations are used to facilitate processing operations 1008 and 1010 by the type-1 PSFs 903 and 905, respectively. In some embodiments, the configuration instructions 1006 include transmission configurations, which include one or more of: ID and / or address of the type-2 PSF 904, and data plane configuration. In some embodiment, the transmission configurations are used to facilitate the sending of the information 1012.
[0128] In some embodiments, each of the type-1 PSFs 903 and 905 processes 1008 and 1010 modules based on PEFT configuration received in the configuration instructions 1006. In some embodiments, processing 1008 and 1010 modules include performing one or more operations indicated by the procedure (s) information element such as training, evaluating, and / or fine-tuning.
[0129] In some embodiments, each of the type-1 PSFs 903 and 905 sends information 1012 to the type-2 PSF 904 via the data plane 920. The information 1012 may include one or more of the PEFT configurations, configuration instructions and results of the processed modules. In some embodiments, type-2 PSF 904 performs one or more operations 1014 based on PEFT configuration including procedure (s) information element. For example, type-2 PSF 904 may combine PEFT modules received from the type-1 PSFs 903 and 905.
[0130] In some embodiments, type-2 PSF 904 sends a notification 1016 to PSC 906. This notification may confirm that the control command 1022 has been executed successfully or report errors for future iterations. In some embodiments, the PSC 906 selects or reselects 1018 a type-2 PSF again and performs the steps indicated by 1002 to 1016 again in an iterative fashion. In some embodiments, procedure 1000 further includes performing operations 1020 to request for influencing traffic routing (receiver selection) .
[0131] In some embodiments, PSC 906 sends a control command 1022 to type-2 PSF 904. The control command 1022 may include transmission configuration for sending updated information 1024. In some embodiments, the transmission configuration includes one or more of: traffic configuration, ID and / or address of the receiving type-1 PSFs. In some embodiments, the control command 1022 includes error messages which may be used for future iterations for resolution thereof. It is to be understood that the above details and operations can provide for the implementation of the instant application in future generation wireless communication network architecture.
[0132] In some embodiments, type-2 PSF 904 sends updated information 1024 to type-1 PSFs 903 and 905 via the data plane 920. In some embodiments, the updated information 1024 includes PEFT data indicative of applying_module (s) information element and PEFT_model information element. In some embodiments, each of the type-1 PSFs 903 and 905 sends a notification 1026 to the PSC 906 indicating the status of the received models, which includes the PEFT modules, any associated errors and in some instances may include the base model. In some embodiments, type-1 PSFs 903 and 905 process 1028 and 1030 modules, based on one or more of the received information elements including applying_module (s) information element and base_model information element.
[0133] As may be appreciated, the order of operations in procedure 1000 can vary depending on system requirements and configuration. Procedure 1000 may enable network-based distributed learning, allowing nodes to collaborate dynamically.
[0134] In the context of procedure 1000, various information elements are used to facilitate communication and execution of operations between type-1 PSFs and type-2 PSFs. In some embodiments, the target_module (s) information element informs the type-1 and type-2 PSF 904 about one or more targets, parts, or portions of a base model (e.g., a machine learning model) to which a PEFT module is to be applied. For example, target_module (s) may specify components such as fully connected (fc) layers, convolutional layers, or attention layers that are to be adjusted during fine-tuning or training operations. In some embodiments, the replacement_module (s) or applying_module (s) information element indicates to the type-1 and type 2PSFs the PEFT module to be applied, Examples of PEFT modules include LoRA, prompt tuning, prefix tuning, etc. In some embodiments, the procedure (s) information element indicates to type-1 and type-2 PSFs one or more operations and type (s) of operations that need to be performed. For instance, procedure (s) might specify tasks such as training, evaluation, or combination operations like averaging, ties merging, DARE merging, or learnable weighted averaging. These instructions ensure that the type-1 and type-2 PSFs can execute the appropriate processing operations 1008, 1010, 1014, 1028, and 1030 as outlined in procedure 1000. In some embodiments, the base_model information element allows a type-1 or type-2 PSF to specify the machine learning model as the foundation for fine-tuning in a distributed learning process. This may be useful when dealing with heterogenous models, where different nodes may use different models for different tasks. For example, a PSF might request base models like Llama2 or Llama3 to perform fine-tuning.
[0135] In some embodiments, parameter-efficient fine-tuning in a distributed learning environment may be provided, which may reduce training and communication costs. In some embodiments, distributed learning without sharing raw data may be enabled, potentially preserving client privacy. In some embodiments, the utilization of pre-trained models may be supported, which may increase the performance of models. In some embodiments, heterogeneous models may be supported, allowing nodes with different resource constraints to participate in the distributed learning process.
[0136] Some embodiments may be applied to any appropriate learning system that includes multiple computing entities, such as multi-core processors, multi-GPU servers, multi-NPU devices, multi-server cloud systems, and similar setups. Some embodiments may also be applied to other network-like systems beyond telecommunication networks, including but not limited to smart grids, IoT devices, networks of connected vehicles, networks of surveillance cameras, sensor networks, and data centers.
[0137] FIG. 11A illustrates a method for fine-tuning a machine learning model in a distributed learning environment, according to an embodiment. In some embodiments, method 1100 is performed by a network node. Method 1100 includes receiving 1101 a request including configuration instructions, and at least one module. The configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model. The method 1100 further includes sending 1102 at least one of the at least one updated module and the updated machine learning model.
[0138] In some embodiments, the sending includes sending, by the network node, at least one of the at least one updated module and the updated machine learning model to at least one other network node, wherein the request includes an indication of the at least one other network node.
[0139] In some embodiments, the configuration instructions indicate at least a portion of the machine learning model to which the at least one module is to be applied. In some embodiments, the configuration instructions are indicated in either a header or a payload of one or more packets, wherein the request includes the one or more packets.
[0140] In some embodiments, the one or more operations include one or more of training the machine learning model, evaluating the machine learning model and combining the at least one module with one or more other modules based on one or more of: a type of combination and one or more hyperparameters. The training of the machine learning model is based on one or more of: training data, a number of epochs, a learning rate, and a loss function. The evaluating of the machine learning model is based on one or more of: evaluation data and one or more metrics for evaluation.
[0141] In some embodiments, the at least one module is based on one of: Low-Rank Adaptation (LoRA) , dynamic Low-Rank Adaptation (DyLoRA) , prompt tuning, and prefix tuning.
[0142] In some embodiments, the request further indicates a data signal representing the at least one module and data related to the at least one module, wherein the one or more operations include combining, by the network node, the at least one module with a second module based on a similarity metric, wherein the similarity metric is determined using the data signal, and wherein the second module is either a local module or a module received from a second network node.
[0143] In some embodiments, the method further includes receiving 1103 from a plurality of other network nodes, a plurality of data signals, each data signal of the plurality of data signals corresponding to a network node of the plurality of other network nodes. The method further includes organizing 1104 the plurality of other network nodes into one or more groups, wherein the organizing is based on similarity metrics associated with the plurality of data signals and selecting 1105 an aggregator for each group of the one or more groups based on the similarity metrics.
[0144] In some embodiments, the method further includes sending 1106 to each network node of the plurality of other network nodes, a notification indicating a corresponding group of the one or more groups to which each network node belongs and a corresponding aggregator of the corresponding group to send at least one module of said each network node.
[0145] In some embodiments, the method further includes receiving 1107, from the plurality of other network nodes, a plurality of modules. Each module of the plurality of modules corresponding to a network node of the plurality of other network nodes, wherein each data signal of the plurality of data signals represents at least one of a corresponding module of the plurality of modules and data related to the corresponding module.
[0146] In some embodiments, the organizing includes organizing the plurality of other network nodes into one or more groups based on similarity metrics of the plurality of modules, and the selecting includes selecting the aggregator for each group based on the similarity metrics of the plurality of modules. In some embodiments, the method further includes sending 1108 to each aggregator of each group of the one or more groups, one or more modules and corresponding one or more data signals of one or more network nodes belonging to each group.
[0147] FIG. 11B illustrates a method for fine-tuning a machine learning model in a distributed learning environment, according to an embodiment. In some embodiments, method 1150 is performed by a first network node. Method 1150 includes sending 1151 to a second network node, a request including configuration instructions and at least one module, wherein the configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model. The method further includes receiving 1152 from the second network node, at least one of the at least one updated module and the updated machine learning model.
[0148] In some embodiments, the method further includes receiving 1153 from a controller, the configuration instructions. In some embodiments, the configuration instructions further indicate the machine learning model to serve as a base model for fine-tuning and instructions for sending at least one of the at least one updated module and the updated machine learning model to the second network node.
[0149] FIG. 12A illustrates an example communication system, according to an embodiment. The communication system 1200 includes a radio access network (RAN) 1220, one or more communication electronic devices (EDs) 1210a, 1210b, 1210c, 1210d, 1210e, 1210f, 1210g, 1210h, 1210i, 1210j (collectively referred to as 1210) , a core network 1230, a Public Switched Telephone Network (PSTN) 1240, the Internet 1250, and other networks 1260 . The RAN 1220 may include, but is not limited to, a future generation RAN, or a legacy RAN such as, but not limited to, 5th generation (5G) , 4th generation (4G) , 3rd generation (3G) or 2nd generation (2G) radio access network. The RAN 1220 may be, for example, an Evolved Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (E-UTRAN) , a NextGen RAN (NG RAN) , or some other type of RAN. Examples of RAN 1220 based on the evolution of telecommunications standards include, but is not limited to, GSM (Global System for Mobile Communications) and CDMA (Code Division Multiple Access) for 2G, UMTS (Universal Mobile Telecommunications System) based on WCDMA (Wideband Code Division Multiple Access) and CDMA2000 for 3G, LTE (Long-Term Evolution) and WiMAX (Worldwide Interoperability for Microwave Access) for 4G, and NR (New Radio) for 5G. In some implementations, The RAN 1220 may use any radio access technology (RAT) in the wireless interface between the one or more EDs 1210 and the RAN 1220. In some implementations, the term “radio access” may refer to the future generation air interface standards which may include both terrestrial networks (TNs) and non-terrestrial networks (NTNs) . The one or more communication EDs 1210 (also referred to as “user equipment” ) are configured to connect (e.g., communicatively couple) with each other or to one or more network nodes 1270a, 1270b (collectively referred to as 1270) in the RAN 1220. The core network (CN) 1230 is a part of the communication system 1200 and may include one or more network nodes (e.g., 1270a , 1270b) which provide support for the network features and telecommunication services. In some implementations, the CN 1230 may be dependent on the RAT used in the communication system 1200. In other implementations, the CN 1230 may be access-agnostic, i.e., the CN 1230 may be independent of the RAT used in the communication system 1200. There are different types of CN 1230, for different 3GPP system generations. For example, the CN 1230 is the Evolved Packet Core (EPC) in 4G, also known as the Evolved Packet System (EPS) . In another example, the CN 1230 is the 5G Core (5GC) which was developed as part of the 5G System (5GS) . The CN 1230 also enables integration of different 3GPP and non-3GPP access types. In some implementations, the CN 1230 also provides the interface towards external networks that may include the PSTN 1240, the Internet 1250, and other networks 1260 in the communication system 1200.
[0150] In general, the communication system 1200 facilitates interaction between multiple wireless or wired elements. The communication system 1200 may transmit different types of content, such as voice, data, video, and / or text, through different transmission methods such as, but not limited to, broadcast, multicast, groupcast, and unicast. Additionally, the communication system 1200 operates by allocating and / or sharing resources, such as carrier spectrum bandwidth, among its constituent elements.
[0151] The communication system 1200 may provide a wide range of communication services and applications including, but not limited to, Enhanced Mobile Broadband (eMBB) services, Ultra-Reliable Low-Latency Communication (URLLC) services, Massive Machine Type Communication (mMTC) services, Integrated Sensing And Communication (ISAC) , immersive communication, Ultra-massive Machine-Type Communication (uMTC) , hyper reliable and low-latency communication, ubiquitous connectivity, integrated AI and communication, and other services that can be provided by a future generation communication system. The communication system 1200 may provide other services and applications such as, but not limited to, earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility and the like.
[0152] The system 1200 may serve as the infrastructure within which various embodiments of the network-based distributed learning environment, including parameter-efficient fine-tuning (PEFT) , can be implemented. In some embodiments, the system 1200 comprises several components, including core networks, RANs, and electronic devices (EDs) , which can facilitate the exchange of PEFT modules and information elements (e.g., target_module (s) information element, applying_module (s) information element, procedure (s) information element, base_model information element, call_back information element) across the network. While the specific way these components interact to implement the embodiments may vary, system 1200 generally supports the transmission and processing of models in a decentralized learning environment. The system 1200 provides the necessary infrastructure for nodes to engage in distributed learning operations, potentially enabling efficient model fine-tuning while preserving data privacy. The overall architecture can be adapted to different deployment scenarios and network configurations without being limited to a specific implementation.
[0153] FIG. 12B illustrates another example communication system, according to an embodiment. One or embodiments described herein may apply to or be implemented by the communication system 1280. The communication system 1280 is similar to and based on the communication system 1200 and includes EDs 1210a, 1210b, 1210c, 1210d (collectively referred to as ED 1210) , RANs 1220a, 1220b, one or more CNs 1230, a PSTN 1240, the Internet 1250, and other networks 1260. Additionally, the communication system 1280 may also include a non-terrestrial network (NTN) 1220c. The RANs 1220a and1220b may include network nodes 1270a and 1270b respectively. Examples of network nodes 1270a, 1270b include base stations, which can be generally referred to as terrestrial network (TN) devices or terrestrial transmit and receive points (T-TRPs) 1270a and 1270b (collectively referred to as 1270) . TRP may refer to a base station in some embodiments. The T-TRPs 1270a, 1270b may be base stations mounted on a building or tower. In one implementation, the NTN 1220c includes a RAN node such as a base station 1272, which may be generally referred to as an NTN device, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, or a non-terrestrial transmit and receive point (NT-TRP) 1272.
[0154] A base station 1270 (which may also refer to as a TRP) is a network element within a radio access network responsible for radio transmission and reception in one or more cells to or from the ED (such as a user equipment) . In different implementations, the base station 1270 may also be known as a base transceiver station (BTS) , a radio base station, a network node, a network device, a device on the network side, a transmit / receive node, a Node B, an evolved NodeB (eNodeB or eNB) , a Home eNodeB, a next Generation NodeB (gNB) , a transmission point (TP) , a site controller, an access point (AP) , a wireless router, a relay station, a terrestrial node, a terrestrial network device, a terrestrial base station, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, and a positioning node, among other possibilities. The base station 1270 may be a macro base station (BS) , a pico BS, a relay node, a donor node, or combinations thereof. When the base station 1270 performs (or is configured to perform) a method described herein, it may be interpreted as the base station itself, one or more modules (or units) in the base station, a circuit or chip, or a combination thereof, performing the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, system in package (SIP) ) , and the like, and may be responsible for one or more communication functions within the base station.
[0155] The EDs 1210a-1210d and TRPs 1270a-1270b, 1272 are examples of communication equipment configured to implement some or all of the operations and / or implementations described herein. The T-TRP 1270a forms part of the RAN 1220a, which may include other TRPs, and / or other devices. Also, the TRP 1270b forms part of the RAN 1220b, which may include other TRPs, and / or devices. Each TRP 1270a, 1270b may transmit and / or receive wireless signals within a particular geographic region or area, sometimes referred to as a “cell” or a “coverage area” . The TRPs 1270a-1270b may be responsible for allocating and / or configuring resources and transmission and / or reception in a set of cell (s) . A cell is a radio network object that can be uniquely identified by a cell identification that is broadcasted over a geographical region or area from base stations associated with the cell. The number of RANs 1220a-1220b shown is merely an example. Any number of RANs may be contemplated when designing the communication system 1280.
[0156] A base station may be a single element, as shown in the figures, or multiple elements distributed throughout the corresponding RAN, or otherwise configured. In some implementations, a plurality of RAN nodes coordinate to assist the ED 1210 in implementing radio access, and different RAN nodes separately implement and handle different functions of the base station. For example, the RAN node may be a central unit (CU) , a distributed unit (DU) , a CU-control plane (CP) , a CU-user plane (UP) , or a radio unit (RU) etc. The CU and the DU may be separately deployed, or included within the same element (i.e., a baseband unit (BBU) ) .
[0157] Furthermore, communication between different devices / apparatuses in various implementations of this disclosure may refer to direct communication (that is, without the need of forwarding by another device / apparatus) , or may refer to communication (s) between different devices / apparatuses via another device / apparatus (that is, requiring forwarding by another device / apparatus) . Alternatively, such communication (s) may involve one functional unit inside a device / apparatus using another functional unit within the device / apparatus to communicate with another device / apparatus. In other words, phrases such as "sending (or transmitting) information to... (an ED or a base station) " in this disclosure may be understood as a destination endpoint of the information being an ED or a base station, including, sending / transmitting information directly or indirectly to an ED or a base station. Similarly, phrases like "receiving information from... (an ED or a base station) " may be understood as a source endpoint of the information being an ED or a base station, including directly or indirectly receiving information from an ED or a base station. Between the source endpoint that sends the information and the destination endpoint, necessary processing such as, but not limited to, format conversion, digital-to-analog conversion, amplification, and filtering may be performed on the information. However, the destination endpoint may understand valid information from the source endpoint. A similar understanding applies to other descriptions in this disclosure without reiterating details already described. In the present disclosure, the terms "send" and "transmit" may be used interchangeably in different implementations of this disclosure.
[0158] The ED 1210 is used to connect people, objects, machines, and other entities. The ED 1210 may be widely used in various scenarios including, but not limited to, cellular communications, device-to-device (D2D) , vehicle to everything (V2X) , peer-to-peer (P2P) , machine-to-machine (M2M) , MTC, internet of things (IoT) , virtual reality (VR) , augmented reality (AR) , mixed reality (MR) , metaverse, digital twin, industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, and autonomous delivery and mobility.
[0159] Each ED 1210 represents any suitable end user device for wireless operation and may include such devices (or may be referred to as, but not limited to) a user equipment (UE) or a user device or a terminal device, a wireless transmit / receive unit (WTRU) , a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA) , an MTC device, a personal digital assistant (PDA) , a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, wearable devices (such as a watch, a pair of glasses, head mounted equipment, etc. ) , an industrial device, or an apparatus (such as a module, modem, or chip) in the forgoing devices, among other possibilities. Future generation EDs 1210 may be referred to by other terms. When an ED 1210 performs (or is configured to perform) a method described herein, it may be interpreted as the ED itself, one or more modules (or units) in the ED, a circuit or chip, or a combination thereof, performing the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, or system in package (SIP) ) , and the like, and may be responsible for one or more communication functions in the ED.
[0160] Each ED 1210 connected to TRPs 1270a-1270b, and / or TRPs 1272 can be dynamically or semi-statically turned-on (i.e., established, activated, or enabled) , turned-off (i.e., released, deactivated, or disabled) and / or configured in response to one of more of:connection availability and connection necessity.
[0161] Any ED 1210 may be alternatively or additionally configured to interface, access, or communicate with any of the TRPs 1270a, 1270b and 1272, the Internet 1250, the CN 1230, the PSTN 1240, the other networks 1260, or any combination thereof. In some examples, the ED 1210a may communicate an uplink (UL) and / or downlink (DL) transmission over a terrestrial air interface 1290a with station-TRP 1270a. In some examples, the EDs 1210a, 1210b, 1210c, and 1210d may also communicate directly with one another via one or more sidelink (SL) air interfaces 1290b. In some examples, the EDs 1210a, 1210d may communicate using an UL and / or DL transmission over a non-terrestrial air interface 1290c with NT-TRP 1272.
[0162] An air interface (such as, for example, 1290a, 1290b, 1290c) generally includes a number of components and associated parameters that collectively specify how a transmission is to be sent and / or received over a wireless communications link between two or more communicating devices such as EDs and base station (s) . For example, an air interface may include one or more components defining the waveform (s) , frame structure (s) , multiple access scheme (s) , protocol (s) , coding scheme (s) and / or modulation scheme (s) for conveying information (such as, data) over a wireless communications link. The air interfaces 1290a and 1290b may use similar communication technology, that may include any suitable radio access technology.
[0163] The non-terrestrial air interface 1290c can enable communication between the EDs 1210a, 1210d and one or more NT-TRPs 1272 via a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs 1210 and one or more NT-TRPs 1272 for multicast transmission.
[0164] The TRPs 1270a-1270b, 1272 may communicate with one another over one or more air interfaces 1290e, 1290f using wireless communication links (such as radio frequency (RF) , microwave, infrared (IR) , etc. ) or wired communication links. The air interfaces 1290e, 1290f may utilize any suitable radio access technology, and may be substantially similar to the air interfaces 1290a, 1290c over which the EDs 1210a-1210d communicate with one or more of the TRP 1270a-1270b, 1272 or they may be substantially different. For example, the communication system 1200 or 1280 may implement one or more channel access methods, such as Time Division Multiple Access (TDMA) , Frequency Division Multiple Access (FDMA) , Code Division Multiple Access (CDMA) , Single Carrier Frequency Division Multiple Access (SC-FDMA) , Low Density Signature Multicarrier Code Division Multiple Access (LDS-MC-CDMA) , Non-Orthogonal Multiple Access (NOMA) , Pattern Division Multiple Access (PDMA) , Lattice Partition Multiple Access (LPMA) , Resource Spread Multiple Access (RSMA) , and Sparse Code Multiple Access (SCMA) .
[0165] The RANs 1220a and 1220b are in communication with the CN 1230 to provide the EDs 1210a 1210b, and 1210c with various services such as voice, data, multimedia, and other services. The RANs 1220a and 1220b and / or the CN 1230 may be in direct or indirect communication with one or more other RANs (not shown) , which may or may not be directly served by the CN 1230, and may employ different radio access technologies from RAN 1220a and / or RAN 1220b. The CN 1230 may also serve as a gateway access between (i) the RANs 1220a and 1220b and / or the EDs 1210a 1210b, and 1210c, and (ii) other networks (such as the PSTN 1240, the Internet 1250, and the other networks 1260) . In addition, some or all of the EDs 1210a 1210b, and 1210c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. For example, the EDs 1210a 1210b, and 1210c communicate using different cellular communications protocols, such as, but not limited to, a Global System for Mobile Communications (GSM) protocol, a code-division multiple access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a Universal Mobile Telecommunications System (UMTS) protocol, a 3GPP Long Term Evolution (LTE) protocol, a fifth generation (5G) protocol, a New Radio (NR) protocol, and the like. Instead of wireless communication (or in addition thereto) , the EDs 1210a 1210b, and 1210c may communicate using wired communication channels to a service provider or switch (not shown) , and / or to the Internet 1250. The PSTN 1240 may include circuit switched telephone networks for providing plain old telephone service (POTS) . The Internet 1250 may include a network of computers and subnets (intranets) or both, and incorporate protocols, such as internet protocol (IP) , transmission control protocol (TCP) , user datagram protocol (UDP) . EDs 1210a 1210b, and 1210c may be multimode devices capable of operation according to multiple radio access technologies, and may incorporate one or multiple transceivers necessary to support such.
[0166] The system 1280 builds upon the architecture of system 1200 (FIG. 12A) and includes both terrestrial and non-terrestrial elements to further enhance the flexibility and reach of the distributed learning system. Similar to system 1200, system 1280 can support the implementation of PEFT and other embodiments by enabling communication between network nodes and devices through terrestrial (T-TRP) and non-terrestrial (NT-TRP) networks. This dual-network setup allows devices in remote or hard-to- reach areas to participate in distributed learning, ensuring that the learning environment is not limited by geographic or infrastructural constraints. The non-terrestrial network components provide additional connectivity options for devices that may not have access to terrestrial networks, facilitating the exchange of PEFT modules and supporting heterogeneous models in different network environments. Thus, just as system 1200 can be adapted to support the embodiments, system 1280 offers a flexible architecture for extending those capabilities to more diverse communication scenarios.
[0167] FIG. 13 illustrates an example of an apparatus 1310 wirelessly communicating with apparatus 1320 in a communication system (e.g., the communication system 1200 or 1280) . The apparatus 1310 may be an electronic device (e.g. ED 1210) . The apparatus 1320 may be a network node (e.g. network node 1270) such as T-TRP 1270 or a NT-TRP 1272. Although there is only one apparatus 1310, and one apparatus 1320 shown in the figure, the number of apparatus 1310 and / or 1320 could be one or more. For example, one ED 1210 may be served by only one T-TRP 1270 (or one NT-TRP 1272) , by more than one T-TRP 1270 (or more than one NT-TRP 1272) . One ED 1210 may be served by one or more T-TRP 1270 and one or more NT-TRP1272. Similarly, one T-TRP 1270 (or one NT-TRP1272) may serve one or more ED 1210.
[0168] Apparatus 1310 includes at least one processor 1312. Only one processor 1312 is illustrated to avoid congestion in the drawing. The apparatus 1310 may further include a transmitter 1301 and a receiver 1303 coupled to one or more antennas 1304. Only one antenna 1304 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 1304 may alternatively be panels. The transmitter 1301 and the receiver 1303 may be integrated, e.g. as a transceiver. The transceiver is configured to modulate data or other content for transmission by at least one antenna 1304 or network interface controller (NIC) . The transceiver is also configured to demodulate data or other content received by the at least one antenna 1304. Each transceiver includes any suitable structure for generating signals for wireless or wired transmission and / or processing signals received wirelessly or by wire. Each antenna 1304 includes any suitable structure for transmitting and / or receiving wireless or wired signals. The apparatus 1310 may include at least one memory 1308. Only the transmitter 1301, receiver 1303, processor 1312, memory 1308, and antenna 1304 is illustrated for simplicity, but the apparatus 1310 may include one or more other components. In present disclosure, the transceiver (or transmitter 1301 and / or receiver1303) may be viewed as an interface circuit.
[0169] The memory 1308 stores instructions used to perform operations described herein. The memory 1308 may also store data used, generated, or collected by the apparatus 1310. For example, the memory 1308 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by one or more processor 1312.
[0170] The apparatus 1310 may further include one or more input / output devices (not shown) or interfaces. The input / output devices or interfaces permit interaction with a user or other devices in the network. Each input / output device or interface includes any suitable structure for providing information to or receiving information from a user, and / or for network interface communications. Suitable structures include, for example, a speaker, microphone, keypad, keyboard, display, touch screen, etc.
[0171] The processor 1312 may perform (or control the apparatus 1310 to perform) operations (or methods) described herein as being performed by the apparatus 1310. For example, the processor 1312 performs or controls the apparatus 1310 to perform receiving transport blocks (TBs) , using a resource for decoding of one of the received TBs, releasing the resource for decoding of another of the received TBs, and / or receiving configuration information configuring a resource. In detail, the operation may include those operations related to preparing a transmission for uplink (UL) transmission to the apparatus 1320; those operations related to processing DL transmissions received from the apparatus 1320; and those operations related to processing SL transmission to and from another apparatus 1310. Processing operations related to preparing a transmission for UL transmission may include operations such as encoding, modulating, transmit beamforming, and generating symbols for transmission. Processing operations related to processing DL transmissions may include operations such as receive beamforming, demodulating and decoding received symbols. Processing operations related to processing SL transmissions may include operations such as transmit / receive beamforming, modulating / demodulating and encoding / decoding symbols. Depending upon the embodiment, a DL transmission may be received by the receiver 1303, possibly using receive beamforming, and the processor 1312 may extract signaling from the DL transmission (e.g. by detecting and / or decoding the signaling) . An example of signaling may be a reference signal transmitted by the apparatus 1320. In some implementations, the processor 1312 implements the transmit beamforming and / or the receive beamforming based on the indication of beam direction, e.g. beam angle information (BAI) , received from the apparatus 1320. In some implementations, the processor 1312 may perform operations relating to network access (e.g. initial access) and / or downlink synchronization, such as operations relating to detecting a synchronization sequence, decoding and obtaining the system information, etc. In some implementations, the processor 1312 may perform channel estimation, e.g. using a reference signal received from the apparatus 1320.
[0172] Although not illustrated, the processor 1312 may form part of the transmitter 1301 and / or part of the receiver 1303. Although not illustrated, the memory 1308 may form part of the processor 1312.
[0173] The processor 1312, the processing components of the transmitter 1301, and the processing components of the receiver 1303 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory (e.g. in the memory 1308) .
[0174] The apparatus 1320 includes one or more processors 1360 (only one processor 1360 is illustrated to in the figure) . The apparatus 1320 may further include at least one transmitter 1352 and at least one receiver 1354 coupled to one or more antennas 1356. Only one antenna 1356 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 1356 may alternatively be panels. The transmitter 1352 and the receiver 1354 may be integrated as a transceiver. The apparatus 1320 may further include at least one memory 1358. The apparatus 1320 may further include scheduler 1353. Only the transmitter 1352, receiver 1354, processor 1360, memory 1358, antenna 1356 and scheduler 1353 are illustrated for simplicity, but the apparatus 1320 may include one or more other components. In present disclosure, the transceiver (or transmitter 1352 and / or receiver1354) may be viewed as an interface circuit.
[0175] In some implementations, the parts of the apparatus 1320 may be distributed. For example, some of the modules of the apparatus 1320 may be located remote from the equipment that houses the antennas 1356 for the apparatus 1320 (thereby also can be viewed as one of more nodes) , and may be coupled to the equipment that houses the antennas 1356 over a communication link (not shown) sometimes known as front haul, such as common public radio interface (CPRI) . Therefore, in some implementations, the term apparatus 1320 may also refer to nodes on the network side that perform processing operations, such as determining the location of the apparatus 1310, resource allocation (scheduling) , message generation, and encoding / decoding, and that are not necessarily part of the equipment that houses the antennas 1356 of the apparatus 1320. The nodes may also be coupled to other apparatus 1320s. In some implementations, the apparatus 1320 may actually be a plurality of nodes that are operating together to serve the apparatus 1310, e.g. through the use of coordinated multipoint transmissions.
[0176] The processor 1360 performs operations including those related to: preparing a transmission for DL transmission to the apparatus 1310, processing an UL transmission received from the apparatus 1310, preparing a transmission for backhaul transmission to another apparatus 1320, and processing a transmission received over backhaul from another apparatus 1320. Processing operations related to preparing a transmission for DL or backhaul transmission may include operations such as encoding, modulating, precoding (e.g. multiple input multiple output (MIMO) precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the UL or over backhaul may include operations such as receive beamforming, demodulating received symbols, and decoding received symbols. The processor 1360 may also perform operations relating to network access (e.g. initial access) and / or DL synchronization, such as generating the content of synchronization signal blocks (SSBs) , generating the system information, etc. In some implementations, the processor 1360 also generates an indication of beam direction, e.g. BAI, which may be scheduled for transmission by a scheduler 1353 which will be described below. In some implementations, the processor 1360 implements the transmit beamforming and / or receive beamforming based on beam direction information (e.g. BAI) received from another apparatus 1320. The processor 1360 performs other network side processing operations described herein, such as determining the location of the apparatus 1310, determining where to deploy another apparatus 1320, etc. In some implementations, the processor 1360 may generate signaling, e.g. to configure one or more parameters of the apparatus 1310 and / or one or more parameters of another apparatus 1320. Any signaling generated by the processor 1360 is sent by the transmitter 1352. In some implementations, the apparatus 1320 implements physical layer processing. In some implementations, the apparatus 1320 may implement higher layer functions such as functions at the medium access control (MAC) or radio link control (RLC) layer in addition to physical layer processing. The apparatus 1320 may further comprise scheduler 1353 coupled to the processor 1360 or integrated in the processor 1360. The scheduler 1353 may be included within or operated separately from the apparatus 1320a. The scheduler 1353 may schedule UL, DL, SL, and / or backhaul transmissions, including issuing scheduling grants and / or configuring scheduling-free (e.g., “configured grant” ) resources.
[0177] The apparatus 1320a may further include a memory 1358 storing instructions used to perform operations described herein. The memory 1358 may also store data used, generated, or collected by the apparatus 1320. For example, the memory 1358 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by the processor 1360.
[0178] Although not illustrated, the processor 1360 may form part of the transmitter 1352 and / or part of the receiver 1354. Also, although not illustrated, the processor 1360 may implement the scheduler 1353. Although not illustrated, the memory 1358 may form part of the processor 1360.
[0179] The processor 1360, the scheduler 1353, the processing components of the transmitter 1352, and the processing components of the receiver 1354 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory, e.g., in the memory 1358.
[0180] The apparatus 1320 and / or the apparatus 1310 may include other components, but these have been omitted for the sake of clarity.
[0181] One or more embodiments may be implemented by apparatus 1310 or 1320. Either of the apparatus 1310 or 1320 may perform one or more operations described herein and can be configured in various capacities within the distributed learning system. Each apparatus may function as a node in a distributed network, which could include roles such as a server, base station, sender node, receiver node, aggregator, or centralized controller. Additionally, each apparatus may serve as a PSC, a type-1 or type-2 PSF, or any other module or component integral to these roles. Each apparatus is capable of supporting PEFT operations and managing information elements such as target_module (s) , applying_module (s) , procedures, call_back, and base_model within a network-based distributed learning environment. The apparatuses may also act as network elements involved in specific functions such as a module within a base station or server, a network interface device, or a processing unit within a larger system architecture. Furthermore, each apparatus can be implemented as an edge device, access point, relay, or any other entity responsible for various communication tasks, processing operations, and network management functions.
[0182] Note that “signaling” , as used herein, may alternatively be called control signaling, control message, control information, or message for simplicity. Signaling between a base station (e.g., the TRP 1270a-b, 1272) and a UE or sensing device (e.g., ED 1210) , or signaling between a different UE or sensing device (e.g., between ED 1210a and ED1210b) may be carried in physical layer signaling (also called as dynamic signaling) , which is transmitted in a physical layer control channel.
[0183] FIG. 14A illustrates an example apparatus 1410, according to an embodiment. The apparatus 1410 may be a communication device or an apparatus implemented in a communication device such as the ED 1210 or the TRPs 1270a, 1270b, 1272. For example, the apparatus 1410 implemented in an ED may be an integrated circuit, which in some instances may be referred to as a chip, a modem, a modem chip, a baseband chip, or a baseband processor. In some implementations, one or more integrated circuits can be packaged into a system-on-chip, a system-in-package, or a multi-chip module. The apparatus 1410 can include one or more integrated circuits and other discrete components. In some implementations, the apparatus 1410 may be a module within one of the TRPs 1270a, 1270b, 1272, or the apparatus 620.
[0184] In an example, the apparatus 1410 may include one or more processors 1411, and an interface circuit 1412. The apparatus 1410 may further include a memory 1413. The one or more processors 1411 are configured to process signals and execute one or more communication protocols. The memory 1413 is configured to store at least a part of corresponding computer program instructions and / or data. In an example, the one or more processors 1411 execute the computer program instructions stored in the memory 1413 to implement related operations (for example, inputting, outputting, receiving, and transmitting) in the method embodiments disclosed herein. In some implementations, the memory 1413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 1413 is configured to store all of the corresponding computer program instructions and / or data for execution by the one or more processors 1411. In some implementations, the memory 1413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 1413 is configured to store a part of the corresponding computer program instructions and / or data. For example, the part of the corresponding computer program instructions and / or data may include computer program instructions and / or data that need to be currently executed by the one or more processors 1411. Thus, the memory 1413 may store different parts of computer program instructions and / or data for a plurality times for the one or more processors 1411 to perform related operations in the method embodiments disclosed herein. As a communication interface, the interface circuit 1412 is configured to implement communication with another component. For example, the interface circuit 1412 may communicate a signal with another apparatus or system, such as a radio frequency processing apparatus or another processor. The signal may include or carry information intended as a payload, such as user data, control information, etc. The signal may also include or carry information useful to a receiver, but not necessarily as a payload, such as a pilot signal or reference signal. Communicating the signal may include transmitting the signal to another component or device. Communicating the signal may additionally or alternatively include receiving the signal from another component or device. Transmitting the signal may include outputting the signal to a component or device that is directly or indirectly coupled to the interface circuit 1412. Receiving the signal may include inputting or obtaining the signal from a component or device that is directly or indirectly coupled to the interface circuit 1412. Optionally, to reduce a load of the one or more processors, a baseband signal processing circuit 1414 may be also disposed to implement processing of at least a part of baseband signals, including signal demodulation, modulation, encoding, decoding, or the like.
[0185] Apparatus 1410 may be processor 1312 (or 1360) in apparatus 1310 (or 1320) , in some scenario, or included in processor 210 1312 (or 1360) in apparatus 1310 (or 1320) in some scenario. Apparatus 1410 may be or include a baseband chip. In some implementations, the apparatus 1410 may be independently packaged into a chip. In some implementations, the apparatus 1310 (or 1320) includes different types of chips. The apparatus 1410 may be packaged into a processor chip (for example, a SoC chip or an SIP chip) with the different types of chips. In some implementations, the apparatus 1410 may be packaged into a chip with some or all of circuits of a radio frequency processing system that may further included in the apparatus 1310 (or 1320) .
[0186] FIG. 14B illustrates another example apparatus 1430 according to an embodiment. The apparatus 1430 may include corresponding modules or units configured to implement methods and / or implementations described herein. In some implementations, the apparatus 1430 includes a processing unit 1432 and a communication unit 1433. Optionally, the apparatus 1430 may further include a storage unit 1431 configured to store apparatus program code (or instructions) and / or data.
[0187] The apparatus 1430 may be an ED side apparatus, for example, an ED or a module in an ED, or a circuit or a chip responsible for a communication function in an ED. In some implementations, apparatus 1430 may be the apparatus 1310. The processing unit 1432 may be the processor 660. The communication unit 1433 may comprise a receiving unit and / or a transmitting unit. The receiving unit and / or the transmitting unit may be the transmitter 652 and / or the receiver 654 respectively. The storage unit 1431 may be the memory 658.
[0188] The apparatus 1430 may be an ED side apparatus, for example, an ED or a module in an ED, or a circuit or a chip responsible for a communication function in an ED. In some implementations, apparatus 1430 may be implemented as apparatus 1310, accordingly, the processing unit 1432 is implemented as processor 1312, the communication unit 1433 is implemented as transmitter 1301 and / or receiver 1303, and the storage unit 1431 is implemented as memory 1308.
[0189] The apparatus 1430 may be a base station side apparatus, for example, a base station or a module in a base station, or a circuit or a chip responsible for a communication function in a base station. In some implementations, apparatus 1430 may be implemented as apparatus 1320, accordingly, the processing unit 1432 is implemented as processor 1360 (the scheduler 1353 may also be included) , the communication unit 1433 is implemented as transmitter 1352 and / or receiver 1354, and the storage unit 1431 is implemented as memory 1358.
[0190] In some implementations, when the apparatus 1430 is an ED 1210 or a module in an ED 1210, a function of the apparatus 1430 may be implemented by one or more processors. The processor may include a modem chip, or a system on chip (SoC) chip or an SIP chip that includes a modem core. A function of the communication unit 1433 may be implemented by a transceiver circuit.
[0191] In some implementations, when the apparatus 1430 is a circuit or a chip that is responsible for a communication function in an ED 1210 –such as a modem chip, a system on chip (SoC) chip or an SIP chip that includes a modem core –a function of the processing unit 1432 may be implemented by a circuit system within the chip which includes one or more processors. A function of the communication unit 1433 may be implemented by an interface circuit or a data transceiver circuit on the chip.
[0192] It may be understood that the units in the apparatus 1430 may be logical or functional. Each function may correspond to one functional unit, or two or more functions may be integrated into a single functional unit. In actual implementation, all or some of the units may be integrated into a single physical entity, or may be distributed across different physical entities. In addition, the functional units may be implemented in the form of hardware, software, or a combination of hardware and software. Whether a function is implemented in the form of hardware or software depends on particular applications and design constraint conditions of the technical solutions. A person skilled in the art may use different methods to implement the described functions for specific applications, but it should not be considered that the implementation goes beyond the scope of this disclosure.
[0193] In an example, a functional unit in any one of the apparatuses may be configured as one or more integrated circuits for implementing the methods disclosed herein, for example, as one or more application-specific integrated circuits (application-specific integrated circuits, ASICs) , one or more central processing units (CPUs) , one or more microprocessors or microprocessor units (MPUs) , one or more microcontrollers or microcontroller units (MCUs) , one or more digital signal processors (DSPs) , one or more field programmable gate arrays (FPGAs) , or a combination of these.
[0194] In an example, the storage unit 1431 may include a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, and / or a register.
[0195] A processor may be referred to as a processor system, an application processor, a baseband processor, a processor circuit, or a processor core. The processor may include one or a combination of one or more central processing units (CPUs) , one or more digital signal processors (DSPs) , one or more microprocessors (microprocessor units, MPUs) , one or more microcontrollers (microcontroller units, MCUs) , one or more graphics processing units (GPUs) , one or more field programmable gate arrays (FPGAs) , one or more artificial intelligence processors (AI processors) , or one or more neural network processing units (NPUs) .
[0196] Memory or a storage unit may include one or more of the following storage media: a random access memory (RAM) , a static random access memory (static RAM, SRAM) , a dynamic random access memory (dynamic RAM, DRAM) , a phase-change memory (PCM) , a resistive random access memory (resistive RAM, ReRAM) , a magnetoresistive random access memory (magnetoresistive RAM, MRAM) , a ferroelectric random access memory (ferroelectric RAM, FRAM) , a cache, a register, a read-only memory (ROM) , a flash memory (flash memory) , an erasable programmable read-only memory (erasable programmable ROM, EPROM) , a hard disk, and the like. In an example, computer program instructions used to execute embodiments may be stored in a non-volatile memory, for example, at least a part of a memory or storage unit (for example, one or more of a ROM, a flash memory, an EPROM, or a hard disk) . When a terminal runs, a part or all of corresponding computer program instructions may be loaded to a memory that has a higher transmission speed with the processor, for example, at least a part of a memory or a storage unit (for example, one or more of a RAM, an SRAM, a DRAM, a PCM, a RERAM, an MRAM, a FRAM, a cache, or a register) , so that the processor executes the computer program instructions to perform the steps in the method embodiments disclosed herein.
[0197] A person skilled in the art should understand that embodiments of this application may be provided as a method, an apparatus (or system) , computer-readable storage medium, or a computer program product. Therefore, this application may use a form of a hardware-only embodiment, a software-only embodiment, or an embodiment with a combination of software and hardware. Moreover, this application may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.
[0198] Embodiments of the present application can be implemented using electronics hardware, software, or a combination thereof. In some embodiments, the application is implemented by one or multiple computer processors executing program instructions stored in memory. In some embodiments, the application is implemented partially or fully in hardware, for example using one or more field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs) to rapidly perform processing operations.
[0199] In the present disclosure, the terms “a” or “an” are defined to mean “at least one” , that is, these terms do not exclude a plural number of items, unless stated otherwise.
[0200] In the present disclosure, terms such as “substantially” , “generally” and “about” , which modify a value, condition or characteristic of a feature of an example embodiment, should be understood to mean that the value, condition or characteristic is defined within tolerances that are acceptable for the proper operation of the example embodiment for its intended application.
[0201] In the present disclosure, unless stated otherwise, the terms “connected” and “coupled” , and derivatives and variants thereof, refer herein to any structural or functional connection or coupling, either direct or indirect, between two or more elements. For example, the connection or coupling between the elements can be acoustical, mechanical, optical, electrical, thermal, logical, or any combinations thereof.
[0202] In the present disclosure, the expression “based on” is intended to mean “based at least partly on” , that is, this expression can mean “based solely on” or “based partially on” , and so should not be interpreted in a limited manner. More particularly, the expression “based on” could also be understood as meaning “depending on” , “representative of” , “indicative of” , “associated with” or similar expressions.
[0203] In the present disclosure, the terms "system" and "network" may be used interchangeably in different embodiments of this application. "At least one" means one or more, and "a plurality of" means two or more. The term "and / or" describes an association relationship of associated objects, and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: Only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character " / " indicates an "or" relationship between associated objects. "At least one of the following items (pieces) " or a similar expression thereof indicates any combination of these items, including a single item (piece) or any combination of a plurality of items (pieces) .
[0204] In the present disclosure, "at least one" means one or more, and "a plurality of" means two or more. "and / or" describes an association relationship of associated objects, and indicates that there may be three relationships. For example, A and / or B may indicate cases includes “only A” , “both A and B” , and “only B” , where A and B may be singular or plural. The character " / " generally indicates that the associated objects are in an OR relationship. "At least one of the following items" or a similar expression thereof refers to any combination of these items, including any combination of a single item or a plurality of items. For example, “at least one of a, b, or c” may represent a, b, c, “a and b” , “a and c” , “b and c” , or “a, b and c” , where a, b, and c may be a single or multiple form.
[0205] It will be appreciated that, although specific embodiments of the technology have been described herein for purposes of illustration, various modifications may be made without departing from the scope of the technology. The specification and drawings are, accordingly, to be regarded simply as an illustration of the application as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present application. In particular, it is within the scope of the technology to provide a computer program product or program element, or a program storage or memory device such as a magnetic or optical wire, tape or disc, or the like, for storing signals readable by a machine, for controlling the operation of a computer according to the method of the technology and / or to structure some or all of its components in accordance with the system of the technology.
[0206] Acts associated with the method described herein can be implemented as coded instructions in a computer program product. In other words, the computer program product is a computer-readable medium upon which software code is recorded to execute the method when the computer program product is loaded into memory and executed on the microprocessor of the wireless communication device.
[0207] Further, each operation of the method may be executed on any computing device, such as a personal computer, server, PDA, or the like and pursuant to one or more, or a part of one or more, program elements, modules or objects generated from any programming language, such as C++, Java, or the like. In addition, each operation, or a file or object or the like implementing each said operation, may be executed by special purpose hardware or a circuit module designed for that purpose.
[0208] Through the descriptions of the preceding embodiments, the present application may be implemented by using hardware only or by using software and a necessary universal hardware platform. Based on such understandings, the technical solution of the present application may be embodied in the form of a software product. The software product may be stored in a non-volatile or non-transitory storage medium, which can be a compact disc read-only memory (CD-ROM) , USB flash disk, or a removable hard disk. The software product includes a number of instructions that enable a computer device (personal computer, server, or network device) to execute the methods provided in the embodiments of the present application. For example, such an execution may correspond to a simulation of the logical operations as described herein. The software product may additionally or alternatively include a number of instructions that enable a computer device to execute operations for configuring or programming a digital logic apparatus in accordance with embodiments of the present application.
[0209] Although the present application has been described with reference to specific features and embodiments thereof, it is evident that various modifications and combinations can be made thereto without departing from the application. The specification and drawings are, accordingly, to be regarded simply as an illustration of the application as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present application.
Claims
1.A method comprising:receiving, at a network node, a request including configuration instructions, and at least one module;wherein the configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model; andsending, by the network node, at least one of the at least one updated module and the updated machine learning model.2.The method of claim 1, wherein the sending comprises:sending, by the network node, at least one of the at least one updated module and the updated machine learning model to at least one other network node, wherein the request comprises an indication of the at least one other network node.3.The method of claim 1, wherein the configuration instructions indicate at least a portion of the machine learning model to which the at least one module is to be applied.4.The method of any one of claims 1 to 3, wherein the configuration instructions are indicated in either a header or a payload of one or more packets, wherein the request comprises the one or more packets.5.The method of claim 1, wherein the one or more operations include one or more of:training the machine learning model based on one or more of: training data, a number of epochs, a learning rate, and a loss function;evaluating the machine learning model, the evaluation based on one or more of: evaluation data and one or more metrics for evaluation; andcombining the at least one module with one or more other modules based on one or more of: a type of combination and one or more hyperparameters.6.The method of claim 1, wherein the at least one module is based on one of: Low-Rank Adaptation (LoRA) , dynamic Low-Rank Adaptation (DyLoRA) , prompt tuning, and prefix tuning.7.The method of claim 1, wherein the request further indicates a data signal representing the at least one module and data related to the at least one module, whereinthe one or more operations comprises combining, by the network node, the at least one module with a second module based on a similarity metric, wherein the similarity metric is determined using the data signal, and wherein the second module is either a local module or a module received from a second network node.8.The method of claim 7, wherein the similarity metric satisfies a similarity threshold.9.The method of claim 1 further comprising:receiving, by the network node from a plurality of other network nodes, a plurality of data signals, each data signal of the plurality of data signals corresponding to a network node of the plurality of other network nodes;organizing, by the network node, the plurality of other network nodes into one or more groups, wherein the organizing is based on similarity metrics associated with the plurality of data signals; andselecting, by the network node, an aggregator for each group of the one or more groups based on the similarity metrics.10.The method of claim 9 further comprising:sending, by the network node to each network node of the plurality of other network nodes, a notification indicating:a corresponding group of the one or more groups to which each network node belongs; anda corresponding aggregator of the corresponding group to send at least one module of said each network node.11.The method of claim 9 further comprising:receiving, by the network node from the plurality of other network nodes, a plurality of modules, each module of the plurality of modules corresponding to a network node of the plurality of other network nodes, wherein each data signal of the plurality of data signals represents at least one of a corresponding module of the plurality of modules and data related to the corresponding module.12.The method of any one of claims 9 to 11, wherein:the organizing comprises organizing, by the network node, the plurality of other network nodes into one or more groups based on similarity metrics of the plurality of modules; andwherein the selecting comprises selecting, by the network node, the aggregator for each group based on the similarity metrics of the plurality of modules.13.The method of claim 12 further comprising:sending, by the network node to each aggregator of each group of the one or more groups, one or more modules and corresponding one or more data signals of one or more network nodes belonging to each group.14.A method comprising:sending, by a first network node to a second network node, a request including configuration instructions and at least one module, wherein the configuration instructions include one or more operations to be performed on the at least one module to obtain at least one updated module, wherein the at least one updated module is used to update a machine learning model; andreceiving, by the first network node from the second network node, at least one of the at least one updated module and the updated machine learning model.15.The method of claim 14 further comprising:receiving, by the first network node from a controller, the configuration instructions.16.The method of claim 14, wherein:the configuration instructions further indicate:the machine learning model to serve as a base model for fine-tuning; andinstructions for sending at least one of the at least one updated module and the updated machine learning model to the second network node.17.An apparatus comprising:one or more processors; andone or more memories storing instructions which, when executed by the one or more processors, cause the apparatus to perform the method of any one of claims 1 to 16.18.A system, wherein the system comprises a first apparatus configured to perform the method of any one of claims 1 to 13 and a second apparatus configured to perform the method of any one of claims 14 to 16.19.A non-transitory computer-readable storage medium having instructions stored thereon which, when executed by an apparatus, cause the apparatus to perform the method of any one of claims 1 to 13 or any one of claims 14 to 16.20.A computer program product storing instructions which, when executed, cause an apparatus to perform the method of any one of claims 1 to 13 or any one of claims 14 to 16.
Citation Information
Patent Citations
Edge deployment of cloud-originated machine learning and artificial intelligence workloads
US12033006B1
Federated learning optimizations
US20230177349A1
Uses of coded data at multi-access edge computing server
US20240155025A1
Federated learning by discovering clients in a visited wireless communication network
WO2024088590A1