Communication method, terminal device, and network device
By receiving global model parameters and selectively transmitting local model aggregate parameters through terminal devices, and using a low-rank adapter for large model fine-tuning, the problem of excessive transmission burden in large model fine-tuning systems is solved, improving communication efficiency and privacy security.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- QUECTEL WIRELESS SOLUTIONS CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-28
AI Technical Summary
In large model fine-tuning systems, the large number of model parameters and frequent interactions can overburden wireless network transmission, affecting communication efficiency and privacy security.
By receiving global model parameters through terminal devices, determining the first parameter, and selectively transmitting local model aggregation parameters, the amount of data transmitted and aggregated is reduced. A low-rank adapter is used for model fine-tuning, reducing communication overhead and resource consumption.
It effectively reduces the amount of data transmitted and aggregated during model training, lowers communication overhead and resource consumption, and improves communication efficiency and privacy security.
Smart Images

Figure CN2024134259_28052026_PF_FP_ABST
Abstract
Description
Communication methods, terminal equipment and network equipment Technical Field
[0001] This application relates to the field of communication technology, and more specifically, to a communication method, terminal equipment, and network equipment. Background Technology
[0002] To meet diverse user needs, artificial intelligence services can be provided based on large models in wireless networks. Large models typically have massive amounts of training data, requiring the introduction of mechanisms for fine-tuning large models to improve the model update rate.
[0003] However, due to the large number of parameters in large models, each fine-tuning round of the related large model fine-tuning system still requires the transmission of a large number of model parameters and frequent parameter interactions. Therefore, how to reduce the transmission overhead in model fine-tuning has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a communication method, a terminal device, and a network device. The following describes various aspects related to the embodiments of this application.
[0005] In a first aspect, a communication method is provided, the communication method being applied to a machine learning model, the communication method comprising: a terminal device receiving global model parameters related to a first model sent by a network device; the terminal device sending a first parameter to the network device, the first parameter being determined based on the global model parameters; the terminal device receiving a plurality of candidate parameters sent by the network device; the terminal device determining local model aggregation parameters to be sent to the network device based on the plurality of candidate parameters; wherein the terminal device is one of a plurality of terminal devices, the plurality of candidate parameters are determined based on a plurality of first parameters sent by the plurality of terminal devices, and the plurality of local model aggregation parameters of the plurality of terminal devices are used to train the first model.
[0006] Secondly, a model training method is provided, wherein the communication method is applied to a machine learning model, the communication method comprising: a network device sending global model parameters related to a first model to multiple terminal devices; the network device receiving multiple first parameters sent by the multiple terminal devices, the multiple first parameters being determined based on the global model parameters; the network device sending multiple candidate parameters to the multiple terminal devices; wherein the multiple candidate parameters are used by the multiple terminal devices to determine multiple local model aggregation parameters sent to the network device; the multiple candidate parameters are determined based on the multiple first parameters, and the multiple local model aggregation parameters are used to train the first model.
[0007] Thirdly, a terminal device is provided, which is applied to a machine learning model. The terminal device includes: a transceiver unit for receiving global model parameters related to a first model sent by a network device; the transceiver unit is further configured to send a first parameter to the network device, the first parameter being determined based on the global model parameters; the transceiver unit is further configured to receive multiple candidate parameters sent by the network device; and a processing unit for determining local model aggregation parameters to be sent to the network device based on the multiple candidate parameters; wherein the terminal device is one of multiple terminal devices, the multiple candidate parameters are determined based on multiple first parameters sent by the multiple terminal devices, and the multiple local model aggregation parameters of the multiple terminal devices are used to train the first model.
[0008] Fourthly, a network device is provided, the network device being applied to a machine learning model, the network device comprising: a transceiver unit, configured to send global model parameters related to a first model to multiple terminal devices; the transceiver unit is further configured to receive multiple first parameters sent by the multiple terminal devices, the multiple first parameters being determined based on the global model parameters; the transceiver unit is further configured to send multiple candidate parameters to the multiple terminal devices; wherein the multiple candidate parameters are used by the multiple terminal devices to determine multiple local model aggregation parameters to be sent to the network device; the multiple candidate parameters are determined based on the multiple first parameters, and the multiple local model aggregation parameters are used to train the first model.
[0009] Fifthly, a communication device is provided, including a memory and a processor, the memory for storing a program, and the processor for calling the program in the memory to perform the method as described in the first or second aspect.
[0010] A sixth aspect provides an apparatus including a processor for calling a program from memory to perform the method as described in the first or second aspect.
[0011] A seventh aspect provides a chip including a processor for calling a program from memory, causing a device on which the chip is mounted to perform the method as described in the first or second aspect.
[0012] Eighthly, a computer-readable storage medium is provided having a program stored thereon that causes a computer to perform the method as described in the first or second aspect.
[0013] Ninth aspect, a computer program product is provided, including a program that causes a computer to perform the method as described in the first or second aspect.
[0014] In a tenth aspect, a computer program is provided that causes a computer to perform the method as described in the first or second aspect.
[0015] In this embodiment, multiple terminal devices can determine multiple first parameters based on global model parameters, and can also determine uploaded local model aggregation parameters based on multiple candidate parameters determined by the network device, thereby updating the global model parameters. The global model parameters are used to train the first model. Therefore, multiple terminal devices can selectively transmit and aggregate multiple local model aggregation parameters through multiple candidate parameters, thereby reducing the amount of data transmitted and aggregated during the training process of the first model, which helps to reduce communication overhead and resource consumption. Attached Figure Description
[0016] Figure 1 shows the wireless communication system used in an embodiment of this application.
[0017] Figure 2 is a flowchart illustrating a communication method provided in an embodiment of this application.
[0018] Figure 3 is a flowchart illustrating one possible implementation of the method shown in Figure 2.
[0019] Figure 4 is a flowchart illustrating another possible implementation of the method shown in Figure 2.
[0020] Figure 5 is a schematic diagram of the structure of a large-scale federal low-rank adaptive fine-tuning system based on over-the-air computation provided in an embodiment of this application.
[0021] Figure 6 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application.
[0022] Figure 7 is a schematic diagram of the structure of a control device for the terminal equipment shown in Figure 7.
[0023] Figure 8 is a schematic diagram of the structure of a network device provided in an embodiment of this application.
[0024] Figure 9 is a schematic diagram of the structure of a control device for the network device shown in Figure 8.
[0025] Figure 10 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0026] Figure 11 is a schematic block diagram of a communication device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0028] The embodiments of this application can be applied to various communication systems. For example, they can be applied to Global System for Mobile Communication (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), General Packet Radio Service (GPRS), Long Term Evolution (LTE), Advanced Long Term Evolution (LTE-A), New Radio (NR), evolution systems of NR, LTE-based access to unlicensed spectrum (LTE-U), NR-based access to unlicensed spectrum (NR-U), NTN systems, Universal Mobile Telecommunications System (UMTS), Wireless Local Area Networks (WLAN), Wireless Fidelity (WiFi), and 5th-generation (5G) systems. The embodiments of this application can also be applied to other communication systems, such as future communication systems. The future communication system could be, for example, a 6th-generation (6G) mobile communication system or a satellite communication system.
[0029] Traditional communication systems support a limited number of connections and are easy to implement. However, with the development of communication technology, communication systems can support not only traditional cellular communication but also one or more other types of communication. For example, a communication system can support one or more of the following communication methods: device-to-device (D2D) communication, machine-to-machine (M2M) communication, machine-type communication (MTC), enhanced machine-type communication (eMTC), vehicle-to-vehicle (V2V) communication, and vehicle-to-everything (V2X) communication. The embodiments of this application can also be applied to communication systems that support the above-mentioned communication methods.
[0030] The communication system in this application embodiment can be applied to carrier aggregation (CA) scenarios, dual connectivity (DC) scenarios, and standalone (SA) network deployment scenarios.
[0031] The communication system in this application embodiment can be applied to unlicensed spectrum. This unlicensed spectrum can also be considered a shared spectrum. Alternatively, the communication system in this application embodiment can also be applied to licensed spectrum. This licensed spectrum can also be considered a dedicated spectrum.
[0032] The embodiments of this application can be applied to NTN systems. As an example, the NTN system may include a 4G-based NTN system, an NR-based NTN system, an Internet of Things (IoT)-based NTN system, and a narrowband Internet of Things (NB-IoT)-based NTN system.
[0033] A communication system may include one or more terminal devices. The terminal devices mentioned in the embodiments of this application may also be referred to as user equipment (UE), access terminal, user unit, user station, mobile station, mobile station (MS), mobile terminal (MT), remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0034] In some embodiments, the terminal device may be a station (ST) in a WLAN. In some embodiments, the terminal device may be a cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA) device, handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, in-vehicle device, wearable device, terminal device in a next-generation communication system (e.g., NR system), or terminal device in a future evolved public land mobile network (PLMN) network, etc.
[0035] In some embodiments, the terminal device may be a device that provides voice and / or data connectivity to a user. For example, the terminal device may be a handheld device, an in-vehicle device, etc., with wireless connectivity. As some specific examples, the terminal device may be a mobile phone, tablet, laptop, PDA, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, etc.
[0036] In some embodiments, the terminal device may be deployed on land. For example, the terminal device may be deployed indoors or outdoors. In some embodiments, the terminal device may be deployed on water, such as on a ship. In some embodiments, the terminal device may be deployed in the air, such as on an airplane, balloon, or satellite.
[0037] In addition to terminal devices, the communication system may also include one or more network devices. In this embodiment, the network device can be a device for communicating with the terminal device; this network device may also be referred to as an access network device or a radio access network device. For example, the network device may be a base station. In this embodiment, the network device may refer to a radio access network (RAN) node (or device) that connects the terminal device to the wireless network. A base station can broadly encompass, or be replaced by, various names including: NodeB, evolved NodeB (eNB), next-generation NodeB (gNB), relay station, transmitting and receiving point (TRP), transmitting point (TP), master MeNB, auxiliary SeNB, multi-mode radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), positioning node, etc. A base station can be a macro base station, micro base station, relay node, donor node, or similar, or a combination thereof. A base station can also refer to a communication module, modem, or chip installed within the aforementioned equipment or apparatus. Base stations can also be mobile switching centers, devices that perform base station functions in D2D, V2X, and M2M communications, network-side devices in 6G networks, and devices that perform base station functions in future communication systems. Base stations can support networks using the same or different access technologies. The embodiments of this application do not limit the specific technologies or device forms used in the network equipment.
[0038] Base stations can be fixed or mobile. For example, a helicopter or drone can be configured to act as a mobile base station, and one or more cells can move depending on the location of the mobile base station. In other examples, a helicopter or drone can be configured as a device to communicate with another base station.
[0039] In some deployments, the network device in this application embodiment may refer to a CU or a DU, or the network device may include both a CU and a DU. The gNB may also include an AAU.
[0040] By way of example and not limitation, in the embodiments of this application, the network device may have mobility characteristics; for example, the network device may be a mobile device. In some embodiments of this application, the network device may be a satellite or a balloon station. In some embodiments of this application, the network device may also be a base station located on land, water, or other similar locations.
[0041] In this embodiment, the network device can provide services to a cell. The terminal device communicates with the network device through the transmission resources (e.g., frequency domain resources, or spectrum resources) used by the cell. The cell can be the cell corresponding to the network device (e.g., a base station). The cell can belong to a macro base station or to a base station corresponding to a small cell. The small cell can include: metro cell, micro cell, pico cell, femto cell, etc. These small cells have the characteristics of small coverage area and low transmission power, and are suitable for providing high-speed data transmission services.
[0042] In some embodiments, this application can also be applied to communication systems based on artificial intelligence (AI). One of the goals of 3GPP Release 18 is to enhance the capabilities of 5G and extend its application to new devices, deployments, and industries. As network designs become increasingly complex, encompassing a wide range of deployment and usage options, traditional methods will not provide rapid solutions. Because manually reconfiguring cellular communication systems is costly and inefficient, it is necessary to use artificial intelligence and machine learning (ML) to automate operational processes, thereby reducing costs by automating functions that require human interaction.
[0043] For example, AI and ML can solve complex and unstructured network problems by using large amounts of data collected from wireless networks.
[0044] As an example, AI can be used in the core network and RAN to enable intelligent network operation. For instance, AI can be used to enhance quality of service (QoS), improve efficiency, simplify deployment, and enhance security.
[0045] As an example, on-device artificial intelligence can benefit the entire communication system. One potential enabling capability of AI is radio sensing. AI can provide valuable knowledge through environmental and contextual awareness, thereby reducing overhead and latency. Through radio sensing, communication systems can support enhanced device experiences, such as intelligent beamforming and power management. Furthermore, AI can help improve system performance, such as reducing interference, achieving better spectrum utilization, and improving radio security. For example, AI can facilitate better detection and prevention of malicious attacks.
[0046] For example, Figure 1 is a schematic diagram of the architecture of a communication system provided in an embodiment of this application. As shown in Figure 1, the communication system 100 may include a network device 110, which may be a device that communicates with a terminal device 120 (or a communication terminal, terminal). The network device 110 can provide communication coverage for a specific geographical area and can communicate with terminal devices located within that coverage area.
[0047] Figure 1 illustrates an exemplary network device and two terminal devices. In some embodiments of this application, the communication system 100 may include multiple network devices and each network device may include other numbers of terminal devices within its coverage area, without limitation.
[0048] In this embodiment of the application, the communication system shown in FIG1 also includes other network entities such as a mobility management entity (MME) and an access and mobility management function (AMF), which are not limited in this embodiment of the application.
[0049] It should be understood that devices with communication functions in the network / system of this application embodiment can be referred to as communication devices. Taking the communication system 100 shown in FIG1 as an example, the communication device may include a network device 110 and a terminal device 120 with communication functions. The network device 110 and the terminal device 120 can be the specific devices described above, which will not be repeated here. The communication device may also include other devices in the communication system 100, such as network controllers, mobility management entities, and other network entities. This application embodiment does not limit this.
[0050] To facilitate a detailed explanation of the innovative aspects of the technical solution, some relevant technical knowledge involved in the embodiments of this application is first introduced. The following related technologies are optional solutions and can be arbitrarily combined with the technical solutions of the embodiments of this application, all of which fall within the protection scope of the embodiments of this application. The embodiments of this application include at least some of the following contents.
[0051] With the continuous development of communication technology, the application scenarios of artificial intelligence services are gradually expanding. For example, to meet diverse user needs, sixth-generation wireless networks are expected to demonstrate enhanced intelligence through the integration of artificial intelligence and wireless communication. In particular, the emergence of large models such as generative pre-trained transformers (GPTs) and large language model Meta AI (Llamas) is driving the use of generative artificial intelligence in mobile edge networks to provide personalized and customized AI services.
[0052] Large models refer to machine learning models with a large number of parameters and complex computational structures. Typically, these models are built from deep neural networks, with the number of parameters reaching billions or even hundreds of billions. The purpose of designing large models is to enhance their expressive power and predictive performance, enabling them to handle more complex tasks and data. These models are widely used in various fields, such as natural language processing, computer vision, speech recognition, and recommendation systems. By training on massive amounts of data, large models can learn complex patterns and features to achieve stronger generalization capabilities, thus enabling accurate predictions on unknown data.
[0053] Large-scale model fine-tuning refers to further training a pre-trained large model on a specific task or dataset to improve its performance on that task. Since pre-trained models learn general features, while specific tasks may have specific requirements, fine-tuning is necessary. Fine-tuning allows the model to better adapt to the task, improving performance. Furthermore, the large-scale datasets used for pre-training may differ from real-world application data; fine-tuning helps the model better understand and process domain-specific data.
[0054] Alternatively, fine-tuning large models typically involves less data and computational resources. This is because the model has already learned a large number of general features, and only some features need to be adjusted to adapt to the new task.
[0055] In large-scale model fine-tuning systems, edge devices can perform centralized model fine-tuning or large-scale federated fine-tuning based on federated learning systems. In centralized model fine-tuning, edge devices directly upload locally stored fine-tuning data samples to a centralized server for centralized model parameter fine-tuning. Federated learning systems, on the other hand, enable distributed training of machine learning models for edge intelligent business applications.
[0056] As an example, a federated learning system in an edge network consists of a network device and multiple terminal devices, and is trained through a multi-round process. In each round, the network device broadcasts a global model (such as the aforementioned machine learning model) to all terminal devices; then, each terminal device trains the global model using local data samples, generates a local model, and uploads its local model to the network device via a wireless channel; the network device then aggregates the local models uploaded by all terminal devices and updates the global model. This process is repeated until the global model meets a preset convergence condition, at which point the federated learning is complete.
[0057] Furthermore, in large-scale model fine-tuning systems, network devices and multiple terminal devices can perform wireless transmission based on over-the-air (OTA) computing to improve transmission efficiency. OTA computing is a novel non-orthogonal access technology. Taking the interaction between a base station and terminal devices as an example, in OTA computing, multiple terminal devices preprocess the original signal and then simultaneously transmit information on the same time-frequency resources using the superposition characteristics of the wireless channel; the signal received by the base station is a superposition of the signals transmitted by multiple terminal devices. The base station then performs post-processing on the received superimposed signal to extract the target signal, thereby realizing signal transmission and computation.
[0058] However, various large-scale model fine-tuning systems have technical problems that need to be solved, as follows.
[0059] First, in a centralized large model fine-tuning system, local fine-tuning data samples contain privacy information related to edge devices. Directly uploading local fine-tuning data samples poses a potential risk of privacy leakage to edge devices.
[0060] Secondly, in large-model federated fine-tuning systems, the large model has a large number of parameters, and the federated fine-tuning process requires extremely frequent parameter exchanges. For example, each fine-tuning round requires wireless transmission of model parameters or gradients, and the massive number of parameters in a large model will place a huge transmission burden on the wireless network, severely restricting the feasibility of fine-tuning large models in wireless networks and affecting the normal operation of other wireless services.
[0061] Secondly, in traditional over-the-air computing systems, computation primarily relies on narrowband single-channel aggregation. This single-channel aggregation method results in significantly lower aggregation rates and prolonged aggregation latency, making it difficult to handle the aggregation of a large number of parameters in large models. Furthermore, traditional over-the-air computing systems are mainly deployed on single-antenna base stations and terminal devices, making spatial reuse difficult and leading to low spectrum utilization.
[0062] It should be noted that the problem of large transmission requirements caused by a large number of model parameters and frequent interactions mentioned above is only an example. The embodiments of this application can be applied to any type of model application scenario where there are many model parameters and frequent interactions are required.
[0063] To address the aforementioned issues, this application proposes a communication method for machine learning models. This method allows terminal devices to determine a first parameter related to importance based on received global model parameters. Multiple first parameters from multiple terminal devices can be used by a network device to determine multiple candidate parameters that can be sent, thereby updating the global model parameters of the first model. Therefore, this method, based on federated learning, allows devices participating in model training to selectively transmit and aggregate local model parameters from multiple terminal devices using multiple candidate parameters, reducing the amount of data transmitted and aggregated, thus lowering communication overhead and resource consumption.
[0064] To facilitate understanding, the communication method will be described in detail below with reference to Figure 2.
[0065] Figure 2 illustrates the interaction between the terminal device and the network device. The terminal device can be any type of communication terminal participating in model training. The terminal device can be one of multiple terminal devices. The network device can communicate with multiple terminal devices to train the first model. In some embodiments, the terminal device can train model parameters related to the first model based on local data samples. In some embodiments, the terminal device can receive global model parameters or other relevant parameters broadcast by the network device.
[0066] In some embodiments, the terminal device can be any terminal device in a wireless edge network.
[0067] In some embodiments, the terminal device can be any one of a plurality of terminal devices participating in the training of the first model. For example, the plurality of terminal devices can train the first model together with a network device.
[0068] As an example, multiple terminal devices can be K terminal devices, where K is an integer greater than or equal to 1. For instance, the terminal devices could be a set of terminal devices. The k-th terminal device in the system, 1≤k≤K.
[0069] In some embodiments, the terminal device may be equipped with one or more radio frequency chains and antennas. The number of radio frequency chains may or may not be equal to the number of antennas. For example, the terminal device may be equipped with N u RF chain and N u A single antenna.
[0070] In some embodiments, the terminal device may store multiple data samples to support training and adaptive fine-tuning of the first model. These data samples used for model training may be referred to as local data samples. The data samples used for fine-tuning may also be referred to as fine-tuning data samples.
[0071] Local data samples can be various data samples stored on the terminal device for training the first model. For example, local data samples include, but are not limited to, text, voice, images, and videos, and this application embodiment does not limit them.
[0072] As an example, a terminal device can store a local data sample set with multiple samples. This local data sample set can include a training subset or a learning subset used to train a first model. For example, when K terminal devices participate in model training, each of the K terminal devices can store K training subsets. The K training subsets can each have D1, ..., D2. K There are samples, specifically represented as: D1 = {x} 1,l},…,D K ={x K,l};
[0073] Where, x k,l This is the l-th sample data at the k-th terminal device.
[0074] In some embodiments, some or all of the terminal devices communicating with the network device participate in model training.
[0075] The network device is any of the communication devices described above that provide services to multiple terminal devices. In some embodiments, the network device is a communication device with powerful computing capabilities. For example, the network device may be a base station that broadcasts a global model to multiple terminal devices.
[0076] For example, the network device can receive signals and data sent by multiple terminal devices, including the terminal device shown in Figure 2. For example, the network device can send global parameters for a certain round to the multiple terminal devices. For example, the network device can send resource allocation schemes and / or reporting schemes for model training to the multiple terminal devices.
[0077] In some embodiments, the network device is equipped with multiple radio frequency chains and multiple antennas. The number of radio frequency chains may or may not be equal to the number of antennas. For example, the network device may be equipped with N f RF chain and N r A single antenna.
[0078] In some embodiments, the network device has certain resources to facilitate communication with multiple terminal devices. For example, a base station may have a set of This represents N sub-channels. Optionally, N can be an integer greater than 1.
[0079] The method shown in Figure 2 is applied to the training process of a machine learning model. This training process can include the pre-training process described above, or it can include the model fine-tuning process; neither is limited here. For simplicity, the following description uses the model fine-tuning process as an example.
[0080] Referring to Figure 2, in step S210, the terminal device receives global model parameters related to the first model sent by the network device. The network device can send global model parameters related to the first model to multiple terminal devices.
[0081] The first model can be any AI / ML model that supports communication services. Optionally, the first model can be applied to a wireless edge network. Optionally, the first model can be a machine learning model that can be applied to the embodiments of this application.
[0082] In some embodiments, the first model can be the large model described above. For example, the first model can be GPT or Llmas. As an example, the first model can be a pre-trained large model with frozen parameters. Fine-tuning can be performed via a low-rank adapter.
[0083] The first model can be any one or more of various neural network models, and this application does not limit this. Optionally, the first model includes, but is not limited to: convolutional neural network model, recurrent neural network model, variational autoencoder neural network model, etc.
[0084] Global model parameters associated with the first model can be used to fine-tune the first model. The global model parameters sent by the network device can be relative to the local model parameters on the terminal device side. Optionally, both the global model parameters and the local model parameters can be low-rank adapters. The global model parameters can be replaced with a global low-rank adapter.
[0085] In some embodiments, the global model parameters may further include a global low-rank adapter for fine-tuning the first model. The global model parameters can be used on all training devices to fine-tune the first model. The global model parameters can be represented by the gradient of the global low-rank adapter. Conversely, the local model parameters may include a local low-rank adapter for fine-tuning the first model on the terminal device. The local model parameters can be represented by the gradient of the local low-rank adapter.
[0086] In some embodiments, the global low-rank adapter associated with the first model may refer to a low-rank adapter used to fine-tune some or all of the models in the first model, or a global low-rank adapter determined according to the first model.
[0087] In some embodiments, the low-rank adapter includes, but is not limited to, convolutional neural network models, linear neural network models, diffusion neural network models, etc. This application does not limit the specific implementation of such models.
[0088] As an example, when the first model is a large model, considering the matrix multiplication characteristics in large models, a low-rank adapter matrix is injected during training. The pre-trained large model W0 with parameters frozen.
[0089] As an example, a low-rank adapter can consist of multiple trainable low-dimensional matrices. That is, a low-rank adapter can include multiple low-dimensional matrices. Optionally, a low-rank adapter can consist of the multiplication of multiple low-dimensional matrices.
[0090] For example, when the first model is At that time, the low-rank adapter ΔW can be composed of two trainable low-dimensional matrices. and Multiplication results in ΔW = BA T .
[0091] As an example, during fine-tuning, the pre-trained large model W0 and the low-rank adapter ΔW are multiplied by the same input x, and their respective output vectors are summed by coordinates to obtain y, i.e., y = W0x + BA. T x. The fine-tuned large model can be expressed as W = W0 + ΔW = W0 + BA T .
[0092] In some embodiments, the global model parameters are determined by the network device. Exemplarily, the network device may integrate the gradients from local low-rank adapters sent by multiple terminal devices in the previous cycle to determine the global low-rank adapter for the current cycle. Exemplarily, for the first cycle of the first model, the network device may determine an initialized global low-rank adapter based on the first model. Exemplarily, for the last cycle, the fine-tuned global low-rank adapter may be the final low-rank adapter for model fine-tuning.
[0093] As an example, a network device can broadcast global model parameters so that all terminal devices involved in model fine-tuning can receive them. In other words, terminal devices can receive global model parameters sent by the network device via broadcast.
[0094] In some embodiments, the training process of the first model may include multiple training epochs. The epoch currently in which the model is being trained may be called the current epoch. The epoch immediately preceding the current epoch may be called the previous epoch. Optionally, the fine-tuning process of the first model may include multiple fine-tuning epochs. A fine-tuning epoch may refer to a fine-tuning round or a training process, also known as a fine-tuning round or learning round. For example, the t-th fine-tuning round is the fine-tuning epoch t. Similarly, in a federated fine-tuning system, the fine-tuning epoch t is the t-th federated fine-tuning round.
[0095] As an example, a fine-tuning cycle may include a global low-rank adapter broadcast cycle, a local fine-tuning cycle on the terminal device side, a federated rank importance analysis cycle on the network device side, a federated gradient aggregation cycle on the network device side, and a global low-rank adapter update cycle. Optionally, the aforementioned global low-rank adapter broadcast cycle, local fine-tuning cycle, federated rank importance analysis cycle, federated gradient aggregation cycle, and global low-rank adapter update cycle are sequentially related.
[0096] For the sake of brevity, the following text will use period t as an example. Period t (the t-th period) can be either the fine-tuning period t or the training period t.
[0097] As an example, upon entering the current period t, K terminal devices receive a global low-rank adapter sent by the network device. The global low-rank adapters sent by network devices in the current period t can be obtained from the previous period (i.e., t-1). For example, during the global low-rank adapter broadcast period in the t-th period, the base station can use the global low-rank adapters aggregated in the (t-1)-th period. The broadcast is sent to all terminal devices.
[0098] In some embodiments, the number of cycles in the entire fine-tuning process can be determined by the network device. The number of cycles can also be referred to as the number of iterations. For example, after determining that the number of fine-tuning rounds has reached the preset maximum number of fine-tuning rounds, the base station broadcasts a training stop instruction to all terminal devices participating in the fine-tuning. As another example, after determining that the fine-tuning has converged, the base station broadcasts a training stop instruction to all terminal devices participating in the fine-tuning.
[0099] For example, when T is the maximum number of fine-tuning rounds, if the current period t = T, the network device determines that the preset maximum number of fine-tuning rounds has been reached.
[0100] In some embodiments, the broadcasting of a global low-rank adapter by a network device to multiple terminal devices may be a process included in the current cycle. For example, the process of receiving a global low-rank adapter may be used to determine the start time of the current cycle. For instance, the network device broadcasting the global low-rank adapter may indicate the start of the current cycle. Alternatively, the global low-rank adapter may be broadcast only after the current cycle has started.
[0101] In some embodiments, the network device broadcasting the global low-rank adapter to multiple terminal devices may not be part of the current cycle's process. For example, the current cycle only begins after the terminal devices participating in model fine-tuning have obtained the global low-rank adapter.
[0102] After executing step S210, the terminal device can also determine local model parameters (not shown in Figure 2) based on the received global model parameters. In model fine-tuning, this process can be referred to as the local fine-tuning process. The local fine-tuning process may also include calculating the first parameter. As an example, the terminal device can train the global model parameters using a local data sample set.
[0103] In some embodiments, when the global model parameter is a global low-rank adapter, the local model parameter may include a local low-rank adapter. Local low-rank adapters include, but are not limited to, convolutional neural network models, linear neural network models, and diffusion neural network models. This application does not limit the scope of these adapters.
[0104] As an example, a local low-rank adapter can include R column vectors, ordered as follows:
[0105] In some embodiments, the terminal devices can train global model parameters using samples from all local data sample sets used for fine-tuning. For example, in the local cycle of the t-th fine-tuning round, K terminal devices can use local data sample sets D1, ..., D K Training a global low-rank adapter
[0106] In some embodiments, the first parameter is used to indicate the importance of some or all of the parameters in the local model parameters corresponding to the terminal device. As an example, when the local model parameters include a local low-rank adapter, the first parameter can indicate the local rank importance corresponding to the local low-rank adapter. For example, the first parameter can be a local rank importance metric parameter. The following mainly uses local rank importance as an example.
[0107] In some embodiments, the terminal device may determine a first dataset for training global model parameters. The first dataset may be a subset of the terminal device's training set. For example, the first dataset may be a subset of a locally fine-tuned data sample set.
[0108] As an example, the data samples in the first dataset can be determined according to certain rules. For instance, the terminal device can select multiple data samples from the training subset based on the type of the first model or fine-tuning requirements to form the first dataset.
[0109] As an example, the data samples in the first dataset can be data samples randomly drawn by the terminal device from the training subset.
[0110] In some embodiments, when training and computing the global low-rank adapter, the loss function of the terminal device can be determined by the local data function of multiple data samples. For example, for the k-th terminal device, terminal device k is randomly sampled. The first dataset consists of several fine-tuned sample data. When, the loss function of the terminal device in the t-th period can be:
[0111] in, It is the global low-rank adapter obtained from the previous cycle aggregation, x k,l It is the l-th sample from terminal device k, 1≤l≤L, f(ΔW) t ;W0,x k,l ) is about sample x k,l The local loss function.
[0112] In some embodiments, the terminal device can determine the local model parameters and the first parameter based on the global model parameters. As an example, the terminal device can determine the local low-rank adapter gradient and the local rank importance based on the global low-rank adapter.
[0113] As an example, by training a global low-rank adapter, a terminal device can obtain the gradient of a local low-rank adapter. For instance, K terminal devices can train using a local data sample set. Afterwards, we can obtain K local low-rank adapter gradients. The K local low-rank adapter gradients are as follows:
[0114] As an example, after determining the first dataset, the terminal device can determine the local rank importance and the local low-rank adapter gradient based on the first dataset and the global low-rank adapter.
[0115] In some embodiments, the local model parameters are local low-rank adapter gradients comprising multiple matrix gradients. For example, the local low-rank adapter gradients include a first matrix gradient and a second matrix gradient, which are sample gradients of the global model parameter correlation matrix. The correlation matrix of the global model parameters can be one or more matrices used to indicate the global model parameters. For example, when the global model parameters are a global low-rank adapter, the correlation matrix can be the low-dimensional matrix that constitutes the low-rank adapter as described above.
[0116] As an example, the local low-rank adapter gradient of the k-th terminal device with respect to low-dimensional matrices At and Bt. and It can be:
[0117] in, and These are the sample gradients with respect to the low-dimensional matrices At and Bt, respectively.
[0118] As an example, the local rank importance metric parameter of the k-th terminal device It can be:
[0119] in, and They represent gradients respectively. and The r-th column vector, 1≤r≤R; L and Q are the gradients. and Quantity. Gradient and It can also be represented as: and
[0120] For K terminal devices, each terminal device can perform local rank importance calculation (first parameter). The local rank importance calculated by each terminal device needs to be sent to the network device (step S220). After multiple terminal devices have sent their respective local rank importance to the network device, the network device can perform federated rank importance analysis.
[0121] In step S220, the terminal device sends a first parameter to the network device. The network device can receive multiple first parameters sent by multiple terminal devices participating in model fine-tuning. For example, K terminal devices can each send their local rank importance to the network device. For instance, terminal device k among the K terminal devices can send its local rank importance metric parameter. Report to the base station.
[0122] The first parameter is determined by the terminal device based on the global model parameters. For example, the local rank importance is determined based on the global low-rank adapter. After receiving the global low-rank adapter, the terminal device can calculate the local rank importance based on the training of the global low-rank adapter.
[0123] In some embodiments, the terminal device may transmit other signals along with the first parameter. For example, the terminal device may transmit the first parameter and a pilot signal to the network device. The pilot signal may be used to determine the channel coefficients.
[0124] As an example, the local rank importance metric parameter and the pilot signal are signals sent by each terminal device and received by the base station.
[0125] As an example, each terminal device sends orthogonal pilot signals, and the base station obtains the channel coefficients of each terminal device based on the received pilot signals.
[0126] As an example, methods for obtaining channel coefficients include, but are not limited to, direct sequence spread spectrum signals, zero-forcing channel estimation, and least squares channel estimation. This application does not impose any limitations on these methods.
[0127] In some embodiments, the network device can determine the channel coefficients of multiple terminal devices based on multiple pilot signals from multiple terminal devices. For example, each terminal device can send a pilot signal when transmitting a first parameter, so that the network device can determine the channel coefficients.
[0128] In step S230, the terminal device receives multiple candidate parameters sent by the network device. The network device may send multiple candidate parameters to multiple terminal devices. In some embodiments, the multiple candidate parameters may be determined based on federated importance analysis performed by the network device. This process may belong to the federated importance analysis cycle described above.
[0129] In some embodiments, multiple candidate parameters are used by the terminal device to determine the local model parameters for aggregation, i.e., local model aggregation parameters. Multiple candidate parameters are used by the terminal device to select the local model aggregation parameter from the local model parameters; therefore, multiple candidate parameters can also be called selection parameters or choice parameters. When the global model parameter is a global low-rank adapter, multiple candidate parameters can be used to indicate the rank selection scheme.
[0130] In some scenarios, multiple candidate parameters can also be replaced by a rank selection scheme. The terminal device can determine the relevant parameters of the local low-rank adapter to be transmitted based on the rank selection scheme sent by the network device, such as the local low-rank adapter aggregated gradient.
[0131] In some embodiments, multiple candidate parameters can be determined based on multiple first parameters sent by multiple terminal devices. As an example, the rank selection scheme is determined based on multiple local rank importances sent by multiple terminal devices. That is, after receiving multiple local rank importances from multiple terminal devices, the network device can perform federated importance analysis based on these multiple local rank importances.
[0132] As an example, during the federated rank importance analysis period in the t-th period, the network device can perform federated rank importance analysis based on the local rank importance metrics reported by the K terminal devices to obtain the rank selection scheme {δt,r}. Optionally, when δt,r = 1, the K terminal devices upload and aggregate the r-th column vector of their local low-rank adapter gradients to the base station; when δt,r = 0, the K terminal devices do not upload the r-th column vector of their local low-rank adapter gradients to the base station. The reverse is also true.
[0133] In some embodiments, network devices can broadcast multiple candidate parameters to facilitate reception by terminal devices participating in model fine-tuning. For example, after determining the rank selection scheme, the base station broadcasts the rank selection scheme to each terminal device.
[0134] As an example, any rank in a rank selection scheme can be represented as the corresponding column vector r. With the introduction of a rank selection scheme, the aggregation typically involves only partial rank information, leading to a lack of gradient rank information. Terminal and network devices can utilize this missing gradient rank information to measure the fine-tuning error caused by the lack of column aggregation of local gradients.
[0135] As an example, the lack of gradient rank information can be caused by Γ t Indicated. Optionally, Γ t It can be determined as follows:
[0136] in,
[0137] In some embodiments, gradient information distortion may occur because only partial gradient information is available during aggregation. Optionally, a balance can be struck between missing gradient rank information and gradient information distortion to improve model fine-tuning performance.
[0138] As an example, gradient information distortion can be represented by Λt. Optionally, Λt can be determined as:
[0139] in, Π represents the mean squared error. t =max r {δ t,r (L|b t,r | 2 +Q|a t,r | 2 )}.
[0140] In some embodiments, multiple local rank importances are also used by network devices to determine resource allocation schemes, which will be explained later with reference to Figure 4.
[0141] In step S240, the terminal device determines the local model aggregation parameter to be sent to the network device based on multiple candidate parameters. When the local model aggregation parameter is a local low-rank adapter aggregation gradient, the multiple local low-rank adapter aggregation gradients sent by multiple terminal devices can be used for gradient aggregation. This process belongs to the federated gradient aggregation cycle described above.
[0142] In some embodiments, the local model aggregation parameters can be some or all of the parameters in the local model parameters. As an example, the local low-rank adapter aggregation gradient can be some or all of the gradients used for aggregation in the local low-rank adapter gradients. When the local low-rank adapter aggregation gradient is only a portion of the gradients in the local low-rank adapter gradients, gradient aggregation may result in a lack of gradient rank information.
[0143] In some embodiments, the local model aggregation parameters include at least one column vector from the local model parameters. Multiple candidate parameters are used to determine this at least one column vector. As an example, the local low-rank adapter aggregated gradient includes at least one column vector from the local low-rank adapter gradient, and a rank selection scheme is used by the terminal device to determine this at least one column vector. That is, the rank selection scheme is used to indicate one or more column vectors of the parameter aggregation so that the terminal device can determine the local low-rank adapter aggregated gradient. For example, when the local low-rank adapter gradient includes R column vectors, if the rank selection scheme of the r-th column vector is 1, then the gradient corresponding to the r-th column vector belongs to the local low-rank adapter aggregated gradient.
[0144] As an example, within the federated gradient aggregation period of the t-th period, K terminal devices construct K local low-rank adapter aggregated gradients according to the rank selection scheme {δt,r}. For example, the local low-rank adapter aggregated gradient constructed by terminal device k according to the rank selection scheme {δt,r}. and It can be represented as:
[0145] In some embodiments, multiple local model aggregation parameters can be aggregated and transmitted via over-the-air computation. As an example, multiple local low-rank adapter aggregation gradients can be aggregated and transmitted via over-the-air computation. That is, multiple terminal devices can transmit multiple local low-rank adapter aggregation gradients to the network device via over-the-air computation. As mentioned earlier, during the transmission process of over-the-air computation, multiple terminal devices can utilize the superposition characteristics of wireless channels to transmit information simultaneously on the same time-frequency resources; the signal received by the network device is already the result of multiple signals superimposed.
[0146] As an example, over-the-air computing can be implemented based on massively multi-input multiple-output and / or orthogonal frequency division multiplexing.
[0147] As an example, since the aggregated gradients of multiple local low-rank adapters used for transmission are selected according to a rank selection scheme, gradient information is susceptible to distortion. Terminal and network devices can utilize this gradient distortion to measure fine-tuning errors caused by distortion in the aggregated signal computed over the air.
[0148] In some embodiments, the aggregation of multiple local model aggregation parameters is achieved through multiple signals. That is, all signals used for aggregation can be referred to as the multiple signals corresponding to the aggregation. These multiple signals can be gradient signals. As an example, gradient aggregation of multiple local low-rank adapter aggregation gradients can be achieved based on multiple gradient signals. That is, multiple terminal devices process the multiple local low-rank adapter aggregation gradients into gradient signals respectively, and then perform aggregation and transmission based on over-the-air computation.
[0149] In some embodiments, the terminal device can normalize multiple local model aggregation parameters to obtain multiple signals corresponding to the aggregation. As an example, the terminal device can normalize the local low-rank adapter aggregation gradient to determine the gradient signal. For multiple terminal devices, the multiple gradient signals used for aggregation are determined by normalizing the aggregated gradients of multiple local low-rank adapters.
[0150] In the above embodiments, the normalization process includes, but is not limited to, data regularization, batch normalization, etc., and is not limited in the embodiments of this application.
[0151] As an example, terminal device k performs local low-rank adapter gradient aggregation. and After normalization, the gradient signal s can be obtained for aggregation and transmission. t,k The terminal device transmits gradient signal s. t,k This is used to achieve the transfer of aggregated gradients in the local low-rank adapter.
[0152] Assumption Represents the gradient signal s t,k It is a vector of length M. The transmitted signal of terminal device k can be represented as: x t,k,n =V t,k,n s t,k,n ;
[0153] in, This represents the transmit precoding matrix of terminal device k on subchannel n, also known as the transmission precoding matrix.
[0154] Accordingly, the signal received by the network device on sub-channel n can be represented as: y t,n =∑ k∈K H t,k,n x t,k,n +n t,n ;
[0155] Among them, H t,k,n Let k be the channel matrix from terminal device k to network device k. Let n be the vector of additive white Gaussian noise on subchannel n.
[0156] In some embodiments, multiple local model aggregated parameters are used to determine global model aggregated parameters. The global model aggregated parameters are used by the network device to update the global model parameters for the current period. The update of the global model parameters is used to train a first model. As an example, multiple local low-rank adapter aggregated gradients are used to determine global low-rank adapter aggregated gradients. The global low-rank adapter aggregated gradients are used by the network device to update the global low-rank adapter, and the update of the global low-rank adapter is used to fine-tune the first model. When multiple local low-rank adapter aggregated gradients are transmitted via over-the-air computation, the network device receives the aggregated signal to determine the global low-rank adapter aggregated gradient.
[0157] As an example, network devices can be based on the digital receive precoding matrix W t,n The network device receives aggregated signals using the analog receive precoding matrix Ut. When receiving signals based on the precoding matrix, the signals received by the network device in sub-channel n are aggregated as follows:
[0158] in, This represents the digital receive precoding matrix of the network device. This represents the simulated receive precoding matrix.
[0159] In the example above, and The average mean square error between them can be given by the following formula:
[0160] The mean square error on each assigned sub-channel is:
[0161] Among them, a t This represents the minimum number of sub-channels. T max T represents the maximum transmission delay. symbol Represents a symbolic number.
[0162] As an example, the digital receive precoding matrix on the network device side can be determined based on the data transmission precoding matrix on the terminal device side and the analog receive precoding matrix on the network device side. For example, for terminal device k, the digital receive precoding matrix of the network device can be:
[0163] in,
[0164] As an example, the data transmission precoding matrix on the terminal device side also needs to consider the digital and analog receive precoding matrices on the network device side. For example, the digital transmission precoding matrix of terminal device k can be:
[0165] in, μ t,k It is a Lagrange multiplier.
[0166] Optionally, the methods for solving for Lagrange multipliers include, but are not limited to, the bisection method, the iterative method, and Newton's method. This application does not limit the specific methods used.
[0167] In some embodiments, when a network device receives the aggregated signal, it can obtain the aggregated gradient vector through an inverse normalization operation. For example, The aggregated gradient vector can be obtained through the inverse normalization operation.
[0168] For example, denormalization processing includes, but is not limited to, data deregulation, batch denormalization, etc., and is not limited in the embodiments of this application.
[0169] In some embodiments, global model parameters are obtained by updating global model aggregated parameters. For example, the global low-rank adapter can be obtained by updating the global low-rank adapter aggregated gradient, which belongs to the global low-rank adapter update cycle described above. Multiple local low-rank adapter aggregated gradients of multiple terminal devices can fine-tune the first model through global low-rank adapter aggregated gradients. That is, the network device can obtain the global low-rank adapter aggregated gradient through multiple local low-rank adapter aggregated gradients, and then update the current global low-rank adapter.
[0170] As an example, during the global low-rank adapter update cycle in the t-th cycle, the network device updates the received aggregate gradient. and Update the current global low-rank adapter to obtain the global low-rank adapter ΔW for the next cycle. t+1 .
[0171] As an example, the global low-rank adapter aggregates the gradient as follows:
[0172] Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and B t This represents the two (low-dimensional) matrices of the global low-rank adapter (obtained in the previous cycle). and Represent matrix A respectively t and B t The aggregation gradient.
[0173] The previous section introduced a method for training a first model using federated learning and multiple candidate parameters on network devices and multiple terminal devices. When this method is combined with a low-rank adapter for fine-tuning the first model, it enables low-rank adaptive fine-tuning of the first model. For ease of understanding, the following example uses a large-scale model low-rank adaptive fine-tuning process based on over-the-air computation, illustrated in Figure 3. Figure 3 is a schematic diagram of the large-scale model federated low-rank adaptive fine-tuning process. The current fine-tuning round in Figure 3 is round t.
[0174] As shown in Figure 3, the process of large-scale federated low-rank adaptive fine-tuning based on over-the-air computing includes steps S310 to S370. The flowchart shown in Figure 3 includes the interaction process between the base station and multiple terminal devices, as well as the process between the base station and the multiple terminal devices themselves.
[0175] In step S310, the current fine-tuning round is entered, and the base station broadcasts the global low-rank adapter obtained from the gradient mixing and aggregation of the previous round. One fine-tuning round is also called one fine-tuning cycle. One fine-tuning cycle includes one global low-rank adapter broadcasting cycle, one local fine-tuning cycle, one federated rank importance analysis cycle, one federated gradient aggregation cycle, and one global low-rank adapter update cycle.
[0176] In step S320, each terminal device randomly selects a local fine-tuning data sample D. k In The mini-batch dataset consists of several fine-tuned sample data. This mini-batch dataset is also known as the first dataset; for the k-th terminal device, the first dataset is... Subsequently, the k-th terminal device can determine the global low-rank adapter based on the received data. Calculate local low-rank adapter gradients with mini-batch datasets and And calculate and report local rank importance measurement parameters.
[0177] In step S330, the base station uses the local rank importance metric parameters reported by each terminal device and the channel state information H at each terminal device as a reference. t,k,n Optimize and obtain the rank selection scheme and resource allocation scheme, and broadcast the rank selection scheme and resource allocation scheme to each terminal device.
[0178] In step S340, each terminal device constructs a local low-rank adapter aggregation gradient according to the rank selection scheme and resource allocation scheme, and transmits it to the base station on the same time-frequency resources based on over-the-air computation.
[0179] In step S350, the base station receives the aggregated low-rank adapter gradient and updates the global low-rank adapter using the aggregated low-rank adapter gradient.
[0180] In step S360, it is determined whether the preset maximum number of fine-tuning rounds has been reached. That is, the base station determines whether the large-scale federated low-rank adaptive fine-tuning of the over-the-air computation model has reached the preset maximum number of training rounds. If yes, proceed to step S370; if no, return to step S320.
[0181] In step S370, the large-model federated low-rank adaptive fine-tuning based on over-the-air computation is terminated. For example, the base station may broadcast a fine-tuning stop command to all terminal devices after determining that the round of fine-tuning has reached a preset maximum value.
[0182] As shown in Figure 3, the large-model federated low-rank adaptive fine-tuning system based on over-the-air computation in this application embodiment can significantly reduce the demand for wireless resources in large-model federated low-rank adaptive fine-tuning in wireless networks through over-the-air computation. Furthermore, by selecting the importance of federated rank, the system can achieve a balance between the lack of gradient rank information and the distortion of gradient information, effectively mitigating the impact of wireless channel fading pre- and noise on the fine-tuning process during over-the-air computation, and improving the performance of the global low-rank adapter. Simultaneously, none of the terminal devices transmit raw data, effectively preventing privacy leaks of the terminal devices.
[0183] The above section, with reference to Figures 2 and 3, introduced the federated model training method and related communication methods. Taking the low-rank adaptive fine-tuning process applying this method as an example, on the one hand, in each round of fine-tuning, the terminal device obtains the local low-rank adapter gradient based on the global low-rank adapter broadcast by the network device and calculates the local rank importance; on the other hand, the network device performs federated rank importance analysis based on the local rank importance reported by the terminal device to obtain the rank selection scheme for low-rank adapter gradient aggregation. In this method, the column importance of the local gradient can be measured based on the scheme of federated rank importance analysis to reduce gradient aggregation.
[0184] Multiple first parameters are also used by the network device to determine a resource allocation scheme for transmitting multiple local model aggregation parameters across multiple terminal devices. Taking local rank importance and local low-rank adapter aggregation gradient as examples, multiple local rank importance can be used by the network device to determine the transmission resources for multiple terminal devices. For example, a terminal device can send its local low-rank adapter aggregation gradient based on the resources allocated by the network device. In other words, the network device can allocate transmission resources for local low-rank adapter aggregation gradients to multiple terminal devices.
[0185] As an example, a network device can allocate transmission resources to multiple terminal devices based on multiple received first parameters. In other words, the multiple first parameters from multiple terminal devices are also used to determine a resource allocation scheme for multiple local model aggregation parameters.
[0186] In some embodiments, the network device determines multiple candidate parameters and a resource allocation scheme, which are then broadcast together. As an example, the resource allocation scheme determined by the network device may be broadcast together with a rank selection scheme.
[0187] As mentioned above, the first parameter is used to determine multiple candidate parameters and resource allocation schemes. These multiple candidate parameters are related to the local model aggregation parameters transmitted by multiple terminal devices and network devices through the transmit / receive precoding matrix. The following section uses the rank selection scheme of the multiple candidate parameters as an example to illustrate the determination method and correlation of the multiple candidate parameters, resource allocation scheme, and transmit / receive precoding matrix.
[0188] As an example, multiple terminal devices can formulate a transmit precoding matrix (e.g., a digital transmission precoding matrix) based on a rank selection scheme and a resource allocation scheme, and network devices can specify a receive precoding matrix based on a rank selection scheme and a resource allocation scheme.
[0189] As an example, the transmission resources allocated by a network device to multiple terminal devices may include orthogonal frequency division multiplexing (OFDM) channels or sub-channels. Exemplarily, the resource allocation scheme may include a channel allocation scheme or a sub-channel allocation scheme. For example, a sub-channel allocation scheme {α} t,r}
[0190] To facilitate network devices in determining resource allocation schemes for multiple terminal devices, this application also proposes a resource allocation mechanism for model fine-tuning. Taking the rank selection scheme as an example, this resource allocation mechanism designs a heuristic backtracking algorithm to optimize the rank selection scheme and the resource allocation scheme. Exemplarily, this resource allocation mechanism can optimize some or all of the information in the rank selection scheme, the resource allocation scheme, and the transmission precoding matrices of the terminal devices and network devices through multiple rounds. For example, when the resource allocation mechanism of this application is applied to a large-model federated low-rank adaptive fine-tuning system based on over-the-air computation, the rank selection scheme {δ...} described above... t,r}, Channel allocation scheme {α t,r}, Transmit precoding matrix V t,k,n Digital receiver precoding matrix W t,n and the analog precoding matrix U t It can be obtained through optimization.
[0191] In some embodiments, the network device optimizes the rank selection scheme, resource allocation scheme, and associated precoding matrix through multiple iterations. When the resource allocation scheme includes a sub-channel allocation scheme, due to the low dimension of R, the backtracking-based allocation method can rank the local ranks according to the joint rank importance analysis and based on the aforementioned mean square error (MSE). t,n To arrange the sub-channels.
[0192] In some embodiments, the rank selection scheme (i.e., multiple candidate parameters) and the resource allocation scheme are related to the first precoding matrix. The first precoding matrix includes one or more of the following: a digital receive precoding matrix of the network device, an analog receive precoding matrix of the network device, and a digital transmit precoding matrix (i.e., a transmit precoding matrix) of the terminal device.
[0193] As an example, before determining the rank selection scheme, resource allocation scheme, and first precoding matrix, the network device can first initialize the resource allocation, rank selection scheme, digital receive precoding matrix, analog receive precoding matrix, and digital transmit precoding matrix.
[0194] In the example above, when the resource allocation scheme is a sub-channel allocation scheme, the sub-channel allocation can be initialized as follows: for any sub-channel n, α t,r =1. The rank selection scheme can be initialized as follows: for any rank r, δ = 1. t,r =1, initialize the iteration round e=0. Each element in the digital and analog receiver precoding matrices can be initialized as a matrix with a constant modulus, an amplitude of 1, and a random phase.
[0195] As an example, Ψ can be set during optimization iterations. t,r ≥Ψ tr′ and Furthermore, when MSE t,n ≤MSE t,n′ , When that happens, update e = e + 1.
[0196] In some embodiments, when the resource allocation scheme includes a sub-channel allocation scheme, the association between the rank selection scheme (i.e., multiple candidate parameters) and the sub-channel allocation scheme with the first precoding matrix includes one or more of the following: the rank selection scheme and the sub-channel allocation scheme are determined based on a given first precoding matrix; the first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and the rank selection scheme; when the rank selection scheme and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix; when the rank selection scheme and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; when the rank selection scheme and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix; when the analog receive precoding matrix, the digital receive precoding matrix, and the digital transmit precoding matrix in the first precoding matrix are determined, the rank selection method and the sub-channel allocation scheme are optimized.
[0197] As an example, the rank selection scheme, sub-channel allocation scheme, and at least one of the first precoding matrices can be optimized using a backtracking algorithm. For instance, given a gradient aggregation rank selection scheme, optimization and resource allocation can be performed on the digital receive precoding matrix, analog receive precoding matrix, and digital transmit precoding matrix in over-the-air computation. Similarly, given the digital receive precoding matrix, analog receive precoding matrix, and digital transmit precoding matrix in over-the-air computation, optimization can be performed on the rank selection scheme for low-rank adapter gradient aggregation and the orthogonal frequency division multiplexing sub-channel allocation scheme.
[0198] As an example, given a subchannel allocation and rank selection scheme, optimize the digital receive precoding matrix based on a given analog receive precoding matrix and a digital transmit precoding matrix.
[0199] As an example, given a subchannel allocation and rank selection scheme, the analog receive precoding matrix is optimized based on a given digital receive precoding matrix and a digital transmit precoding matrix.
[0200] As an example, methods for optimizing the analog reception precoding matrix include, but are not limited to, manifold optimization and continuous convex approximation optimization. This application does not limit these methods in its embodiments.
[0201] As an example, given a subchannel allocation and rank selection scheme, optimize the digital transmission precoding matrix based on a given digital receive precoding matrix and an analog receive precoding matrix.
[0202] As an example, given a digital receive precoding matrix, an analog receive precoding matrix, and a digital transmit precoding matrix, we optimize the subchannel allocation scheme and the rank selection scheme.
[0203] In some embodiments, the rank selection scheme, resource allocation scheme, and first precoding matrix are further determined based on the missing gradient rank information and gradient distortion of the first model. When training or optimizing the first model, the missing gradient rank information and gradient distortion of the first model can be information obtained in any training epoch or any optimization epoch.
[0204] As an example, gradient rank information loss and gradient information distortion can be used to determine whether an optimization iteration has converged. For instance, a network device can determine whether the joint result of gradient rank information loss and gradient information distortion has reached a minimum; if not, further optimization is needed. The joint result of gradient rank information loss and gradient information distortion can be represented by the optimal upper bound (Γ). t +Λ t ) (e) To express.
[0205] As an example, a t This represents the minimum number of sub-channels allocated given a maximum transmission delay, which can be obtained during the allocation process. Middle front a t Sub-channels and aggregation The optimal upper bound (Γ) of the first Re column t +Λ t ) (e) .
[0206] As an example, the current optimal upper bound (Γ) t +Λ t ) (e) >(Γ t +Λ t ) (e-1) When, the optimal upper bound (Γ) can be determined. t +Λ t ) (e-1) It reaches the minimum value.
[0207] As an example, when determining the optimal upper bound (Γ) t +Λ t ) (e) After reaching the minimum value, the rank selection scheme {δ} obtained by the network device t,r} and channel allocation scheme {α t,r This refers to the rank selection and resource allocation scheme. Network devices can also obtain an optimized digital transmission precoding matrix V. t,k,n Receive digital precoding matrix W t,n and the analog precoding matrix U t .
[0208] For ease of understanding, the resource allocation mechanism of this application embodiment is described below using a large-scale federated low-rank adaptive fine-tuning resource allocation mechanism based on aerial computing as an example, in conjunction with the flowchart in Figure 4. The current fine-tuning round in Figure 4 is the t-th round.
[0209] As shown in Figure 4, the resource allocation mechanism includes steps S410 to S470. The resource allocation method shown in Figure 4 is mainly executed by the base station, while the terminal devices can coordinate resource allocation.
[0210] In step S410, the base station receives local rank importance metric parameters transmitted from each terminal device and obtains the channel coefficients of each terminal device based on the received pilot signals. For example, the local rank importance metric parameter reported by the k-th terminal device is...
[0211] In step S420, given the sub-channel allocation and rank selection scheme, the given analog receive precoding matrix and digital transmit precoding matrix, the digital receive precoding matrix is optimized.
[0212] In step S430, given the sub-channel allocation and rank selection scheme, the given digital receive precoding matrix and digital transmit precoding matrix, the analog receive precoding matrix is optimized.
[0213] In step S440, given the sub-channel allocation and rank selection scheme, the given digital receive precoding matrix and analog receive precoding matrix, the digital transmit precoding matrix is optimized.
[0214] In step S450, given the digital receive precoding matrix, the analog receive precoding matrix, and the digital transmit precoding matrix, the subchannel allocation and rank selection scheme is optimized.
[0215] In step S460, determine whether convergence or the maximum number of iterations has been reached. If yes, proceed to step S470; otherwise, continue iterating from step S420 to step S450.
[0216] In step S470, the base station obtains the sub-channel allocation and rank selection scheme, as well as the digital receive precoding matrix, the analog receive precoding matrix, and the digital transmit precoding matrix.
[0217] As shown in Figure 4, the resource allocation mechanism of the large-scale federated low-rank adaptive fine-tuning system based on over-the-air computing in this application embodiment can achieve collaborative resource allocation between terminal devices and base stations, improving computational and transmission efficiency. Specifically, the digital receive precoding matrix, analog receive precoding matrix, and digital transmit precoding matrix are obtained iteratively in each optimization round of sub-channel allocation and rank selection. For example, the sub-channel allocation and rank selection optimization is obtained from the given digital receive precoding matrix, analog receive precoding matrix, and digital transmit precoding matrix using a backtracking-based optimization method, thereby achieving the joint minimization of gradient rank information loss and gradient information distortion.
[0218] The above section, with reference to Figure 4, introduces a resource allocation method. This method enables collaborative resource allocation among terminal devices, improving computational and transmission efficiency. Taking the rank selection scheme as an example, on the one hand, given a gradient aggregation rank selection scheme, the method optimizes and allocates resources for the digital receiver precoding matrix, analog receiver precoding matrix, and digital transmitter precoding matrix in large-scale multiple-input multiple-output (MLMI) and orthogonal frequency division multiplexing (OFDM) over-the-air computation. On the other hand, given the digital receiver precoding matrix, analog receiver precoding matrix, and digital transmitter precoding matrix in over-the-air computation, the method optimizes the low-rank adapter gradient aggregation rank selection scheme and the OFDM sub-channel allocation scheme to efficiently and accurately complete the large-scale federated low-rank adaptive fine-tuning of the over-the-air computation model. Furthermore, this method can also achieve a balance between gradient rank information loss and gradient information distortion while ensuring maximum aggregation delay, improving the performance of the first model after fine-tuning.
[0219] This application also proposes a federated low-rank adaptive fine-tuning system. The fine-tuning system includes a network device and multiple terminal devices. Any one of the terminal devices executes the method described above for terminal device execution, and the network device executes the method described above for network device execution.
[0220] For ease of understanding, the following example uses a large-scale federated low-rank adaptive fine-tuning system based on aerial computation, illustrated in Figure 5. It should be noted that the examples in Figures 2 to 5 are merely to assist those skilled in the art in understanding the embodiments of this application, and are not intended to limit the embodiments to the specific numerical values or scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or variations based on the examples in Figures 2 to 5, and such modifications or variations also fall within the scope of the embodiments of this application.
[0221] As shown in Figure 5, this low-rank adaptive fine-tuning system consists of one base station and K terminal devices. The K terminal devices are terminal device 1, ..., terminal device k, ..., terminal device K. The fine-tuning process shown in Figure 5 can be any of several fine-tuning cycles.
[0222] Referring to Figure 5, in step S51, the base station broadcasts a global adapter, which is received by K terminal devices. The global adapter is the previously described global low-rank adapter.
[0223] In step S52, the K terminal devices perform local fine-tuning based on the received global adapter. During the local fine-tuning process, the K terminal devices determine their local low-rank adapters ΔW1, ..., ΔW respectively based on the global low-rank adapter. k ..., ΔW K .
[0224] In step S53, K terminal devices report their local rank importance.
[0225] In step S54, the base station performs a federal rank importance analysis to determine the rank selection scheme.
[0226] In step S55, K terminal devices report the aggregation gradient according to the rank selection scheme, and the base station performs federated gradient aggregation.
[0227] In step S56, the base station performs a global adapter update, namely a global low-rank adapter ΔW update.
[0228] The method embodiments of this application have been described in detail above with reference to Figures 1 to 5. The apparatus embodiments of this application will now be described in detail below with reference to Figures 6 to 11. It should be understood that the descriptions of the apparatus embodiments correspond to the descriptions of the method embodiments; therefore, any parts not described in detail can be referred to the preceding method embodiments.
[0229] Figure 6 is a schematic block diagram of a terminal device according to an embodiment of this application. The terminal device 600 can be any terminal device applied to a machine learning model. The terminal device 600 shown in Figure 6 includes a transceiver unit 610 and a determination unit 620.
[0230] The transceiver unit 610 can be used to receive global model parameters related to the first model sent by the network device; the transceiver unit 610 is also used to send a first parameter to the network device, the first parameter being determined based on the global model parameters; the transceiver unit 610 is also used to receive multiple candidate parameters sent by the network device.
[0231] The processing unit 620 can be used to determine the local model aggregation parameters to be sent to the network device based on multiple candidate parameters; wherein, the terminal device is one of multiple terminal devices, the multiple candidate parameters are determined based on multiple first parameters sent by multiple terminal devices, and the multiple local model aggregation parameters of the multiple terminal devices are used to train the first model.
[0232] Optionally, the local model aggregation parameters include at least one column vector from the local model parameters, and multiple candidate parameters are used by the terminal device to determine at least one column vector.
[0233] Optionally, the processing unit 620 is also used to determine a first dataset, which belongs to the training subset of the terminal device; and to determine a first parameter and a local model parameter based on the first dataset and the global model parameters.
[0234] Optionally, the local model parameters are local low-rank adapter gradients, which include the gradients of the first matrix and the second matrix. The gradients of the first matrix and the second matrix are sample gradients of the correlation matrix of the global model parameters.
[0235] Optionally, multiple local model aggregation parameters are aggregated and transmitted based on over-the-air computation.
[0236] Optionally, the multiple signals corresponding to the aggregation are determined by normalizing the aggregation parameters of multiple local models.
[0237] Optionally, multiple local model aggregation parameters are used to determine global model aggregation parameters, the global model aggregation parameters are used by the network device to update global model parameters, and the update of global model parameters is used to train the first model.
[0238] Optionally, the global model parameters are global low-rank adapters, the global model aggregation parameters are global low-rank adapter aggregated gradients, and the global low-rank adapter aggregated gradients are:
[0239] Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and B t The two matrices represent the global low-rank adapter. and Represent matrix A respectively t and B t The aggregation gradient.
[0240] Optionally, multiple first parameters are also used to determine a resource allocation scheme for multiple local model aggregation parameters, and multiple candidate parameters and resource allocation schemes are sent via broadcast.
[0241] Optionally, multiple candidate parameters and resource allocation schemes are related to the first precoding matrix, which includes one or more of the following: the digital receive precoding matrix of the network device, the analog receive precoding matrix of the network device, and the digital transmit precoding matrix of the terminal device.
[0242] Optionally, the resource allocation scheme includes a sub-channel allocation scheme, and the association between multiple candidate parameters and the sub-channel allocation scheme and the first precoding matrix includes one or more of the following: multiple candidate parameters and the sub-channel allocation scheme are determined according to a given first precoding matrix; the first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and multiple candidate parameters; when multiple candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix; when multiple candidate parameters and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; when multiple candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix.
[0243] Optionally, multiple candidate parameters, resource allocation schemes, and the first precoding matrix are also determined jointly based on the lack of gradient rank information and the distortion of gradient information in the first model.
[0244] Optionally, the transceiver unit 610 is further configured to send a first parameter and a pilot signal to the network device; wherein the pilot signal is used to determine the channel coefficient.
[0245] Optionally, the first model is GPT or a large language model.
[0246] Figure 7 is a schematic diagram of a control device for the terminal device shown in Figure 6. Figure 7 uses a large-model federated low-rank adaptive fine-tuning system based on over-the-air computation as an example. The control device 700 is the control device for any terminal device in this fine-tuning system. As shown in Figure 7, the control device 700 of the terminal device may include a large-model low-rank adaptive fine-tuning terminal device computation module 710, a local low-rank adapter gradient aggregation construction module 720, and a large-model low-rank adaptive fine-tuning terminal device transmission module 730.
[0247] The large model low-rank adaptive fine-tuning terminal device computing module 710 can be used to control the terminal device to calculate the local low-rank adapter gradient and the local rank importance metric parameter using the received global low-rank adapter.
[0248] The local low-rank adapter aggregation gradient construction module 720 can be used to control the terminal device to construct the local low-rank adapter aggregation gradient using the received rank selection scheme and the locally calculated low-rank adapter gradient.
[0249] The large-model low-rank adaptive fine-tuning terminal device transmission module 730 can be used to control the local low-rank adapter constructed by the terminal device transmission to aggregate gradients in the same time-frequency resources and aggregate them on the base station side.
[0250] Figure 8 is a schematic block diagram of a network device according to an embodiment of this application. The network device 800 can be any of the network devices described above for model fine-tuning. The network device 800 shown in Figure 8 includes a transceiver unit 810.
[0251] The transceiver unit 810 can be used to send global model parameters related to the first model to multiple terminal devices; the transceiver unit 810 is also used to receive multiple first parameters sent by multiple terminal devices, the multiple first parameters being determined based on the global model parameters; the transceiver unit 810 is also used to send multiple candidate parameters to multiple terminal devices; wherein, the multiple candidate parameters are used by multiple terminal devices to determine multiple local model aggregation parameters to be sent to the network device; the multiple candidate parameters are determined based on the multiple first parameters, and the multiple local model aggregation parameters are used to train the first model.
[0252] Optionally, any one of the multiple local model aggregation parameters includes at least one column vector from the corresponding local model parameters, and the multiple candidate parameters are used by any one of the multiple terminal devices to determine at least one column vector.
[0253] Optionally, global model parameters and multiple first datasets from multiple terminal devices are used to determine multiple first parameters and multiple local model parameter gradients, with the multiple first datasets belonging to multiple training subsets from multiple terminal devices respectively.
[0254] Optionally, the multiple local model parameters are multiple local low-rank adapter gradients, and any one of the multiple local low-rank adapter gradients includes a first matrix gradient and a second matrix gradient, wherein the first matrix gradient and the second matrix gradient are sample gradients of the global model parameter correlation matrix.
[0255] Optionally, multiple local model aggregation parameters are aggregated and transmitted based on over-the-air computation.
[0256] Optionally, the aggregation of multiple signals is determined by normalizing the aggregation parameters of multiple local models.
[0257] Optionally, multiple local model aggregation parameters are used to determine global model aggregation parameters, which are then used by the network device to update model parameters. The update of the global model parameters is then used to train the first model.
[0258] Optionally, the global model parameters are global low-rank adapters, the global model aggregation parameters are global low-rank adapter aggregated gradients, and the global low-rank adapter aggregated gradients are:
[0259] Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and Bt The two matrices represent the global low-rank adapter. and Represent matrix A respectively t and B t The aggregation gradient.
[0260] Optionally, multiple first parameters are also used to determine a resource allocation scheme for multiple local model aggregation parameters, and multiple candidate parameters and resource allocation schemes are sent via broadcast.
[0261] Optionally, multiple candidate parameters and resource allocation schemes are related to the first precoding matrix, which includes one or more of the following: the digital receive precoding matrix of the network device, the analog receive precoding matrix of the network device, and the digital transmit precoding matrix of the terminal device.
[0262] Optionally, the resource allocation scheme includes a sub-channel allocation scheme, and the association between multiple candidate parameters and the sub-channel allocation scheme and the first precoding matrix includes one or more of the following: multiple candidate parameters and the sub-channel allocation scheme are determined according to a given first precoding matrix; the first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and multiple candidate parameters; when multiple candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix; when multiple candidate parameters and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; when multiple candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix.
[0263] Optionally, multiple candidate parameters, resource allocation schemes, and the first precoding matrix are also determined jointly based on the lack of gradient rank information and the distortion of gradient information in the first model.
[0264] Optionally, the transceiver unit 810 is further configured to receive multiple first parameters and multiple pilot signals sent by multiple terminal devices; wherein the multiple pilot signals are used to determine the channel coefficients.
[0265] Optionally, the first model is GPT or a large language model.
[0266] Figure 9 is a schematic diagram of a control device for the network device shown in Figure 8. Figure 9 uses a large-model federated low-rank adaptive fine-tuning system based on over-the-air computing as an example. Control device 900 is the control device for the base station in this fine-tuning system. As shown in Figure 9, control device 900 may include a global low-rank adapter broadcast module 910, a rank selection and resource allocation optimization module 920, a low-rank adapter gradient signal aggregation module 930, and a global low-rank adapter update module 940.
[0267] The Global Low-Rank Adapter Broadcast Module 910 can be used to broadcast the global low-rank adapter to various terminal devices.
[0268] The rank selection and resource allocation optimization module 920 can receive local rank importance measurement parameters transmitted from each terminal device and obtain the channel coefficients of each terminal device based on the received pilot signals. Furthermore, this module generates a rank selection and resource allocation strategy according to the aforementioned method.
[0269] The low-rank adapter gradient signal aggregation module 930 can be used to receive local low-rank adapter aggregated gradients transmitted by various terminal devices and aggregate them in the air to generate a global low-rank adapter aggregated gradient.
[0270] The global low-rank adapter update module 940 can be used to update the global low-rank adapter by aggregating gradients received from the global low-rank adapter.
[0271] Optionally, the global low-rank adapter can be obtained by aggregating gradient updates from the global low-rank adapter.
[0272] Figure 10 is a schematic diagram of an electronic device provided in an embodiment of this application. This electronic device is used to implement any step in the model training method described above. The following description uses a large model federated low-rank adaptive fine-tuning process based on over-the-air computation as an example. As shown in Figure 10, the electronic device structure includes a processor 1010, a memory 1020, a communication interface 1030, and a communication bus 1040.
[0273] The processor 1010 can be used to execute the program stored in the memory 1020 to implement any step in the large model federated low-rank adaptive fine-tuning process based on over-the-air computing provided in the embodiments of this application.
[0274] In this embodiment, the processor 1010 can be a general-purpose processor, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be other general-purpose processors or special-purpose processors, such as digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0275] The memory 1020 can be used for large-scale federated low-rank adaptive fine-tuning related procedures based on in-flight computing.
[0276] In this embodiment, the memory 1020 may be random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor. This embodiment does not impose specific limitations in this regard.
[0277] The communication bus 1040 can be used to complete communication between the processor 1010, the memory 1020 and the communication interface 1030.
[0278] In this embodiment, the communication bus 1040 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not mean that there is only one bus or one type of bus. This embodiment does not impose specific limitations in this regard.
[0279] The communication interface 1030 can be used for communication between the electronic device 1000 and other devices. These other devices include, but are not limited to, maintenance personnel and management equipment for large-scale federated low-rank adaptive fine-tuning models based on over-the-air computing, etc. This application embodiment does not specifically limit the scope of these devices.
[0280] In this embodiment, the communication interface 1030 can be an interface circuit for direct digital communication between a computer system and other systems. It typically includes a serial communication interface and a parallel communication interface. The serial communication interface is, for example, an asynchronous transmission standard interface (EIA-RS-232, RS232) or a Universal Serial Bus (USB). The parallel communication interface is, for example, a peripheral component interconnect express (PCI Express). This embodiment does not impose specific limitations on this aspect.
[0281] Figure 11 is a schematic diagram of the structure of a communication device according to an embodiment of this application. The dashed lines in Figure 11 indicate that the unit or module is optional. This device 1100 can be used to implement the methods described in the above method embodiments. Device 1100 can be a chip, a terminal device, or a network device.
[0282] Apparatus 1100 may include one or more processors 1110. The processor 1110 may support apparatus 1100 in implementing the methods described in the preceding method embodiments. Similar to processor 1010, processor 1110 may also be a general-purpose processor or a special-purpose processor, which will not be described further. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0283] The apparatus 1100 may further include one or more memories 1120. The memories 1120 store a program that can be executed by the processor 1110, causing the processor 1110 to perform the methods described in the preceding method embodiments. The memories 1120 may be independent of the processor 1110 or integrated within the processor 1110.
[0284] The device 1100 may also include a transceiver 1130. The processor 1110 can communicate with other devices or chips via the transceiver 1130. For example, the processor 1110 can send and receive data with other devices or chips via the transceiver 1130.
[0285] This application also provides a computer-readable storage medium for storing a program. This computer-readable storage medium can be applied to a terminal device or network device provided in this application embodiment, and the program causes a computer to execute the methods performed by the terminal device or network device in the various embodiments of this application.
[0286] The aforementioned computer-readable storage media can be any usable medium that a computer can read, or a data storage device such as a server or data center that integrates one or more usable media. The usable medium can be magnetic, optical, or semiconductor media, etc. Examples of computer storage media include, but are not limited to: phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of RAM, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc-read-only memory (CD-ROM), solid-state disk (SSD), digital video disc (DVD) or other optical storage, magnetic tape, magnetic tape / disk storage or other magnetic storage devices, or any other non-transfer medium. As defined herein, computer-readable media does not include transient media, such as modulated data signals and carrier waves.
[0287] Computer-readable media include both permanent and non-permanent, removable and non-removable media, and can be used to store information by any method or technology. Computer-readable media can be used to store information that can be accessed by computing devices. The information can be computer-readable instructions, data structures, program modules, or other data.
[0288] This application also provides a readable storage medium storing a program or instructions that, when executed by a processor, can implement various processes of the above method embodiments or achieve the same technical effects. To avoid repetition, these will not be described again here. Optionally, when the program or instructions are executed by a processor, they can implement any step in any of the above embodiments, such as any step in the large-model federated low-rank adaptive fine-tuning process based on over-the-air computing.
[0289] This application also provides a computer program product. The computer program product includes a program. The computer program product can be applied to a terminal device or network device provided in this application embodiment, and the program causes a computer to perform the methods executed by the terminal or network device in the various embodiments of this application. Optionally, the computer program product includes instructions that, when run on a computer, cause the computer to perform any step in any of the above embodiments, such as any step in the large-model federated low-rank adaptive fine-tuning process based on over-the-air computation.
[0290] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments, electronic device embodiments, computer-readable storage medium embodiments, and computer program product embodiments are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0291] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0292] This application also provides a computer program. This computer program can be applied to a terminal device or network device provided in this application, and the computer program causes the computer to execute the methods performed by the terminal or network device in various embodiments of this application.
[0293] In this application, the terms "system" and "network" are used interchangeably. Furthermore, the terminology used in this application is for illustrative purposes only and is not intended to limit the scope of the application.
[0294] Relational terms such as “first,” “second,” and “third” in this application are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0295] It should be understood that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0296] In the embodiments of this application, the term "instruction" can be a direct instruction, an indirect instruction, or an indication of a relationship. For example, A instructing B can mean that A directly instructs B, such as B being able to obtain information through A; it can also mean that A indirectly instructs B, such as A instructing C, so B can obtain information through C; or it can mean that there is a relationship between A and B.
[0297] In the embodiments of this application, the term "correspondence" may indicate a direct or indirect correspondence between two things, or an association between two things, or a relationship such as instruction and being instructed, configuration and being configured.
[0298] In the embodiments of this application, the term "protocol" may refer to standard protocols in the field of communications, such as LTE protocols, NR protocols, and related protocols applied in future communication systems. This application does not limit the scope of these protocols.
[0299] In the embodiments of this application, determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information.
[0300] In the embodiments of this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0301] In the embodiments of this application, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0302] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0303] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0304] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0305] Through the above description of the embodiments, those skilled in the art can clearly understand that the above method embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a service classification device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0306] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any modifications, equivalent substitutions, or improvements that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A communication method, characterized in that, The communication method is applied to a machine learning model, and the communication method includes: The terminal device receives global model parameters related to the first model sent by the network device; The terminal device sends a first parameter to the network device, the first parameter being determined based on the global model parameter; The terminal device receives multiple candidate parameters sent by the network device; The terminal device determines the local model aggregation parameters to be sent to the network device based on the multiple candidate parameters; The terminal device is one of a plurality of terminal devices, the plurality of candidate parameters are determined based on a plurality of first parameters sent by the plurality of terminal devices, and the plurality of local model aggregation parameters of the plurality of terminal devices are used to train the first model.
2. The communication method according to claim 1, characterized in that, The local model aggregation parameters include at least one column vector among the local model parameters, and the plurality of candidate parameters are used by the terminal device to determine the at least one column vector.
3. The communication method according to claim 2, characterized in that, After the terminal device receives the global model parameters related to the first model sent by the network device, the communication method further includes: The terminal device determines a first dataset, which belongs to the training subset of the terminal device. The terminal device determines the first parameter and the local model parameter based on the first dataset and the global model parameter.
4. The communication method according to claim 3, characterized in that, The local model parameters are local low-rank adapter gradients, which include a first matrix gradient and a second matrix gradient. The first matrix gradient and the second matrix gradient are sample gradients of the global model parameter correlation matrix.
5. The communication method according to any one of claims 1-4, characterized in that, The multiple local model aggregation parameters are aggregated and transmitted based on over-the-air computation.
6. The communication method according to claim 5, characterized in that, The multiple signals corresponding to the aggregation are determined by normalizing the aggregation parameters of the multiple local models.
7. The communication method according to any one of claims 1-6, characterized in that, The multiple local model aggregation parameters are used to determine the global model aggregation parameters, which are then used by the network device to update the global model parameters. The update of the global model parameters is then used to train the first model.
8. The communication method according to claim 7, characterized in that, The global model parameters are global low-rank adapter parameters, and the global model aggregation parameters are global low-rank adapter aggregated gradients. The global low-rank adapter aggregated gradients are: Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and B t The two matrices represent the global low-rank adapter. and Represent matrix A respectively t and B t The aggregation gradient.
9. The communication method according to any one of claims 1-8, characterized in that, The plurality of first parameters are also used to determine the resource allocation scheme of the plurality of local model aggregation parameters, and the plurality of candidate parameters and the resource allocation scheme are sent by broadcast.
10. The communication method according to claim 9, characterized in that, The plurality of candidate parameters and the resource allocation scheme are related to the first precoding matrix, which includes one or more of the following: the digital receive precoding matrix of the network device, the analog receive precoding matrix of the network device, and the digital transmit precoding matrix of the terminal device.
11. The communication method according to claim 10, characterized in that, The resource allocation scheme includes a sub-channel allocation scheme, and the association between the plurality of candidate parameters and the sub-channel allocation scheme and the first precoding matrix includes one or more of the following: The plurality of candidate parameters and the sub-channel allocation scheme are determined according to a given first precoding matrix; The first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and the multiple candidate parameters; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix. When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix.
12. The communication method according to claim 10 or 11, characterized in that, The multiple candidate parameters, the resource allocation scheme, and the first precoding matrix are also determined based on the missing gradient rank information and the distortion of gradient information in the first model.
13. The communication method according to any one of claims 1-12, characterized in that, The terminal device sends a first parameter to the network device, including: The terminal device sends the first parameter and pilot signal to the network device; The pilot signal is used to determine the channel coefficients.
14. The communication method according to any one of claims 1-13, characterized in that, The first model is either a generative pre-trained converter (GPT) or a large language model.
15. A communication method, characterized in that, The communication method is applied to a machine learning model, and the communication method includes: The network device sends global model parameters related to the first model to multiple terminal devices; The network device receives multiple first parameters sent by the multiple terminal devices, and the multiple first parameters are determined according to the global model parameters; The network device sends multiple candidate parameters to the multiple terminal devices; The plurality of candidate parameters are used by the plurality of terminal devices to determine a plurality of local model aggregation parameters to be sent to the network device; the plurality of candidate parameters are determined based on the plurality of first parameters, and the plurality of local model aggregation parameters are used to train the first model.
16. The communication method according to claim 15, characterized in that, Each of the plurality of local model aggregation parameters includes at least one column vector in the corresponding local model parameters, and the plurality of candidate parameters are used by any of the plurality of terminal devices to determine the at least one column vector.
17. The communication method according to claim 16, characterized in that, The global model parameters and the multiple first datasets of the multiple terminal devices are used to determine the multiple first parameters and the multiple local model parameters, wherein the multiple first datasets belong to multiple training subsets of the multiple terminal devices.
18. The communication method according to claim 17, characterized in that, The plurality of local model parameters are plurality of local low-rank adapter gradients. Each local low-rank adapter gradient includes a first matrix gradient and a second matrix gradient. The first matrix gradient and the second matrix gradient are sample gradients of the global model parameter correlation matrix.
19. The communication method according to any one of claims 15-18, characterized in that, The multiple local model aggregation parameters are aggregated and transmitted based on over-the-air computation.
20. The communication method according to claim 19, characterized in that, The multiple signals corresponding to the aggregation are determined by normalizing the aggregation parameters of the multiple local models.
21. The communication method according to any one of claims 15-20, characterized in that, The multiple local model aggregation parameters are used to determine the global model aggregation parameters, which are then used by the network device to update the global model parameters. The update of the global model parameters is then used to train the first model.
22. The communication method according to claim 21, characterized in that, The global model parameters are global low-rank adapter parameters, and the global model aggregation parameters are global low-rank adapter aggregated gradients. The global low-rank adapter aggregated gradients are: Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and B t The two matrices represent the global low-rank adapter. and Represent matrix A respectively t and B t The aggregation gradient.
23. The communication method according to any one of claims 15-22, characterized in that, The plurality of first parameters are also used to determine the resource allocation scheme of the plurality of local model aggregation parameters, and the plurality of candidate parameters and the resource allocation scheme are sent by broadcast.
24. The communication method according to claim 23, characterized in that, The plurality of candidate parameters and the resource allocation scheme are related to the first precoding matrix, which includes one or more of the following: the digital receive precoding matrix of the network device, the analog receive precoding matrix of the network device, and the digital transmit precoding matrix of the terminal device.
25. The communication method according to claim 24, characterized in that, The resource allocation scheme includes a sub-channel allocation scheme, and the association between the plurality of candidate parameters and the sub-channel allocation scheme and the first precoding matrix includes one or more of the following: The plurality of candidate parameters and the sub-channel allocation scheme are determined according to a given first precoding matrix; The first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and the multiple candidate parameters; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix. When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix.
26. The communication method according to claim 24 or 25, characterized in that, The multiple candidate parameters, the resource allocation scheme, and the first precoding matrix are also determined based on the missing gradient rank information and the distortion of gradient information in the first model.
27. The communication method according to any one of claims 15-26, characterized in that, The network device receives multiple first parameters sent by the multiple terminal devices, including: The network device receives the multiple first parameters and multiple pilot signals sent by the multiple terminal devices; The plurality of pilot signals are used to determine the channel coefficients.
28. The communication method according to any one of claims 15-27, characterized in that, The first model is either a generative pre-trained converter (GPT) or a large language model.
29. A terminal device, characterized in that, The terminal device is used in a machine learning model, and the terminal device includes: The transceiver unit is used to receive global model parameters related to the first model sent by the network device; The transceiver unit is also used to send a first parameter to the network device, the first parameter being determined based on the global model parameter; The transceiver unit is also used to receive multiple candidate parameters sent by the network device; The processing unit is configured to determine the local model aggregation parameters to be sent to the network device based on the plurality of candidate parameters; The terminal device is one of a plurality of terminal devices, the plurality of candidate parameters are determined based on a plurality of first parameters sent by the plurality of terminal devices, and the plurality of local model aggregation parameters of the plurality of terminal devices are used to train the first model.
30. The terminal device according to claim 29, characterized in that, The local model aggregation parameters include at least one column vector among the local model parameters, and the plurality of candidate parameters are used by the terminal device to determine the at least one column vector.
31. The terminal device according to claim 30, characterized in that, After the terminal device receives the global model parameters related to the first model sent by the network device, the processing unit is further configured to: A first dataset is determined, which belongs to the training subset of the terminal device; The first parameter and the local model parameter are determined based on the first dataset and the global model parameters.
32. The terminal device according to claim 31, characterized in that, The local model parameters are local low-rank adapter gradients, which include a first matrix gradient and a second matrix gradient. The first matrix gradient and the second matrix gradient are sample gradients of the global model parameter correlation matrix.
33. The terminal device according to any one of claims 29-32, characterized in that, The multiple local model aggregation parameters are aggregated and transmitted based on over-the-air computation.
34. The terminal device according to claim 33, characterized in that, The multiple signals corresponding to the aggregation are determined by normalizing the aggregation parameters of the multiple local models.
35. The terminal device according to any one of claims 29-34, characterized in that, The multiple local model aggregation parameters are used to determine the global model aggregation parameters, which are then used by the network device to update the global model parameters. The update of the global model parameters is then used to train the first model.
36. The terminal device according to claim 35, characterized in that, The global model parameters are global low-rank adapter parameters, and the global model aggregation parameters are global low-rank adapter aggregated gradients. The global low-rank adapter aggregated gradients are: Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and B t The two matrices represent the global low-rank adapter. and Represent matrix A respectively t and B t The aggregation gradient.
37. The terminal device according to any one of claims 29-36, characterized in that, The plurality of first parameters are also used to determine the resource allocation scheme of the plurality of local model aggregation parameters, and the plurality of candidate parameters and the resource allocation scheme are sent by broadcast.
38. The terminal device according to claim 37, characterized in that, The plurality of candidate parameters and the resource allocation scheme are related to the first precoding matrix, which includes one or more of the following: the digital receive precoding matrix of the network device, the analog receive precoding matrix of the network device, and the digital transmit precoding matrix of the terminal device.
39. The terminal device according to claim 38, characterized in that, The resource allocation scheme includes a sub-channel allocation scheme, and the association between the plurality of candidate parameters and the sub-channel allocation scheme and the first precoding matrix includes one or more of the following: The plurality of candidate parameters and the sub-channel allocation scheme are determined according to a given first precoding matrix; The first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and the multiple candidate parameters; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix. When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix.
40. The terminal device according to claim 38 or 39, characterized in that, The multiple candidate parameters, the resource allocation scheme, and the first precoding matrix are also determined based on the missing gradient rank information and the distortion of gradient information in the first model.
41. The terminal device according to any one of claims 29-40, characterized in that, The transceiver unit is also used for: Send the first parameter and pilot signal to the network device; The pilot signal is used to determine the channel coefficients.
42. The terminal device according to any one of claims 29-41, characterized in that, The first model is either a generative pre-trained converter (GPT) or a large language model.
43. A network device, characterized in that, The network device is used in a machine learning model, and the network device includes: The transceiver unit is used to send global model parameters related to the first model to multiple terminal devices. The transceiver unit is also used to receive multiple first parameters sent by the multiple terminal devices, the multiple first parameters being determined based on the global model parameters; The transceiver unit is also used to send multiple candidate parameters to the multiple terminal devices; The plurality of candidate parameters are used by the plurality of terminal devices to determine a plurality of local model aggregation parameters to be sent to the network device; the plurality of candidate parameters are determined based on the plurality of first parameters, and the plurality of local model aggregation parameters are used to train the first model.
44. The network device according to claim 43, characterized in that, Each of the plurality of local model aggregation parameters includes at least one column vector in the corresponding local model parameters, and the plurality of candidate parameters are used by any of the plurality of terminal devices to determine the at least one column vector.
45. The network device according to claim 44, characterized in that, The global model parameters and the multiple first datasets of the multiple terminal devices are used to determine the multiple first parameters and the multiple local model parameters, wherein the multiple first datasets belong to multiple training subsets of the multiple terminal devices.
46. The network device according to claim 45, characterized in that, The plurality of local model parameters are plurality of local low-rank adapter gradients. Each local low-rank adapter gradient includes a first matrix gradient and a second matrix gradient. The first matrix gradient and the second matrix gradient are sample gradients of the global model parameter correlation matrix.
47. The network device according to any one of claims 43-46, characterized in that, The multiple local model aggregation parameters are aggregated and transmitted based on over-the-air computation.
48. The network device according to claim 47, characterized in that, The multiple signals corresponding to the aggregation are determined by normalizing the aggregation parameters of the multiple local models.
49. The network device according to any one of claims 43-48, characterized in that, The multiple local model aggregation parameters are used to determine the global model aggregation parameters, which are then used by the network device to update the global model parameters. The update of the global model parameters is then used to train the first model.
50. The network device according to claim 49, characterized in that, The global model parameters are global low-rank adapter parameters, and the global model aggregation parameters are global low-rank adapter aggregated gradients. The global low-rank adapter aggregated gradients are: Where t represents the current period, η represents the learning rate of the global low-rank adapter, and A t and B t The two matrices represent the global low-rank adapter. and Represent matrix A respectively t and B t The aggregation gradient.
51. The network device according to any one of claims 43-50, characterized in that, The plurality of first parameters are also used to determine the resource allocation scheme of the plurality of local model aggregation parameters, and the plurality of candidate parameters and the resource allocation scheme are sent by broadcast.
52. The network device according to claim 51, characterized in that, The plurality of candidate parameters and the resource allocation scheme are related to the first precoding matrix, which includes one or more of the following: the digital receive precoding matrix of the network device, the analog receive precoding matrix of the network device, and the digital transmit precoding matrix of the terminal device.
53. The network device according to claim 52, characterized in that, The resource allocation scheme includes a sub-channel allocation scheme, and the association between the plurality of candidate parameters and the sub-channel allocation scheme and the first precoding matrix includes one or more of the following: The plurality of candidate parameters and the sub-channel allocation scheme are determined according to a given first precoding matrix; The first precoding matrix is determined through multiple optimizations of the sub-channel allocation scheme and the multiple candidate parameters; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the digital receive precoding matrix. When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the digital receive precoding matrix and the digital transmit precoding matrix in the first precoding matrix are used to optimize the analog receive precoding matrix; When the plurality of candidate parameters and the sub-channel allocation scheme are determined, the analog receive precoding matrix and the digital receive precoding matrix in the first precoding matrix are used to optimize the digital transmit precoding matrix.
54. The network device according to claim 52 or 53, characterized in that, The multiple candidate parameters, the resource allocation scheme, and the first precoding matrix are also determined based on the missing gradient rank information and the distortion of gradient information in the first model.
55. The network device according to any one of claims 43-54, characterized in that, The transceiver unit is also used for: Receive the plurality of first parameters and the plurality of pilot signals sent by the plurality of terminal devices; The plurality of pilot signals are used to determine the channel coefficients.
56. The network device according to any one of claims 43-55, characterized in that, The first model is either a generative pre-trained converter (GPT) or a large language model.
57. A communication device, characterized in that, It includes a memory and a processor, the memory being used to store a program, and the processor being used to invoke the program in the memory to execute the communication method as described in any one of claims 1-28.
58. An apparatus, characterized in that, Includes a processor for calling a program from memory to perform the communication method as described in any one of claims 1-28.
59. A chip, characterized in that, Includes a processor for calling a program from memory, causing a device on which the chip is mounted to perform the communication method as described in any one of claims 1-28.
60. A computer-readable storage medium, characterized in that, It stores a program that causes a computer to perform the communication method as described in any one of claims 1-28.
61. A computer program product, characterized in that, Includes a program that causes a computer to perform the communication method as described in any one of claims 1-28.
62. A computer program, characterized in that, The computer program causes the computer to perform the communication method as described in any one of claims 1-28.