Network quantization method and device and related equipment
Patent Information
- Application Number
- CN202380090649.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-19
- Publication Date
- 2025-09-12
AI Technical Summary
The learning performance of existing collaborative learning algorithms is significantly reduced in dynamically changing wireless network environments, and the transmission bandwidth consumption is large, making it difficult to meet the accuracy requirements of local and global models.
By transmitting the quantization model and KL divergence between terminal equipment and access network equipment, the local quantization model is updated based on global aggregation information, and the quantization level is optimized based on wireless network resources and KL divergence to achieve dynamic updating and aggregation, reducing transmission bandwidth consumption. .
It maintains good learning performance in dynamically changing wireless networks, meets local model accuracy requirements, significantly reduces transmission bandwidth consumption, and improves training efficiency.
Smart Images

Figure CN120642307A_ABST
Abstract
Description
A network quantization method, device and related equipment Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a network quantization method, apparatus, and related equipment. Background Art
[0002] In recent years, to promote the development of intelligent wireless networks in the upcoming B5G / 6G era, scholars in academia and industry have conducted extensive research on developing collaborative learning algorithms for efficient communication and based on heterogeneous models. Quantization is a potential approach. For example, dual quantization schemes can be used to ensure convergence speed while simultaneously reducing communication costs by quantizing model parameters and gradients. Alternatively, lossy collaborative learning algorithms can be introduced by quantizing global and local models to reduce communication costs. Wireless network environments are constantly changing, for example, due to factors such as terminal mobility, interference levels, and changes in service load and idle times. However, current collaborative learning algorithms are designed based on relatively stable network environments. Direct application to dynamically changing wireless network environments can significantly degrade learning performance.
[0003] Summary of the Invention
[0004] The present application provides a network quantization method, apparatus, and related equipment. The method is applicable to dynamically changing wireless networks, which can not only meet the local / global model accuracy requirements, but also reduce the consumption of transmission bandwidth.
[0005] In a first aspect, the present application provides a network quantization method, which is executed by a first device, or may be executed by a component of the first device (such as a processor, a chip, or a chip system, etc.), or may be implemented by a logic module or software that can realize all or part of the functions of the first device. For example, the first device is a terminal device (such as a user in collaborative learning). The first device receives the first model and the first KL divergence of the t-th aggregation processing from the second device; wherein the first model of the t-th aggregation processing is obtained by the second device performing aggregation processing on multiple second models of the t-th updated quantization from multiple first devices, and the first KL divergence of the t-th aggregation processing represents the average difference between the multiple second models of the t-th updated quantization and the first model of the t-th aggregation processing; t is a positive integer. The first device determines the second model of the t+1-th updated quantization based on the first model of the t-th aggregation processing and the first KL divergence, and sends the second model of the t+1-th updated quantization to the second device.
[0006] In this method, the first device can receive the global aggregate information / model (i.e., the first model and the first KL divergence) sent by the second device, and update the local quantization model based on the global aggregate information / model. Due to the integration of global knowledge, the first device can achieve faster convergence when updating the local quantization model (i.e., the second model), thereby improving training efficiency. In addition, the first device can continue to upload the second model to the second device, thereby performing a new round of updated quantization. Even if the use of quantization results in a certain loss of accuracy, this method still achieves good learning performance, can meet the local model accuracy requirements, and is conducive to reducing the consumption of transmission bandwidth.
[0007] In one possible implementation, the first device determines a third model trained for the t+1th time based on the first model from the tth aggregation process and the third model trained for the tth time; determines a quantization level for the third model trained for the t+1th time based on the first KL divergence from the tth aggregation process. The first device determines a second model for the t+1th quantization update based on the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time.
[0008] In this method, when updating a local quantization model based on global aggregated information / model, the first device can first train the local model based on the global model to obtain a new local model. This helps achieve better learning performance and meets local model accuracy requirements. Furthermore, the first device can update the quantization level of the local model and quantize the new local model based on the new quantization level, which helps reduce transmission bandwidth consumption.
[0009] In one possible implementation, the first device determines an enhanced loss function for the tth training based on one or more of the following information: a cross-entropy loss function corresponding to the third model of the tth training, a cross-entropy loss function corresponding to the second model of the tth updated quantization, wireless resources consumed by transmitting the tth updated quantization second model, an upper limit of wireless resources that can be allocated by the first device, and a second KL divergence for the tth training; wherein the second KL divergence for the tth training represents the difference between the third model of the tth training and the first model of the tth aggregation processing, and / or the difference between the second model obtained by the tth updated quantization and the first model of the tth aggregation processing. The first device trains the first model of the tth aggregation processing based on the enhanced loss function of the tth training to obtain a third model of the t+1th training.
[0010] In this method, when updating the local model, the first device needs to additionally consider parameters such as KL divergence and wireless network resources, so that the local model is applicable to dynamically changing wireless networks and meets local / global model accuracy requirements.
[0011] In one possible implementation, the first device updates the quantization level of the third model trained for the tth time based on one or more information including the first KL divergence of the tth aggregation processing, the wireless resources consumed in transmitting the second model obtained by the tth updated quantization, and the cross-entropy loss function of the second model obtained by the tth updated quantization, to obtain the quantization level of the third model trained for the t+1th time.
[0012] In this method, when the first device updates the quantization level of the local model, it also needs to consider parameters such as KL divergence and wireless network resources, so that the local quantization model is also applicable to dynamically changing wireless networks and meets local / global model accuracy requirements.
[0013] In a possible implementation, the first device quantizes the third model trained for the t+1th time using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
[0014] In this method, the first device can directly use the quantization level updated in this round to quantize the local model trained in this round, which is beneficial to reducing the consumption of transmission bandwidth.
[0015] In one possible implementation, the first device updates the third model trained at the t+1th time based on the third KL divergence of the tth updated quantization to obtain a fourth model trained at the t+1th time; wherein the third KL divergence of the tth updated quantization represents the difference between the third model trained at the t+1th time and the third model trained at the tth time, and / or the difference between the fourth model updated quantization at the t+1th time and the third model trained at the tth time; the fourth model updated quantization at the t+1th time is a model obtained by quantizing the third model trained at the t+1th time using the quantization level of the third model trained at the t+1th time. The first device quantizes the fourth model trained at the t+1th time using the quantization level of the third model trained at the t+1th time to determine the second model updated quantization at the t+1th time.
[0016] In this method, since the global aggregation model may aggregate heterogeneous models (for example, the local models of the first device under different base stations are heterogeneous), the first device also considers the differences between the heterogeneous models when updating the local model based on the global aggregation model and KL divergence, which is conducive to narrowing the gap between heterogeneous models, improving model accuracy, and reducing transmission bandwidth consumption.
[0017] In the second aspect, the present application provides another network quantization method, which is executed by a second device, and can also be executed by a component of the second device (such as a processor, a chip, or a chip system, etc.), and can also be implemented by a logic module or software that can realize all or part of the functions of the second device. For example, the second device is an access network device or a core network device (such as an aggregator in collaborative learning). Among them, the second device receives multiple second models of the t-th updated quantization from multiple first devices, and t is a positive integer. The second device aggregates the multiple second models of the t-th updated quantization to obtain the first model and the first KL divergence of the t-th aggregated processing, and the first KL divergence of the t-th aggregated processing represents the average difference between the multiple second models of the t-th updated quantization and the first model of the t-th aggregated processing. The second device sends the first model and the first KL divergence of the t-th updated quantization to multiple first devices respectively, and the first model and the first KL divergence of the t-th updated quantization are used by the first device to determine the second model of the t+1-th updated quantization.
[0018] In this method, the second device can receive the local quantization model (i.e., the second model) from the first device. Even if the quantization results in a certain loss of accuracy, this method still achieves good learning performance, can meet the local model accuracy requirements, and is conducive to reducing transmission bandwidth consumption. In addition, the second device aggregates multiple local quantization models to obtain and distribute global aggregate information / model (i.e., the third model and the first KL divergence). Due to the integration of global knowledge, it is beneficial for the first device to achieve faster convergence when updating the local quantization model, thereby improving training efficiency.
[0019] In one possible implementation, the second device restores the multiple second models updated and quantized for the tth time to obtain multiple sixth models restored for the tth time; aggregates the multiple sixth models restored for the tth time to obtain a first model aggregated for the tth time. The second device calculates a first KL divergence for the tth aggregation based on the first model aggregated for the tth time and the multiple second models updated and quantized for the tth time.
[0020] In this method, the second device can obtain global aggregation information / model and obtain KL divergence, thereby determining the difference between the local quantization model and the global aggregation model, which is conducive to improving model accuracy.
[0021] In one possible implementation, the second device receives multiple second models updated and quantized for the t+1th time from multiple first devices; the second device aggregates the multiple second models updated and quantized for the t+1th time to obtain the first model and the first KL divergence of the t+1th aggregation processing.
[0022] In this method, the second device can receive the local model updated based on the global aggregation model and KL divergence from the first device, and the second device can start a new round of global aggregation processing, which is conducive to improving model accuracy.
[0023] In a third aspect, the present application provides a network quantization device, which may be a terminal device, a device within a terminal device, or a device capable of being used in conjunction with a terminal device. In one possible implementation, the network quantization device may include a module corresponding to the method / operation / step / action described in the first aspect and any possible implementation of the first aspect. The module may be implemented as a hardware circuit, software, or a combination of hardware circuits and software. In one possible implementation, the network quantization device may include a processing unit and a communication unit.
[0024] It can be understood that the network quantization device can also achieve the effects that can be achieved in the first aspect and any possible implementation of the first aspect.
[0025] In a fourth aspect, the present application provides another network quantization device, which may be a network device, a device within a network device, or a device capable of being used in conjunction with a network device. In one possible implementation, the network quantization device may include a module corresponding to the method / operation / step / action described in the second aspect and any possible implementation of the second aspect. The module may be implemented as a hardware circuit, software, or a combination of hardware circuits and software. In one possible implementation, the network quantization device may include a processing unit and a communication unit.
[0026] It can be understood that the network quantization device can also achieve the effects that can be achieved in the second aspect and any possible implementation of the second aspect.
[0027] In a fifth aspect, the present application provides a terminal device, comprising: a processor and a memory, the memory being configured to store instructions, which, when executed by the processor, enable the terminal device to implement the method of the first aspect and any possible implementation of the first aspect. Optionally, the processor and the memory are coupled.
[0028] In a sixth aspect, the present application provides a network device, comprising: a processor and a memory, the memory being configured to store instructions, wherein when the instructions are executed by the processor, the network device implements the method of the second aspect and any possible implementation of the second aspect. Optionally, the processor and the memory are coupled.
[0029] In the seventh aspect, the present application provides a communication system, which includes multiple devices or equipment in the above-mentioned third to sixth aspects, so that the devices or equipment execute the methods in the first and second aspects, as well as any possible implementation of the first and second aspects.
[0030] In an eighth aspect, the present application provides a computer-readable storage medium storing instructions, which, when executed on a computer, enables the computer to execute the method of the first aspect and the second aspect, as well as any possible implementation of the first aspect and the second aspect.
[0031] In a ninth aspect, the present application provides a chip system comprising a processor and an interface, and may also include a memory, for implementing the method of the first and second aspects, as well as any possible implementation of the first and second aspects. The chip system may be composed of a chip, or may include a chip and other discrete devices.
[0032] In a tenth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer, enable the computer to execute the method of the first aspect and the second aspect, as well as any possible implementation of the first aspect and the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] FIG1 is a schematic diagram of a communication system provided by the present application;
[0034] FIG2 is a schematic diagram of a global iteration process provided by this application;
[0035] FIG3 is a flow chart of a network quantization method provided by the present application;
[0036] FIG4 is a schematic diagram showing a comparison of the accuracy of a model of the network quantization method provided by this application and other collaborative learning schemes;
[0037] FIG5 is a schematic diagram showing a comparison of the accuracy of another model of the network quantization method provided by this application and other collaborative learning schemes;
[0038] FIG6 is a schematic diagram showing a comparison of bandwidth resources of the network quantization method provided by this application and other collaborative learning solutions;
[0039] FIG7 is a schematic diagram showing a comparison of training time between the network quantization method provided by the present application and other collaborative learning schemes;
[0040] FIG8 is a schematic diagram showing another comparison of training time between the network quantization method provided by the present application and other collaborative learning schemes;
[0041] FIG9 is a schematic diagram of a device provided by the present application;
[0042] FIG10 is a schematic diagram of a device provided in this application. DETAILED DESCRIPTION
[0043] In the embodiments of this application, " / " can indicate that the associated objects are in an "or" relationship. For example, A / B can mean A or B. "And / or" can be used to describe the existence of three relationships between associated objects. For example, "A and / or B" can mean: A exists alone, A and B exists simultaneously, or B exists alone. A and B can be singular or plural. To facilitate the description of the technical solutions of the embodiments of this application, the words "first" and "second" may be used in the embodiments of this application to distinguish between technical features with the same or similar functions. The words "first" and "second" do not limit the number or order of execution, and the words "first" and "second" do not necessarily mean different. In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" should not be construed as preferred or advantageous over other embodiments or designs. The use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for easier understanding.
[0044] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0045] In order to solve the problem of how to achieve collaborative learning with high learning performance and reduce communication overhead in a dynamically changing wireless network environment, the present application provides a network quantization method, which is suitable for dynamically changing wireless networks and can not only meet the local / global model accuracy requirements but also reduce the consumption of transmission bandwidth.
[0046] The network quantization method provided in this application can be applied to the communication system shown in Figure 1. For example, the communication system includes a network device and a terminal device. In one possible implementation, the network device is an access network device such as a base station that provides user access functions. In another possible implementation, the network device is a device in the core network such as an access and management function (AMF), which communicates with the terminal device through the base station.
[0047] Among them, the communication system of the present application may include but is not limited to communication systems of various radio access technologies (RAT), for example, it may be an LTE communication system, it may also be a 5G (or new radio (NR)) communication system, it may also be a transition system between the LTE communication system and the 5G communication system, the transition system may also be called a 4.5G communication system, and of course it may also be a future communication system, such as the sixth generation (6G) or even the seventh generation (7G) system. The network architecture and service scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. It is known to those skilled in the art that with the evolution of the communication network architecture and the emergence of new service scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0048] Among them, terminal devices, also known as user equipment (UE), mobile station (MS), mobile terminal (MT), etc., refer to devices that provide voice and / or data connectivity to users. For example, handheld devices and in-vehicle devices with wireless connection capabilities. Currently, some examples of terminal devices include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, drones, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, terminal devices in 5G networks, terminal devices in future evolved PLMN networks, or terminal devices in future communication systems.
[0049] The access network device refers to a radio access network (RAN) node (or device) that connects a terminal device to a wireless network, and may also be referred to as a base station. For example, some examples of RAN nodes include: a gNB, a transmission reception point (TRP), an evolved Node B (eNB), a radio network controller (RNC), a Node B (NB), a base station controller (BSC), a base transceiver station (BTS), a home base station (e.g., a home evolved NodeB, or a home Node B, HNB), a base band unit (BBU), or a wireless fidelity (Wifi) access point (AP), a satellite in a satellite communication system, a radio controller in a cloud radio access network (CRAN) scenario, a wearable device, a drone, or a device in an Internet of Vehicles (e.g., vehicle to everything (V2X)), or a communication device in device to device (D2D) communication. In addition, in one network structure, the access network equipment may include a centralized unit (CU) node, a distributed unit (DU) node, or a RAN device including a CU node and a DU node. The RAN device including the CU node and the DU node splits the protocol layer of the eNB in the long-term evolution (LTE) system, centrally controlling some protocol layer functions in the CU, and distributing some or all of the remaining protocol layer functions in the DU, which is then centrally controlled by the CU.
[0050] The core network includes devices such as AMF and user plane function (UPF). The AMF is used for user access management, security authentication, and mobility management. The AMF communicates with base stations, enabling communication with terminal devices through the base stations. The UPF manages user plane data transmission and traffic statistics.
[0051] 1. To facilitate understanding, the definitions of relevant terms involved in this application are introduced in detail below.
[0052] 1. Collaborative learning:
[0053] Collaborative learning has been widely recognized as one of the most promising enablers of networked intelligence. It not only facilitates the collaborative training of machine learning models, but also protects user privacy and data security. For example, collaborative learning employs multiple global iterations to train a collaborative (global) model, with each agent exchanging information / models once during each global iteration. Traditional collaborative learning falls into two main categories: multi-agent learning and distributed learning.
[0054] In multi-agent learning, agents (i.e., smart devices) interact with their environment, and with each other, collaborating, coordinating, competing, or collectively learning to complete specific learning tasks. The content of these interactions can be configured based on training requirements (e.g., data and rewards). In distributed machine learning, agents actively interact and collaborate to build consensus by sharing lightweight experiences (e.g., local training parameters and gradients) and assuming asymmetric roles. Compared to traditional centralized machine learning methods, collaborative learning can maintain a globally shared learning model to solve large-scale machine learning problems while reducing network burden. Distributed machine learning also protects user privacy and data security.
[0055] Despite these attractive and valuable advantages, collaborative learning still faces numerous challenges in achieving network intelligence, particularly when deployed in wireless networks. For one thing, agents can consume significant transmission resources when transmitting interactive content. For example, in distributed machine learning, even when transmitting lightweight models (e.g., parameters and gradients) rather than raw data, the wireless resource cost of this transmission can reach billions of bytes when the number of participating user devices or the size of the local model is sufficiently large. Furthermore, the heterogeneity of wireless systems (heterogeneous computing and communication resources) can degrade collaborative learning performance (training time and model accuracy). For example, if the wireless link over which an agent transmits its model is unstable or has low computational power, the agent's transmission rate and / or computational speed will be limited. Furthermore, due to concerns about agent privacy and data security, many companies and institutions are developing their own models to address different tasks, hoping to explore collaboration without sharing learning model details or local data.
[0056] For example, federated learning, a typical collaborative learning framework, is widely considered one of the most promising collaborative learning methods for supporting networked collaborative intelligence. Federated learning can be used to address the issue of users updating models locally. Its design goal is to enable efficient machine learning across multiple parties or computing nodes while ensuring information security during big data exchange, protecting the privacy of end-user and personal data, and ensuring legal and regulatory compliance. The machine learning algorithms that can be used in federated learning are not limited to neural networks but also include important algorithms such as random forests. Federated learning is expected to become the foundation of next-generation AI collaborative algorithms and collaborative networks.
[0057] 2. Quantification:
[0058] Quantization is an effective and efficient method for reducing communication costs by transmitting a quantized model instead of the original full-precision model, while maintaining similar learning accuracy to the original full-precision model. In the transmission interaction task of collaborative learning, quantization can reduce communication overhead by reducing the number of bits representing the transmission model / transmitted data. For example, dual quantization schemes can be used to ensure convergence speed while reducing communication overhead by quantizing model parameters and gradients. Alternatively, lossy collaborative learning algorithms can be introduced to reduce communication overhead by quantizing global and local models. Quantization may introduce model errors when reducing the arithmetic precision of model updates, which may hinder the convergence of collaborative learning algorithms. In addition, due to the strict wireless resource constraints in wireless networks, the transmission interaction content varies not only in size, structure, and task, but also in quantization level and quantization accuracy.
[0059] 3. KL divergence (Kullback-Leibler divergence, KLD):
[0060] Relative entropy is also known as KL divergence, information divergence, or information gain. KL divergence measures the asymmetry between two probability distributions P and Q. It measures the average number of extra bits required to encode samples from P using a code based on Q. Typically, P represents the true distribution of the data, and Q represents the theoretical distribution, model distribution, or an approximation of P.
[0061] 2. Network quantization method provided by this application:
[0062] The network quantization method provided in the present application is an adaptive quantization scheme based on collaborative learning implemented on the basis of federated learning. Among them, collaborative learning uses a method of multiple global iterations to train a collaborative (global) model, and in each global iteration, the agents (for example, the first device and the second device) interact with each other once for information / model. In the present application, during the interaction process, compressed and quantized information / models are transmitted between the agents, and the quantization level of the transmitted information / model can be adjusted according to their own network environment. Among them, the adaptive quantization scheme based on collaborative learning is mainly composed of local updates (local model updates, quantization level updates) and global aggregation. The process of a local update and global aggregation can be called a global iterative process. Among them, in each global iterative process, the first device performs multiple local model updates locally, and sends the local update results to the second device for aggregation. For example, Figure 2 is a schematic diagram of a global iterative process provided in the present application. A global iteration process includes: global model transmission (including global model, KL divergence), local model updating, local model quantization (including model quantization and quantization level update), local quantized model transmission, and global model update (global aggregation, including model recovery and global model aggregation). Assuming that the t-th global model transmission is the starting point of the t-th global iteration process, the t-th global iteration process is completed after the local model update, local model quantization, local quantized model transmission, and global model update, where t is a positive integer. Optionally, when the t+1-th global model transmission starts, the t+1-th global iteration process starts, and so on, for multiple cycles of iteration until convergence.
[0063] The network quantization method provided in this application is described below through specific step examples.
[0064] For example, FIG3 is a flow chart of a network quantization method provided by the present application. The network quantization method is applied to the communication system shown in FIG1. For example, the network quantization method can be implemented by interaction between a first device and a second device, wherein the first device is a terminal device (e.g., a user in collaborative learning) and the second device is an access network device or a core network device (e.g., an aggregator in collaborative learning). The method includes the following steps:
[0065] S101, the second device sends the first model and the first KL divergence of the t-th aggregation process to the first device; correspondingly, the first device receives the first model and the first KL divergence of the t-th aggregation process from the second device.
[0066] Among them, the first model of the t-th aggregation processing is obtained by the second device performing aggregation processing on multiple second models updated and quantized for the t-th time from multiple first devices, that is, the global model in the global model transmission in the iterative process corresponding to Figure 2. The second model of the t-th updated and quantized first device is the local quantized model obtained by the first device updating and quantizing the local model, that is, the local quantized model obtained by the local model update and local model quantization in the iterative process corresponding to Figure 2. It can be understood that if the second device is an aggregator in collaborative learning, the second device can obtain the second models of multiple first devices respectively, and perform aggregation processing on the multiple second models to obtain a global model. Optionally, the multiple first devices can be multiple first devices in the same cluster (for example, accessing the same base station), or multiple first devices in different clusters (for example, accessing different base stations), and this application is not limited thereto. Optionally, S101 is only an example of the second device sending the first model and the first KL divergence of the t-th aggregation processing to any one of the multiple first devices. The second device can also send the first model and the first KL divergence of the t-th aggregation processing to the other multiple first devices respectively. This application does not limit this.
[0067] For the convenience of description, the first model of the t-th aggregation process is represented as g t , the second model of the t-th updated quantization of any first device among the multiple first devices (assuming it is the i-th first device) is expressed as Where t represents the tth global iteration process (including the tth aggregation process, the tth update quantization, etc.), Indicates quantization processing, w i (t) represents the third model trained for the tth time by the i-th first device (that is, the local model in the local model update in the one-iteration process shown in FIG2 ), where t and i are both positive integers. Optionally, when multiple first devices are devices in different clusters (for example, devices connected to different base stations), the third model trained for the tth time by the i-th first device in the u-th cluster is expressed as Then the second model of the t-th updated quantization of the i-th first device in the u-th cluster is expressed as
[0068] Specifically, g tThe global model is obtained by aggregating multiple local models for the second device, and the aggregation process includes model recovery and global model aggregation. For example, before S101, the second device receives the second model updated and quantized for the tth time from multiple first devices, and restores the multiple second models updated and quantized for the tth time from the multiple first devices respectively to obtain multiple sixth models restored for the tth time; and aggregates the multiple sixth models restored for the tth time to obtain the first model aggregated for the tth time. Among them, the multiple sixth models restored for the tth time are also multiple restored local models (for example, the local model restored by the second device for the i-th first device can be expressed as ). Then, the second device aggregates the multiple restored local models to obtain g t .
[0069] The first KL divergence of the t-th aggregation process represents the average difference between the multiple second models updated and quantized for the t-th time and the first model of the t-th aggregation process, that is, the KL divergence in the global model transmission in one iteration process shown in Figure 2. For example, the second device can calculate the average difference between the multiple local quantization models and the global model (that is, the first KL divergence of the t-th aggregation process) according to formula (1) or formula (2).
[0070]
[0071]
[0072] When the multiple first devices are multiple devices in the same cluster, the second device uses formula (1) to calculate the average difference between the multiple local quantization models and the global model. In formula (1), represents the first KL divergence (average difference) of the t-th aggregation process, N represents the number of first devices, represents the difference between the first model processed by the tth aggregation and the second model updated and quantized by the tth first device. When the multiple first devices are multiple devices in different clusters, the second device uses formula (2) to calculate the average difference between the multiple local quantization models and the global model. In formula (2), represents the first KL divergence (average difference) of the t-th aggregation process, N u represents the number of first devices in the u-th cluster (assuming there are M clusters in total), represents the number of all first devices in M clusters, It represents the difference between the first model of the t-th aggregation process in the u-th cluster and the second model of the t-th updated quantization of the i-th first device.
[0073] When the second device completes a global model update, the second device sends the first model g of the tth aggregation process to the first device. t and the first KL divergence ( or ).
[0074] S102 : The first device determines a second model updated and quantized for the t+1th time according to the first model processed by the tth aggregation and the first KL divergence.
[0075] The second model updated and quantized for the t+1th time corresponds to the local quantization model obtained by the local model update and local model quantization in one iteration process shown in FIG2 , and is compared with the second model updated and quantized for the tth time in S101 , which is the local quantization model in the new iteration process. For ease of description, the second model updated and quantized for the t+1th time of any one of the multiple first devices (assuming it is the i-th first device) is represented in this application as w i (t+1) represents the third model trained for the t+1th time by the i-th first device (that is, the local model in the local model update in the one-iteration process shown in FIG2 ). Optionally, when multiple first devices are devices in different clusters (for example, connected to different base stations), the third model trained for the t+1th time by the i-th first device in the u-th cluster is represented as Then the second model of the t+1th updated quantization of the i-th first device in the u-th cluster is expressed as
[0076] Specifically, when S102 is executed, it may include the following process:
[0077] (1) The first device determines the third model trained for the t+1th time based on the first model processed by the tth aggregation process and the third model trained for the tth time.
[0078] When the first device receives the first model (i.e., the global model g t ), the first device can use the global model to train the local model first to obtain the third model of the t+1th training (that is, the local model after the new round of training). In addition, when the first device trains the local model, it is necessary to additionally consider wireless network resources and KL divergence. Specifically, when the first device determines the third model of the t+1th training based on the first model of the tth aggregation processing and the third model of the tth training, the following training process can be included:
[0079] The first device determines the enhanced loss function for the t-th training based on one or more of the following information: a cross-entropy loss function corresponding to the third model trained for the t-th time, a cross-entropy loss function corresponding to the second model updated and quantized for the t-th time, wireless resources consumed by transmitting the second model updated and quantized for the t-th time, an upper limit of wireless resources that can be allocated by the first device, and a second KL divergence for the t-th training; the second KL divergence for the t-th training represents the difference between the third model trained for the t-th time and the first model processed for the t-th aggregation, and / or the difference between the second model obtained by the t-th update and quantization and the first model processed for the t-th aggregation;
[0080] The first device trains the first model of the t-th aggregation processing based on the enhanced loss function of the t-th training to obtain a third model of the t+1-th training.
[0081] For example, the above training process can be expressed by formula (3) and formula (4):
[0082]
[0083]
[0084] Among them, w i (t+1) represents the third model trained for the t+1th time, c t represents the quantization level of the third model of the t-th training, represents the enhanced loss function of the t-th training, f i (w i ( t )) represents the cross entropy loss function corresponding to the third model of the t-th training, represents the cross entropy loss function corresponding to the second model of the t-th update quantization, represents the wireless resources consumed by transmitting the second model of the tth updated quantization, B i represents the upper limit of wireless resources that can be allocated to the i-th first device, f i KD (w i ( t ), g t ) represents the difference between the third model trained at the tth time and the first model processed at the tth time. represents the difference between the second model obtained by the t-th update quantization and the first model of the t-th aggregation process. The second KL divergence of the t-th training represents the difference between the third model of the t-th training and the first model of the t-th aggregation process, and / or the difference between the second model obtained by the t-th update quantization and the first model of the t-th aggregation process; that is, the second KL divergence of the t-th training includes f iKD (w i (t),g t )and Wherein, η1, λ1, λ2 are coefficients of the above parameters, and the specific values are not limited in this application.
[0085] (2) The first device determines the quantization level of the third model trained for the t+1th time according to the first KL divergence of the tth aggregation process.
[0086] Among them, when the first device obtains the third model trained for the t+1th time (that is, the local model is updated), it can also determine the quantization level of the third model trained for the t+1th time (that is, the quantization level of the local model is updated). In addition, when the first device updates the quantization level of the local model, it also needs to consider the wireless network resources and KL divergence. Specifically, when the first device determines the quantization level of the third model trained for the t+1th time based on the first KL divergence of the tth aggregation process, the following update process can be included:
[0087] The first device updates the quantization level of the third model trained for the tth time based on one or more information including the first KL divergence of the tth aggregation processing, the wireless resources consumed by transmitting the second model obtained by the tth updated quantization, and the cross-entropy loss function of the second model obtained by the tth updated quantization, and obtains the quantization level of the third model trained for the t+1th time.
[0088] For example, the above update process can be expressed by formula (5):
[0089]
[0090] Among them, c t+1 represents the quantization level of the third model trained for the t+1th time, represents the difference between the second model obtained by the t-th update quantization and the first model obtained by the t-th aggregation process, represents the cross entropy loss function corresponding to the second model of the t-th update quantization, It represents the wireless resources consumed by the second model for transmitting the t-th updated quantization, η2, λ1 are the coefficients of the above parameters, and the specific values are not limited in this application.
[0091] (3) The first device determines the second model for updating quantization for the t+1th time based on the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time.
[0092] When the first device determines the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time (i.e., updates the local model and the quantization level of the local model), the first device may determine the second model updated and quantized for the t+1th time (i.e., updates the local quantized model). For example, the first device quantizes the third model trained for the t+1th time using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
[0093] Optionally, when multiple first devices are devices in different clusters, due to differences in models between different clusters, the first device may also consider the differences in heterogeneous models between different clusters when determining the second model for the t+1th update quantization, thereby narrowing the gap between heterogeneous models between different clusters. For example, the first device updates the third model trained at the t+1th update quantization based on the third KL divergence of the tth update quantization to obtain a fourth model trained at the t+1th update; wherein the third KL divergence of the tth update quantization represents the difference between the third model trained at the t+1th update and the third model trained at the tth training, and / or the difference between the fourth model trained at the t+1th update quantization and the third model trained at the tth training; the fourth model trained at the t+1th update quantization is a model obtained by quantizing the third model trained at the t+1th training using the quantization level of the third model trained at the t+1th training. The first device quantizes the fourth model trained at the t+1th training using the quantization level of the third model trained at the t+1th training to determine the second model for the t+1th update quantization. The above process can be expressed by formula (6) and formula (7):
[0094]
[0095]
[0096] in, represents the second model of the t+1th update quantization, w i (t+1)′ represents the fourth model trained for the t+1th time, f i KD (w i (t+1),w i (t)) represents the difference between the third model trained at time t+1 and the third model trained at time t, It represents the difference between the fourth model updated and quantized for the t+1th time and the third model trained for the tth time. η3 and λ2 are the coefficients of the above parameters. The specific values are not limited in this application.
[0097] It is understood that the calculation method adopted by any of the multiple first devices when executing S102 can refer to the description in formulas (3) to (7), which will not be repeated here. Optionally, if the multiple first devices include devices in different clusters, the above formulas (3) to (7) can be modified by referring to the corresponding descriptions of formulas (1) and (2) to obtain analogy, which will not be repeated here.
[0098] S103 , the first device sends the second model updated and quantized for the t+1th time to the second device; correspondingly, the second device receives the second model updated and quantized for the t+1th time from the first device.
[0099] For example, after the first device completes the t+1th update quantization processing, it can send the second model of the t+1th update quantization to the second device, that is, corresponding to the next global iteration process shown in Figure 2, so as to continue to execute the next global iteration process until the convergence condition is reached.
[0100] In this embodiment, the first device can receive the global aggregate information / model (i.e., the first model and the first KL divergence) sent by the second device, and update the local quantization model based on the global aggregate information / model. Due to the integration of global knowledge, the first device can achieve faster convergence when updating the local quantization model (i.e., the second model), thereby improving training efficiency. In addition, the first device can continue to upload the second model to the second device, thereby performing a new round of update quantization. Even if the use of quantization results in a certain loss of accuracy, this method still achieves good learning performance, can meet the local model accuracy requirements, and is conducive to reducing the consumption of transmission bandwidth.
[0101] III. Performance Analysis of the Network Quantization Method of This Application:
[0102] 1. Comparison of model accuracy:
[0103] For example, Figures 4 and 5 are schematic diagrams showing the comparison of the model accuracy of the network quantization method provided by this application with other collaborative learning schemes. The horizontal axis in Figure 4 represents the number of iterations, the vertical axis represents the training accuracy, and it is assumed that there are two clusters (including cluster 1 and cluster 2), and the dataset used is the MNIST dataset. The horizontal axis in Figure 5 represents the number of iterations, the vertical axis represents the training accuracy (%), and it is assumed that there is one cluster (cluster 1), and the dataset used is the CIFAR 10 dataset. According to Figures 4 and 5, it can be deduced that even though a certain loss of accuracy is caused by the use of quantization, the network quantization method provided by this application still achieves better learning performance. And according to Figures 4 and 5, it can be deduced that the training accuracy of the network quantization method provided by this application (abbreviated as AQeD) is close to that of the collaborative learning scheme without quantization (abbreviated as GCFL), and is significantly higher than the collaborative learning scheme (abbreviated as FL8Q) in which all transmitted models are quantized to 8 bits (8-bit).
[0104] 2. Bandwidth resource comparison:
[0105] For example, Figure 6 is a schematic diagram comparing the bandwidth resources of the network quantization method provided by this application with other collaborative learning schemes. The horizontal axis in Figure 6 represents the number of user devices in each cluster, and the vertical axis represents the total bandwidth consumption (in megahertz, MHz). Assume there are two clusters (cluster 1 and cluster 2), and the datasets used include the MNIST dataset and the CIFAR-10 dataset. Based on Figure 6, it can be deduced that the network quantization method provided by this application uses fewer bandwidth resources. Figure 6 shows the total bandwidth consumption within a cluster when the training model accuracy reaches 70%. Figure 8 shows how the total bandwidth consumption within a cluster varies with the number of users on the MNIST dataset and the CIFAR-10 dataset. Based on Figure 6, it can be deduced that as the number of user devices increases, bandwidth consumption begins to increase; when a certain number of user devices is reached, bandwidth consumption reaches a maximum, and then decreases as the number of user devices increases. Specifically, when the number of user devices is approximately less than 300 (that is, the number of user devices in each cluster is approximately less than 150), total bandwidth consumption increases with the number of users. When the number of UEs exceeds approximately 300 (that is, the number of UEs per cluster exceeds approximately 150), total bandwidth consumption decreases (because poor channel quality prevents some UEs from transmitting local models). Furthermore, Figure 6 shows that schemes with quantized models (such as AQeD and FL8Q) consume less total bandwidth than collaborative learning schemes without quantization (such as GCFL).
[0106] 3. Training time comparison:
[0107] For example, Figures 7 and 8 are schematic diagrams showing the comparison of the training time of the network quantization method provided by the present application with other collaborative learning schemes. Wherein, the abscissa in Figure 7 represents the number of user devices, and the ordinate represents the training time (seconds, s). It is assumed that there are two clusters (including cluster 1 and cluster 2), and the data set used is the MNIST data set. The abscissa in Figure 8 represents the number of user devices, and the ordinate represents the training time. It is assumed that there are two clusters (including cluster 1 and cluster 2), and the data set used is the CIFAR 10 data set. It can be deduced from Figures 7 and 8 that the network quantization method provided by the present application can achieve faster convergence because it integrates global knowledge. Moreover, the training time of the network quantization method provided by the present application is less than that of other collaborative learning schemes (such as FL8Q and GCFL). This is because local model training is based on the global experience of heterogeneous models. Furthermore, when the number of UEs is less than approximately 300 (i.e., less than approximately 150 UEs per cluster), training time increases as the number of UEs increases. However, when the number of UEs exceeds approximately 300 (i.e., more than approximately 150 UEs per cluster), training time decreases as the number of UEs increases. This is because poor channel quality prevents the server from successfully receiving the local quantization model, and transmission time increases with the number of UEs.
[0108] In order to realize the various functions in the method provided by the present application, the device or equipment provided by the present application may include a hardware structure and / or a software module, and realize the above-mentioned various functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether a certain function among the above-mentioned functions is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution. The division of modules in the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, the various functional modules in the various embodiments of the present application can be integrated into a processor, or they can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0109] FIG9 is a schematic diagram of an apparatus provided in the present application. The apparatus may include a module corresponding to each of the methods / operations / steps / actions described in any of the embodiments shown in FIG2 and FIG3 , and the module may be implemented as a hardware circuit, software, or a combination of hardware circuit and software.
[0110] The apparatus 900 includes a communication unit 901 and a processing unit 902, and is configured to implement the methods executed by the various devices in the aforementioned embodiments. For example, the apparatus may be referred to as a network quantization apparatus or a communication apparatus.
[0111] In one possible embodiment, the apparatus is a terminal device, or is located in a terminal device. Specifically, the communication unit 901 is used to receive the first model and the first KL divergence of the t-th aggregation processing from the second device; wherein, the first model of the t-th aggregation processing is obtained by the second device performing aggregation processing on multiple second models of the t-th updated quantization from multiple first devices, and the first KL divergence of the t-th aggregation processing represents the average difference between the multiple second models of the t-th updated quantization and the first model of the t-th aggregation processing; t is a positive integer. The processing unit 902 is used to determine the second model of the t+1-th updated quantization based on the first model of the t-th aggregation processing. The communication unit 901 is also used to send the second model of the t+1-th updated quantization to the second device.
[0112] Optionally, the processing unit 902 is configured to determine a second model for updating and quantizing for the t+1th time based on the first model and the first KL divergence obtained through the tth aggregation process, including:
[0113] Determine a third model for the t+1th training according to the first model for the tth aggregation process and the third model for the tth training;
[0114] Determine the quantization level of the third model for the t+1th training according to the first KL divergence of the tth aggregation process;
[0115] The second model updated and quantized for the t+1th time is determined according to the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time.
[0116] Optionally, the processing unit 902 is configured to determine a third model trained for the t+1th time based on the first model processed by the tth aggregation process and the third model trained for the tth time, including:
[0117] Determine an enhanced loss function for the t-th training according to one or more of the following information: a cross-entropy loss function corresponding to the third model of the t-th training, a cross-entropy loss function corresponding to the second model updated and quantized for the t-th time, wireless resources consumed by transmitting the second model updated and quantized for the t-th time, an upper limit of wireless resources that can be allocated by the first device, and a second KL divergence for the t-th training; wherein the second KL divergence for the t-th training represents a difference between the third model of the t-th training and the first model of the t-th aggregation process, and / or a difference between the second model obtained by the t-th update and quantization and the first model of the t-th aggregation process;
[0118] The first model of the t-th aggregation process is trained based on the enhanced loss function of the t-th training to obtain the third model of the t+1-th training.
[0119] Optionally, the processing unit 902 is configured to determine a quantization level of a third model trained for the t+1th time according to the first KL divergence of the tth aggregation process, including:
[0120] According to one or more information of the first KL divergence of the t-th aggregation processing, the wireless resources consumed by transmitting the second model obtained by the t-th updated quantization, and the cross-entropy loss function of the second model obtained by the t-th updated quantization, the quantization level of the third model trained for the t-th time is updated to obtain the quantization level of the third model trained for the t+1-th time.
[0121] Optionally, the processing unit 902 is configured to determine the second model for updating quantization for the t+1th time according to the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time, including:
[0122] The third model trained for the t+1th time is quantized using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
[0123] Optionally, the processing unit 902 is configured to determine the second model for updating quantization for the t+1th time according to the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time, including:
[0124] The third model trained for the t+1th time is updated based on the third KL divergence of the t-th updated quantization to obtain a fourth model trained for the t+1th time; wherein the third KL divergence of the t-th updated quantization represents the difference between the third model trained for the t+1th time and the third model trained for the t-th time, and / or the difference between the fourth model trained for the t+1th updated quantization and the third model trained for the t-th time; the fourth model trained for the t+1th updated quantization is a model obtained by quantizing the third model trained for the t+1th time using the quantization level of the third model trained for the t+1th time;
[0125] The fourth model trained for the t+1th time is quantized using the quantization level of the third model trained for the t+1th time, and the second model updated and quantized for the t+1th time is determined.
[0126] The specific execution process of the communication unit 901 and the processing unit 902 in this embodiment can also refer to the corresponding description in the previous method embodiment, which will not be repeated here. The network quantization method implemented by the device can receive the global aggregation information / model (i.e., the first model and the first KL divergence) issued by the second device, and update the local quantization model based on the global aggregation information / model. Since the global knowledge is integrated, the first device can achieve faster convergence when updating the local quantization model (i.e., the second model), thereby improving training efficiency. Even if a certain loss of accuracy is caused by the use of quantization, the method still achieves good learning performance, can meet the local model accuracy requirements, and is conducive to reducing the consumption of transmission bandwidth.
[0127] In another possible embodiment, the device is a network device, or is located in a network device. Specifically, the communication unit 901 is used to receive multiple second models updated and quantized for the tth time from multiple first devices, where t is a positive integer. The processing unit 902 is used to aggregate the multiple second models updated and quantized for the tth time to obtain the first model and the first KL divergence of the tth aggregation process, and the first KL divergence of the tth aggregation process represents the average difference between the multiple second models updated and quantized for the tth time and the first model of the tth aggregation process. The communication unit 901 is also used to send the first model updated and quantized for the tth time and the first KL divergence to the multiple first devices respectively, and the first model updated and quantized for the tth time and the first KL divergence are used by the first device to determine the second model updated and quantized for the t+1th time.
[0128] Optionally, the processing unit 902 is configured to aggregate the multiple second models updated and quantized for the tth time to obtain the first model and the first KL divergence for the tth aggregation process, including:
[0129] Performing a restoration process on the multiple second models updated and quantized for the tth time to obtain multiple sixth models restored for the tth time;
[0130] Aggregate the multiple sixth models processed for the tth restoration to obtain the first model for the tth aggregation;
[0131] The first KL divergence of the t-th aggregation process is calculated based on the first model of the t-th aggregation process and the multiple second models updated and quantized for the t-th time.
[0132] Optionally, the communication unit 901 is also used to receive multiple second models updated and quantized for the t+1th time from multiple first devices; the processing unit 902 is also used to aggregate the multiple second models updated and quantized for the t+1th time to obtain the first model and first KL divergence of the t+1th aggregation processing.
[0133] The specific execution process of the communication unit 901 and the processing unit 902 in this embodiment can also refer to the corresponding description in the previous method embodiment, which will not be repeated here. The network quantization method implemented by the device can receive the local quantization model (i.e., the second model) from the first device. Even if the quantization results in a certain loss of accuracy, the method still achieves good learning performance, can meet the local model accuracy requirements, and is conducive to reducing the consumption of transmission bandwidth. In addition, due to the integration of global knowledge, it is beneficial for the first device to achieve faster convergence when updating the local quantization model, thereby improving training efficiency.
[0134] Figure 10 is a schematic diagram of a device provided by the present application, which is used to implement the network quantization method in the above method embodiments. The device 1000 can be a chip system or the device described in the above method embodiments.
[0135] Device 1000 includes a communication interface 1001 and a processor 1002. Communication interface 1001 may be, for example, a transceiver, an interface, a bus, a circuit, or a device capable of performing transceiver functions. Communication interface 1001 is used to communicate with other devices via a transmission medium, thereby enabling device 1000 to communicate with other devices. Processor 1002 is used to execute processing-related operations.
[0136] In one possible implementation, the device 1000 may be a terminal device or may be located in a terminal device. Specifically, the communication interface 1001 is configured to receive the first model and first KL divergence of the t-th aggregation process from the second device; wherein the first model of the t-th aggregation process is obtained by the second device aggregating multiple second models of the t-th updated quantization from multiple first devices, and the first KL divergence of the t-th aggregation process represents the average difference between the multiple second models of the t-th updated quantization and the first model of the t-th aggregation process; t is a positive integer. The processor 1002 is configured to determine the second model of the t+1-th updated quantization based on the first model of the t-th aggregation process and the first KL divergence. The communication interface 1001 is further configured to send the second model of the t+1-th updated quantization to the second device.
[0137] The specific execution process of the communication interface 1001 and the processor 1002 in this embodiment can also refer to the method performed by the communication unit and the processing unit in the embodiment of Figure 9, as well as the description in the previous method embodiment, which will not be repeated here. The network quantization method implemented by the device can receive the global aggregation information / model (i.e., the first model and the first KL divergence) issued by the second device, and update the local quantization model based on the global aggregation information / model. Since global knowledge is integrated, the first device can achieve faster convergence when updating the local quantization model (i.e., the second model), thereby improving training efficiency. Even if a certain loss of accuracy is caused by the use of quantization, the method still achieves better learning performance, can meet the local model accuracy requirements, and is conducive to reducing the consumption of transmission bandwidth.
[0138] In another possible implementation, the device 1000 may be a network device or may be located in a network device. Specifically, the communication interface 1001 is used to receive multiple second models updated and quantized for the tth time from multiple first devices, where t is a positive integer. The processor 1002 is used to aggregate the multiple second models updated and quantized for the tth time to obtain the first model and the first KL divergence of the tth aggregation process. The first KL divergence of the tth aggregation process represents the average difference between the multiple second models updated and quantized for the tth time and the first model of the tth aggregation process. The communication interface 1001 is also used to send the first model updated and quantized for the tth time and the first KL divergence to the multiple first devices respectively. The first model updated and quantized for the tth time and the first KL divergence are used by the first device to determine the second model updated and quantized for the t+1th time.
[0139] The specific execution process of the communication interface 1001 and the processor 1002 in this embodiment can also refer to the method executed by the communication unit and the processing unit in the embodiment of Figure 9, as well as the description in the previous method embodiment, which will not be repeated here. The network quantization method implemented by the device can receive the local quantization model (i.e., the second model) from the first device. Even if a certain accuracy loss is caused by the use of quantization, the method still achieves good learning performance, can meet the local model accuracy requirements, and is conducive to reducing the consumption of transmission bandwidth. In addition, due to the integration of global knowledge, it is beneficial for the first device to achieve faster convergence when updating the local quantization model, thereby improving training efficiency.
[0140] Optionally, the device 1000 may further include at least one memory 1003 for storing program instructions and / or data. In one embodiment, the memory and the processor are coupled. Coupling in this application is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, and is used for information exchange between devices, units or modules. The processor may operate in conjunction with the memory. The processor may execute program instructions stored in the memory. The at least one memory and the processor are integrated together.
[0141] The specific connection medium between the communication interface, processor, and memory described above is not limited in this application. For example, the memory, processor, and communication interface are connected via a bus. Bus 1004 is represented by a bold line in Figure 10 . The connection methods between other components are merely illustrative and not limiting. The bus can be categorized as an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 10 uses only a single bold line, but this does not imply that there is only one bus or only one type of bus.
[0142] In this application, a processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in this application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this application may be directly executed by a hardware processor, or by a combination of hardware and software modules within the processor.
[0143] In the present application, the memory may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or a volatile memory, such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the present application may also be a circuit or any other device that can implement a storage function, for storing program instructions and / or data.
[0144] The present application provides another device, which includes a processor coupled to a memory, and the processor is configured to read and execute computer instructions stored in the memory to implement the network quantization method in any one of the embodiments shown in FIG2 and FIG3 .
[0145] The present application provides a communication system, which includes the network device and terminal device in any one of the embodiments shown in Figures 2 and 3.
[0146] The present application provides a computer-readable storage medium. The computer-readable storage medium stores a program or instruction. When the program or instruction is executed on a computer, the computer executes the network quantization method in any of the embodiments shown in Figures 2 and 3.
[0147] The present application provides a computer program product. The computer program product includes instructions. When the instructions are executed on a computer, the computer executes the network quantization method in any one of the embodiments shown in FIG2 and FIG3 .
[0148] The present application provides a chip or a chip system, which includes at least one processor and an interface, wherein the interface and the at least one processor are interconnected by lines, and the at least one processor is used to run a computer program or instruction to execute the network quantization method in any of the embodiments shown in Figures 2 and 3.
[0149] The interface in the chip may be an input / output interface, a pin, or a circuit.
[0150] The chip system may be a system on chip (SOC) or a baseband chip, wherein the baseband chip may include a processor, a channel encoder, a digital signal processor, a modem, an interface module, and the like.
[0151] In one implementation, the chip or chip system described above in this application further includes at least one memory, in which instructions are stored. The memory may be a storage unit within the chip, such as a register, a cache, etc., or a storage unit of the chip (e.g., a read-only memory, a random access memory, etc.).
[0152] The technical solutions provided in this application can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a terminal device, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium.
[0153] In this application, under the premise that there is no logical contradiction, the various embodiments may reference each other, for example, the methods and / or terms between method embodiments may reference each other, for example, the functions and / or terms between device embodiments may reference each other, for example, the functions and / or terms between device embodiments and method embodiments may reference each other.
[0154] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.
Claims
1. A network quantization method, characterized in that: include: The first device receives the first model and the first KL divergence of the t-th aggregation process from the second device; The first model of the t-th aggregation process is obtained by aggregating multiple second models updated and quantized for the t-th time from multiple first devices by the second device, and the first KL divergence of the t-th aggregation process represents the average difference between the multiple second models updated and quantized for the t-th time and the first model of the t-th aggregation process; t is a positive integer; The first device determines, based on the first model and the first KL divergence of the t-th aggregation process, a second model for updating and quantizing the t+1-th time; The first device sends the second model updated and quantized for the t+1th time to the second device.
2. The method according to claim 1, characterized in that The first device determines, based on the first model and the first KL divergence of the t-th aggregation processing, a second model for updating and quantizing the t+1-th time, including: Determine a third model trained for the t+1th time based on the first model of the tth aggregation process and the third model trained for the tth time; Determining a quantization level of a third model trained for the t+1th time according to the first KL divergence of the tth aggregation process; The second model for quantization update for the t+1th time is determined according to the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time.
3. The method according to claim 2, characterized in that The determining of the third model trained for the t+1th time based on the first model of the tth aggregation processing and the third model trained for the tth time includes: Determine an enhanced loss function for the t-th training according to one or more of the following information: a cross-entropy loss function corresponding to the third model of the t-th training, a cross-entropy loss function corresponding to the second model updated and quantized for the t-th time, wireless resources consumed by transmitting the second model updated and quantized for the t-th time, an upper limit of wireless resources that can be allocated by the first device, and a second KL divergence for the t-th training; the second KL divergence for the t-th training represents a difference between the third model of the t-th training and the first model of the t-th aggregation process, and / or a difference between the second model obtained by the t-th update and quantization and the first model of the t-th aggregation process; The first model of the t-th aggregation process is trained based on the enhanced loss function of the t-th training to obtain a third model of the t+1-th training.
4. The method according to claim 2, characterized in that The determining, according to the first KL divergence of the t-th aggregation process, the quantization level of the third model trained for the t+1th time includes: According to one or more information of the first KL divergence of the t-th aggregation processing, the wireless resources consumed by transmitting the second model obtained by the t-th updated quantization, and the cross-entropy loss function of the second model obtained by the t-th updated quantization, the quantization level of the third model trained for the t-th time is updated to obtain the quantization level of the third model trained for the t+1th time.
5. The method according to any one of claims 2 to 4, characterized in that The determining, based on the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time, of the second model for updating quantization for the t+1th time comprises: The third model trained for the t+1th time is quantized using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
6. The method according to claim 5, characterized in that The step of quantizing the third model trained for the t+1th time using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time includes: The third model of the t+1th training is updated based on the third KL divergence of the tth updated quantization to obtain the t+1th The fourth model of the t+1th training; the third KL divergence of the t-th updated quantization represents the difference between the third model of the t+1th training and the third model of the t-th training, and / or the difference between the fourth model of the t+1th updated quantization and the third model of the t-th training; the fourth model of the t+1th updated quantization is a model obtained by quantizing the third model of the t+1th training using the quantization level of the third model of the t+1th training; The fourth model trained for the t+1th time is quantized using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
7. A network quantization method, characterized in that: include: The second device receives a plurality of second models updated and quantized for the tth time from the plurality of first devices, where t is a positive integer; The second device aggregates the multiple second models updated and quantized for the tth time to obtain a first model and a first KL divergence for the tth aggregation process, where the first KL divergence for the tth aggregation process represents an average difference between the multiple second models updated and quantized for the tth time and the first model for the tth aggregation process; The second device sends the first model and the first KL divergence of the t-th updated quantization to the multiple first devices respectively, and the first model and the first KL divergence of the t-th updated quantization are used by the first device to determine the second model of the t+1-th updated quantization.
8. The method according to claim 7, characterized in that The second device aggregates the multiple second models updated and quantized for the tth time to obtain a first model and a first KL divergence for the tth aggregation process, including: Restoring the plurality of second models updated and quantized for the tth time to obtain a plurality of sixth models restored for the tth time; Aggregating the plurality of sixth models processed for the t-th restoration process to obtain a first model processed for the t-th aggregation process; A first KL divergence of the t-th aggregation process is calculated based on the first model of the t-th aggregation process and the multiple second models updated and quantized for the t-th time.
9. The method according to claim 7 or 8, characterized in that The method further comprises: The second device receives a plurality of second models quantized for the t+1th time from the plurality of first devices; The second device aggregates the multiple second models updated and quantized for the t+1th time to obtain the first model and the first KL divergence of the t+1th aggregation process.
10. A network quantization device, characterized in that: include: A communication unit, configured to receive the first model and the first KL divergence of the t-th aggregation process from the second device; The first model of the t-th aggregation process is obtained by aggregating multiple second models updated and quantized for the t-th time from multiple first devices by the second device, and the first KL divergence of the t-th aggregation process represents the average difference between the multiple second models updated and quantized for the t-th time and the first model of the t-th aggregation process; t is a positive integer; A processing unit, configured to determine a second model for t+1th updated quantization based on the first model and the first KL divergence obtained by the tth aggregation process; The communication unit is further configured to send the second model updated and quantized for the t+1th time to the second device.
11. The device according to claim 10, characterized in that The processing unit is configured to determine the second model for the t+1th updated quantization based on the first model and the first KL divergence of the tth aggregation processing, including: Determine a third model trained for the t+1th time based on the first model of the tth aggregation process and the third model trained for the tth time; Determining a quantization level of a third model trained for the t+1th time according to the first KL divergence of the tth aggregation process; The second model for quantization update for the t+1th time is determined according to the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time.
12. The device according to claim 11, characterized in that The processing unit is configured to determine a third model trained for the t+1th time based on the first model of the tth aggregation processing and the third model trained for the tth time, including: Determine an enhanced loss function for the t-th training according to one or more of the following information: a cross-entropy loss function corresponding to the third model trained for the t-th time, a cross-entropy loss function corresponding to the second model updated and quantized for the t-th time, wireless resources consumed by transmitting the second model updated and quantized for the t-th time, an upper limit of wireless resources that can be allocated by the first device, and a second KL divergence; the second KL divergence represents the difference between the third model trained for the t-th time and the first model processed by the t-th aggregation, and / or the difference between the second model obtained by the t-th update and quantization and the first model processed by the t-th aggregation; The first model of the t-th aggregation process is trained based on the enhanced loss function of the t-th training to obtain a third model of the t+1-th training.
13. The device according to claim 11, characterized in that The processing unit is configured to determine a quantization level of a third model trained for the t+1th time according to the first KL divergence of the tth aggregation process, including: According to one or more information of the first KL divergence, the wireless resources consumed by transmitting the second model obtained by the t-th updated quantization, and the cross-entropy loss function of the second model obtained by the t-th updated quantization, the quantization level of the third model trained for the t-th time is updated to obtain the quantization level of the third model trained for the t+1-th time.
14. The device according to any one of claims 11 to 13, characterized in that The processing unit is configured to determine the second model for t+1th updated quantization according to the quantization level of the third model trained for the t+1th time and the third model trained for the t+1th time, including: The third model trained for the t+1th time is quantized using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
15. The device according to claim 14, characterized in that The processing unit is configured to perform quantization processing on the third model trained for the t+1th time using the quantization level of the third model trained for the t+1th time, and determine the second model updated and quantized for the t+1th time, including: The third model trained for the t+1th time is updated based on the third KL divergence to obtain a fourth model trained for the t+1th time; the third KL divergence represents the difference between the third model trained for the t+1th time and the third model trained for the tth time, and / or the difference between the fourth model updated and quantized for the t+1th time and the third model trained for the tth time; the fourth model updated and quantized for the t+1th time is a model obtained by quantizing the third model trained for the t+1th time using the quantization level of the third model trained for the t+1th time; The fourth model trained for the t+1th time is quantized using the quantization level of the third model trained for the t+1th time to determine the second model updated and quantized for the t+1th time.
16. A network quantization device, characterized in that: include: a communication unit, configured to receive a plurality of second models updated and quantized for a t-th time from a plurality of first devices, where t is a positive integer; a processing unit, configured to aggregate the plurality of second models updated and quantized for the tth time to obtain a first model and a first KL divergence obtained through the tth aggregation process, wherein the first KL divergence obtained through the tth aggregation process represents an average difference between the plurality of second models updated and quantized for the tth time and the first model obtained through the tth aggregation process; The communication unit is further configured to send the first model and the first KL divergence of the t-th updated quantization to the multiple first devices respectively, and the first model and the first KL divergence of the t-th updated quantization are used by the first device to determine the second model of the t+1-th updated quantization.
17. The device according to claim 16, characterized in that The processing unit is configured to aggregate the plurality of second models updated and quantized for the tth time to obtain a first model and a first KL divergence for the tth aggregation process, including: Restoring the plurality of second models updated and quantized for the tth time to obtain a plurality of sixth models restored for the tth time; Aggregating the plurality of sixth models processed for the t-th restoration process to obtain the first model processed for the t-th aggregation process; The first KL divergence of the t-th aggregation process is calculated based on the first model of the t-th aggregation process and the multiple second models updated and quantized for the t-th time.
18. The device according to claim 16 or 17, characterized in that The communication unit is further configured to receive a plurality of second models updated and quantized for the t+1th time from the plurality of first devices; The processing unit is further configured to aggregate the plurality of second models updated and quantized for the t+1th time to obtain a first model and a first KL divergence for the t+1th aggregation process.
19. A terminal device, characterized in that: include: A processor and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the processor, the network device executes the method according to any one of claims 1 to 6.
20. A network device, characterized in that: include: A processor and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the processor, the terminal device executes the method according to any one of claims 7 to 9.
21. A communication system, characterized in that: The method comprises a terminal device and a network device, wherein the terminal device executes the method according to any one of claims 1 to 6, and the network device executes the method according to any one of claims 7 to 9.
22. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to perform the method according to any one of claims 1 to 9.
23. A chip system, characterized in that: The chip system includes a processor, a memory, and an interface, and the processor and the interface are used to execute the method according to any one of claims 1 to 9.
24. A computer program product, characterized in that The method comprises instructions which, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 9.