Self-adaptive quantization radius distributed optimization method and device, equipment and storage medium

Through the adaptive quantization radius distributed optimization method, the quantization radius and gradient information transmission is dynamically adjusted, which solves the problems of wasted communication resources and slow convergence speed in traditional methods, and achieves efficient distributed optimization.

CN120358522APending Publication Date: 2025-07-22Shenzhen Big Data Research Institute Wuxi Innovation Center
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510629013.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing distributed optimization methods have problems such as waste of communication transmission volume and limited convergence speed under limited communication resources, especially in sensor networks, where traditional quantization radius fixation leads to redundancy and excessive communication overhead.

Method used

Adaptive quantization radius distributed optimization method is adopted, and the quantization radius is dynamically adjusted by constructing a convex optimization model and feasibility premise assumption, combining gradient descent operations and quantization functions, optimize gradient information transmission, reduce redundant areas, and improve convergence speed.

Benefits of technology

On the basis of ensuring optimization quality, communication resource consumption is reduced, gradient information transmission efficiency is improved, the iterative convergence process of the algorithm is significantly accelerated, and information redundancy is reduced. It is suitable for actual distributed optimization systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358522A_ABST
    Figure CN120358522A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive quantization radius distributed optimization method and device, equipment and a storage medium, and relates to the field of communication. Aiming at a distributed optimization problem, constructing a convex optimization model and a feasibility premise hypothesis; performing iterative calculation on the convex optimization model to obtain a previous iteration step result and gradient information of each current sub-node, performing gradient descent operation in combination with a communication protocol, and updating a decision variable; respectively calculating loss function gradient information by each sub-node according to the decision variable obtained in the previous step, and determining a radius value of a current iteration step based on a gradient difference value between a current gradient value and a quantization gradient in the previous step; and judging the gradient difference value, quantizing the radius value and gradient information of the iteration step based on a quantization function, and transmitting the quantization information to a central node. According to the scheme, the quantization precision is dynamically adjusted, and on the basis of ensuring the optimization quality, the gradient information is ensured to be smoothly transmitted to the central node, so that the gradient information can be better put into an actual distributed optimization system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of communications, and particularly to an adaptive quantization radius distributed optimization method, apparatus, device, and storage medium under limited communication resources. Background Art

[0002] In today's information age, with the rapid development of the Internet of Things and big data technologies, sensor networks, as an important means of data collection and processing, have been widely applied in various fields. However, the data processing and optimization problems in sensor networks are becoming increasingly complex, and traditional centralized optimization methods are facing bottleneck problems such as large communication volume and slow convergence speed. Therefore, distributed optimization methods have gradually become a research hotspot and shown great application potential in distributed systems such as sensor networks.

[0003] The core idea of the distributed optimization method is to decompose complex optimization problems into multiple sub-problems, which are processed in parallel by multiple processors or nodes, and information sharing and collaboration are achieved through network communication. This method can significantly reduce communication overhead and improve computational efficiency, especially suitable for large-scale, distributed sensor networks. In a sensor network, each sensor node usually can only obtain local observation information, and the distributed optimization method can utilize this local information to achieve the global optimization goal through collaboration and iteration. As an important research direction in distributed optimization, the key of the quantization algorithm lies in how to achieve efficient optimization through limited information exchange under communication constraints. The quantization algorithm usually involves converting continuous observation information or gradient information into discrete quantization values to reduce communication volume. However, the information loss introduced during the quantization process may lead to a decline in optimization performance. Therefore, how to design a reasonable quantization strategy to balance communication overhead and optimization performance has become an urgent problem to be solved.

[0004] In the related art, for the classical distributed optimization finite communication quantization algorithm, while ensuring the linear convergence of the algorithm, it compresses the important gradient information required for algorithm iteration into bits for transmission, where the quantization radius controls the quantization accuracy. Usually, the quantization grid radius is set to a fixed value, and this setting rule is only related to the number of set bits for transmission and the Lipschitz constant related to the decision variable, without considering the information of the specific objective function, and there is a certain redundancy in the projection area during actual execution, which may cause waste of total communication transmission resources and also limit the convergence speed of the distributed gradient descent algorithm to a certain extent. Summary of the Invention

[0005] The embodiments of the present application provide an adaptive quantization radius distributed optimization method under limited communication resources to solve the problems of waste of communication transmission resources and limitation of the convergence speed of distributed gradient descent.

[0006] On the one hand, the present application provides an adaptive quantization radius distributed optimization method, and the method includes:

[0007] For a convex optimization problem under limited communication resources, construct a corresponding convex optimization model and construct a feasibility premise assumption;

[0008] Initialize and iteratively calculate the decision variables of the convex optimization model, obtain the results of the previous iteration step and the gradient information of each current child node, and perform gradient descent operations in combination with the communication protocol to update the decision variable x k ;

[0009] Each child node calculates the gradient information of the loss function according to the decision variable obtained in the previous step, and determines the radius value of the current iteration step based on the gradient difference between the current gradient value and the quantization gradient of the previous step;

[0010] Judge and correct the gradient difference, quantize the radius value and gradient information of this iteration step based on the quantization function, and transmit the quantization information to the central node to restore the gradient information through the central node.

[0011] Specifically, the construction of the feasibility premise assumption includes:

[0012] For i = 1, 2,..., N, the objective function f i (·) is a μ-strongly convex function and has a continuous L-Lipschiz gradient; i represents the node index, and N is the total number of nodes;

[0013] The distributed gradient descent algorithm satisfies σ-function linear convergence;

[0014] The initialization and iterative calculation of the convex optimization model include:

[0015] Set the initial values of the variables x 0 and the radius r 0 , and let each child node's initial quantization gradient each child node's initial quantization radius

[0016] The execution of the gradient descent operation to update the decision variable includes:

[0017] In the k-th step of the iteration process, use the x obtained from the previous iteration k-1 and the quantization gradients collected from each child node At the central node, perform gradient descent operations according to the gradient descent algorithm where γ is the specified step size.

[0018] Specifically, the calculation of the gradient information of the loss function includes:

[0019] Traverse the 1st, .., Nth child nodes, and according to the data of the child nodes themselves, through the formula independently calculate the true gradient value of the i-th node at the current iteration step, where f i is the loss function of the i-th node, is the gradient information of f i .

[0020] Specifically, determining the radius value of the current iteration step based on the gradient difference between the current gradient value and the previous-step quantized gradient includes:

[0021] Traverse the 1st, .., Nth child nodes. After calculating the gradient of the current iteration step, subtract it from the quantized gradient of the previous iteration step, and determine the radius value of the current iteration step based on the norm of the difference; it is expressed as follows:

[0022]

[0023] where represents the radius value of the previous iteration step, represents the quantized gradient value of the (k - 1)-th iteration step, represents the gradient value of the current k-th iteration step, ||·|| ∞ is the infinity norm of the vector.

[0024] Specifically, judging the gradient difference and quantizing the radius value and gradient information of this iteration step based on the quantization function includes:

[0025] Define the radius quantization function quant. If the radius satisfies the set one-step judgment condition, then quantize it to according to the radius quantization function. If it does not satisfy, then correct the radius until the condition is met, and then quantize it to

[0026] Determine the quantized radius , and then use the gradient quantization function of the vector to perform quantization processing on the gradient vector of the current iteration step.

[0027] Specifically, constructing the gradient quantization function includes:

[0028] Construct a vector q, establish a square with side length 2r centered on the vector q, and divide the square into (2 b - 1) × (2 b - 1) square grids, where b is the given number of bits;

[0029] Project the vector c that satisfies ||c - q|| ∞ ≤ r onto the nearest grid point in the grid;

[0030] Define the gradient quantization function for vectors, which is defined component-wise as follows:

[0031]

[0032] where represents the b-bit quantization of the j-th dimensional component of c, and δ = r / (2 b - 1);

[0033] Define the radius quantization function r + = quant(Δ, r, r, b), and obtain the quantization radius expression as follows:

[0034]

[0035] Perform binary encoding on the quantized gradient information, quantization radius, and traversal multiple, and transmit them to the central node.

[0036] Specifically, the one-step judgment condition and correction are as follows:

[0037] Let the number of bits for storing the quantization radius be b r , and at each moment k, for each child node i, if then directly use the radius quantization function to quantize to to obtain the quantization result;

[0038] If then perform a loop operation, traverse the integer m starting from 1 until m satisfies and stop, find the smallest m that satisfies the condition , and correct the parameter of the quantization grid radius to Perform the quantization operation

[0039] On the other hand, the present application provides an adaptive quantization radius distributed optimization device, and the device includes:

[0040] A construction module for constructing a corresponding convex optimization model and constructing a feasibility premise assumption for a distributed convex optimization problem under limited communication resources;

[0041] An update module for initializing and iteratively calculating the decision variables of the convex optimization model, obtaining the results of the previous iteration step and the gradient information of each current child node, and performing a gradient descent operation in combination with the communication protocol to update the decision variable x k ;

[0042] A calculation module, configured to enable each sub-node to calculate the gradient information of the loss function according to the decision variables obtained in the previous step, and determine the radius value of the current iteration step based on the gradient difference between the current gradient value and the quantized gradient of the previous step;

[0043] A quantization module, configured to judge and correct the gradient difference, quantize the radius value and gradient information of this iteration step based on a quantization function, and transmit the quantization information to the central node, and the central node restores the gradient information.

[0044] In another aspect, the present application provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the adaptive quantization radius distributed optimization method described in any of the above aspects.

[0045] In another aspect, the present application provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the adaptive quantization radius distributed optimization method described in any of the above aspects.

[0046] The beneficial effects brought by the technical solutions provided by the embodiments of the present application at least include: The adaptive quantization radius distributed optimization algorithm designed in the present application solves to a certain extent the problems in the existing quantization algorithms that the grid quantization radius is fixed, there is redundancy in the projection area, resulting in the convergence rate of the gradient descent algorithm not meeting the actual needs and consuming excessive total communication bits. Especially in the case of limited communication resources, this situation is very common in actual needs. In this solution, by dynamically adjusting the quantization accuracy, on the basis of ensuring the optimization quality, it can ensure the smooth transmission of gradient information to the central node (less affected by the limited communication resources), and comprehensively considers potential error influencing factors, enabling it to be better applied to the actual distributed optimization system. Description of the Drawings

[0047] Figure 1 is a flowchart of the adaptive quantization radius distributed optimization method provided by the embodiments of the present application;

[0048] Figure 2 is an explanatory diagram of the two-dimensional situation of the quantization function used in the embodiments of the present application;

[0049] Figure 3 shows a comparison diagram of the error performance of transmitting information using different quantization algorithms;

[0050] Figure 4It is a comparison diagram of error performance for transmitting gradient information with a fixed quantization radius bit number and different quantization vector bit numbers;

[0051] Figure 5 The structural block diagram of the adaptive quantization radius distributed optimization device provided by the embodiment of the present application;

[0052] Figure 6 It shows the structural block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0053] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0054] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0055] Figure 1 It is the flowchart of the adaptive quantization radius distributed optimization method provided by the embodiment of the present application, including the following steps:

[0056] S1. For the convex optimization problem under limited communication resources, construct the corresponding convex optimization model and construct the feasibility premise assumptions;

[0057] In engineering practice and management decision-making, many application problems can be modeled as a mathematical optimization problem. In this embodiment, the convex optimization problem mainly discussed is the sum form of multiple convex functions as the objective function. In fact, since the total loss function is the sum of each individual loss function, many optimization problems can be converted into this form for solution.

[0058] For the distributed convex optimization problem under limited communication resources, to solve the optimization problem of minimizing the loss function, we consider a distributed system with N nodes to cooperate to solve the problem together. The distributed gradient descent algorithm is an efficient solution method, which distributes the computational burden of the algorithm through a series of computationally capable nodes. In this case, the N nodes calculate the corresponding gradients for their respective loss functions f i , and transmit the gradient information through the communication protocol. In the process of distributed iterative update of the gradient information, it involves transmitting the gradient information at the k-th moment of each child node to the center for the next iteration.

[0059] Constructing the feasibility premise assumptions is for the premise assumptions or premise parameter settings to implement the quantization algorithm of the embodiment of the present application. The premise assumptions in this solution can be set as follows:

[0060] Hypothesis 1: For \(i = 1, 2, \ldots, N\), \(f i (\cdot)\) is a \(\mu\)-strongly convex function and has a continuous \(L\)-Lipschitz gradient;

[0061] Hypothesis 2: The distributed gradient descent algorithm satisfies \(\sigma\)-function linear convergence.

[0062] S2. Initialize the decision variables of the convex optimization model and perform iterative calculations to obtain the results of the previous iteration step and the gradient information of each current child node. Combine the communication protocol to perform gradient descent operations and update the decision variable \(x k ;

[0063] This step requires first giving the initialization data of the decision variables and specifying the parameter information at the initialization time \(k = 0\). According to the hypothesis conditions and actual situations, etc., set the initial value of the variable \(x 0 and the radius \(r 0 , and let the initial quantized gradient of each child node The initial quantization radius of each child node where \(i\) is the node index and \(N\) is the total number of nodes. The subsequent iterative steps are to continuously update the decision variable \(x k .

[0064] S3. Each child node calculates the gradient information of the loss function respectively according to the decision variable obtained in the previous step, and determines the radius value of the current iteration step based on the gradient difference between the current gradient value and the quantized gradient of the previous step;

[0065] For each iterative step of the model, the true gradient of the loss function needs to be calculated separately. Gradient update is a necessary operation for model iteration, and the prerequisite for update is to obtain the decision variable \(x k-1 of the previous iterative step, then calculate the gradient difference between the true gradient of the current step and the quantized gradient of the previous step, and then determine the radius value of the current iteration step.

[0066] S4. Judge and correct the gradient difference, quantize the radius value and gradient information of this iteration step based on the quantization function, and transmit the quantization information to the central node to restore the gradient information through the central node.

[0067] For the judgment logic of the gradient difference, if the radius value of the current iteration step meets the set one-step judgment condition, then quantize it to by the quantization function. If it does not meet the condition, then correct the radius according to the designed algorithm until the condition is met, and then quantize it to to reduce the error and improve the convergence accuracy of the algorithm. Determine the quantization radius After that, use the quantization function of the vector to quantize the gradient vector of the current iteration step.

[0068] Furthermore, after binary encoding the information such as the quantized gradient and radius, transmit it to the central node, continue to execute the gradient descent step, and perform subsequent iterative training.

[0069] In summary, the adaptive quantization radius distributed optimization algorithm designed in this application solves to a certain extent the problems in the existing quantization algorithms, such as the fixed quantization radius of grid quantization, redundancy in the projection area, resulting in the convergence rate of the gradient descent algorithm not meeting the actual needs and consuming excessive total communication bits. Especially in the case of limited communication resources, this situation is very common in actual needs. In this solution, by dynamically adjusting the quantization accuracy, on the basis of ensuring the optimization quality, it can ensure the smooth transmission of gradient information to the central node (less affected by the limited communication resources), and comprehensively consider potential error influencing factors, enabling it to be better applied to the actual distributed optimization system.

[0070] In some specific embodiments, perform the gradient descent operation to update the decision variable, which can be specifically implemented through the following steps:

[0071] In any k-th iteration step during execution, use the x obtained in the previous iteration step (step k - 1) k-1 and the quantized gradients collected from each child node At the central node, according to the gradient descent algorithm Perform the gradient descent operation, where γ is the specified step size. Here, N represents the total number of child nodes.

[0072] When calculating the gradient information of the loss function in each iteration step, first traverse the 1st,.., N-th child nodes, and according to the data of the child nodes themselves, through the formula independently calculate the true gradient value of the i-th node at the current iteration step. Where f i is the loss function of the i-th node, is the gradient information of f i

[0073] ​After obtaining the gradient information of the loss function, we need to transmit this information to the central node. In practical applications, the gradient is transmitted in a way that does not exceed 64 - bit binary digits. If the true value is directly transmitted, it requires the network channel to have infinite bandwidth and the algorithm to be executed with infinite precision. Limited by factors such as communication capacity, the child nodes can only transmit a limited number of bytes of information. Therefore, we perform quantization operations on the gradient information so that it can be approximated by a value that can be represented by a limited number of bits. The commonly used quantization grid radius has a certain redundancy in the projection area during actual execution. To accelerate the convergence rate of the gradient descent algorithm and effectively reduce the consumption of the total communication bits to a certain extent, this application reasonably proposes to design the quantization radius according to the difference between the current - step gradient and the previous - step quantization value, thereby realizing a dynamic adaptive adjustment mechanism for the radius. This mechanism can accurately adjust the radius size according to the actual calculation process, minimize the redundant part of the quantization area jointly defined by the center point and the radius, effectively improve the quantization efficiency and reduce the information redundancy. This optimization strategy not only significantly accelerates the iterative convergence process of the algorithm, promotes the refinement and accuracy of the quantization area division, but also consumes fewer total communication bits at a specific precision, enabling the entire algorithm to exhibit higher efficiency and better performance.

[0074] This application provides an implementation method for calculating the radius value of the current iteration step, which is as follows:

[0075] Traverse the 1st,..., Nth child nodes. After calculating the gradient of the current iteration step, use the quantized gradient of the previous step and subtract it from the quantized gradient of the previous iteration step. Based on the norm of the difference value, determine the radius value of the current iteration step; it is expressed as follows:

[0076]

[0077] Among them, represents the radius value of the previous iteration step, represents the quantized gradient value of the (k - 1)th iteration step, represents the gradient value of the current kth iteration step, ||·|| ∞ is the infinity norm of the vector. This mechanism can accurately define the quantization area composed of the center point and the dynamic radius according to the radius size obtained by real - time calculation, thereby minimizing the redundant part within this area. This feature not only significantly improves the quantization efficiency but also effectively reduces the information redundancy, providing strong support for the performance optimization of the distributed optimization algorithm.

[0078] In some embodiments, the gradient difference is judged, and the radius value and gradient information of this iteration step are quantized based on the quantization function, including:

[0079] If the radius If the set one-step judgment condition is satisfied, it is quantified according to the radius quantization function as If not, the radius is corrected until the condition is satisfied, and then it is quantified as Determine the quantization radius After that, the gradient quantization function of the vector is used to quantize the gradient vector of the current iteration step.

[0080] To support this transmission process, this application adopts a quantization function to process the gradient and radius. Through the quantization operation, the bit approximation values of the true updated gradient information and radius are found, and this bit data is transmitted to the central node, and then the gradient information is restored through the central node. This function plays a crucial role in the transmission of the gradient.

[0081] As Figure 2 shown, considering a vector q in, a square with a side length of 2r can be established with it as the center. Then, the square is evenly divided into (2 b -1) × (2 b -1) small square grids, where b is the given number of bits, so a grid centered at q can be obtained. We can project any vector ∞ c that satisfies ||c - q|| ≤ r onto the nearest lattice point in the grid, and then construct the gradient quantization function for the vector Defined component by component as follows:

[0082]

[0083] where represents the b-bit quantization of the j-th dimensional component of c, and δ = r / (2 b -1);

[0084] Similarly, we introduce the radius quantization function: r + = quant(Δ, r, r, b). Further, we obtain the following specific expression for the quantization radius:

[0085]

[0086] Regarding the quantization error, the following conditions need to be satisfied:

[0087] Lemma 1: Let c ∈ R d be given, then for all c ∈ R ∞ that satisfy ||c - q|| d ≤ r, there is

[0088] Lemma 2: Let Δ ∈ R, then for all that satisfy ||Δ - r|| ∞For all Δ ∈ R such that Δ ≤ r, there is

[0089] According to Lemma 2, in the embodiments applied in this application, only When it is in the quantization grid with as the grid center and as the grid radius, can the error between the quantized radius and the true value be controlled within a very small range. Therefore, we adopt one-step judgment and correction, and continue quantization after correction. The specific content is as follows:

[0090] Let the number of bits storing the quantized radius be b r , at each moment k, for each child node i, if Then directly use the radius quantization function to quantize to to obtain the quantization result.

[0091] If Then perform a loop operation, traverse the integer m starting from 1 until m satisfies to stop, that is, find the smallest m that satisfies the condition , and correct the parameter of the quantization grid radius to Perform the quantization operation This step ensures that the size of the quantized radius meets the specific requirements of the iterative process, thereby effectively reducing the errors that may occur during the algorithm convergence process, and providing an important guarantee for improving the accuracy and stability of the distributed optimization algorithm.

[0092] After determining the quantized radius , then according to the definition of the gradient quantization function of the vector quantize the gradient vector of the current iteration step, where b q is the number of bits storing the quantized radius.

[0093] Finally, binary encode the quantized gradient information, quantized radius, one-step judgment, and multiple m, and transmit them to the central node to perform the subsequent gradient descent step.

[0094] Figure 3 Shows a comparison graph of the error performance of transmitting information using different quantization algorithms; specifically shows the comparison of the error performance of transmitting information between the classical quantization method, the adaptive quantization algorithm of this scheme, and non-quantization when the quantization bit radius number b r is 12 and the quantization vector bit number b q is 10. Among them, k represents the number of iteration steps, ||x k -x* || represents the error magnitude. When not quantifying, since the limitation of the number of communication bits is not considered, the effect is the best. It can be seen that the convergence speed of this method is significantly better than that of the classical quantization method, and it is almost the same as the convergence effect without quantization.

[0095] Figure 4 It is a comparison graph of the error performance of transmitting gradient information with a fixed quantization radius bit number and different quantization vector bit numbers. In this embodiment, the fixed quantization radius bit number is b r , considering different quantization vector bit numbers b q The error performance of transmitting information is considered. The figure shows the fixed quantization radius bit number b r is 4, considering different quantization vector bit numbers b q are 6, 8, 10, and 12 respectively. The comparison of the error performance of transmitting information between the classical quantization method, the adaptive quantization algorithm of this scheme, and without quantization is shown. Among them, k represents the number of iteration steps, ||x k -x * || represents the error magnitude. When not quantifying, since the limitation of the number of communication bits is not considered, the effect is the best. It can be seen that when the number of iteration steps is before 200, the convergence rates with different quantization vector bit numbers are almost the same as that without quantization. When the number of iteration steps is after 200, they are also very close. And when b q becomes larger and larger, the quantization vector gets closer and closer to the true vector, so the convergence effect is better and better, almost the same as the convergence effect without quantization.

[0096] Figure 5 It is the structural block diagram of the adaptive quantization radius distributed optimization device provided by the embodiment of the present application, including the following:

[0097] A construction module 510, configured to construct a corresponding convex optimization model and construct a feasibility premise assumption for a convex optimization problem under limited communication resources;

[0098] An update module 520, configured to initialize and iteratively calculate the decision variables of the convex optimization model, obtain the results of the previous iteration step and the gradient information of each current sub-node, perform a gradient descent operation in combination with the communication protocol, and update the decision variable x k ;

[0099] A calculation module 530, configured to calculate the gradient information of the loss function by each sub-node according to the decision variables obtained in the previous step, and determine the radius value of the current iteration step based on the gradient difference between the current gradient value and the previous quantization gradient;

[0100] A quantization module 540 is configured to judge and correct the gradient difference, quantize the radius value and gradient information of this iteration step based on a quantization function, and transmit the quantization information to a central node, and the gradient information is restored by the central node.

[0101] The adaptive quantization radius distributed optimization device provided by the embodiments of the present application can be applied to the adaptive quantization radius distributed optimization method provided in the above embodiments. For related details, refer to the above method embodiments. Their implementation principles and technical effects are similar, and will not be elaborated here.

[0102] It should be noted that the adaptive quantization radius distributed optimization device provided in the embodiments of the present application is only illustrated by the above division of each functional module / functional unit. In actual applications, the above functions can be allocated to different functional modules / functional units according to needs, that is, the internal structure of the adaptive quantization radius distributed optimization device is divided into different functional modules / functional units to complete all or part of the functions described above. In addition, the implementation manner of the adaptive quantization radius distributed optimization method provided in the above method embodiments and the implementation manner of the adaptive quantization radius distributed optimization device provided in this embodiment belong to the same concept. For the specific implementation process of the adaptive quantization radius distributed optimization device provided in this embodiment, refer to the above method embodiments, and will not be elaborated here.

[0103] Figure 6 The block diagram of the structure of a computer device provided by an exemplary embodiment of the present application is shown. It is a computer device such as a desktop computer, a laptop computer, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. Among them, the processor and the memory can be connected through a bus or other means. Among them, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, graphics processing units (GPUs), embedded neural network processors (NPUs) or other dedicated deep learning coprocessors, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.

[0104] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0105] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the above embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, to implement the methods in the above method embodiments. The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0106] In some embodiments, the computer device may also optionally include: a peripheral device interface and at least one peripheral device. The processor, the memory, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit, a display screen, and a keyboard.

[0107] The peripheral device interface can be used to connect at least one I / O (Input / Output) related peripheral device to the processor and the memory. In some embodiments, the processor, the memory, and the peripheral device interface are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor, the memory, and the peripheral device interface can be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0108] The display screen is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen is a touch display screen, the display screen also has the ability to collect touch signals on or above the surface of the display screen. The touch signals can be input to the processor as control signals for processing. At this time, the display screen can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen, which is set on the front panel of the computer device; in some other embodiments, there can be at least two display screens, which are respectively set on different surfaces of the computer device or in a folding design; in some other embodiments, the display screen can be a flexible display screen, which is set on the curved surface or the folding surface of the computer device. Even, the display screen can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0109] The power supply is used to supply power to each component in the computer device. The power supply can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0110] Those skilled in the art can understand that the structure shown in this embodiment does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0111] An embodiment of the present application also discloses a computer-readable storage medium. Specifically, the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the methods in the above method embodiments are implemented. Those skilled in the art can understand that all or part of the processes in the above method embodiments of the present application can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0112] This specific embodiment is only an interpretation of the present invention and does not limit the present invention. After reading this specification, those skilled in the art can make modifications to this embodiment without creative contributions as needed, but as long as it is within the scope of the claims of the present invention, it is protected by the patent law.

Claims

1. An adaptive quantization radius distributed optimization method, characterized in that The method includes: For the convex optimization problem under limited communication resources, constructing the corresponding convex optimization model and constructing the feasibility premise hypothesis; Initialize the decision variables of the convex optimization model and perform iterative calculations to obtain the results of the previous iteration step and the gradient information of each current child node. Combine the communication protocol to perform gradient descent operations and update the decision variable x k ; Each child node calculates the gradient information of the loss function respectively according to the decision variables obtained in the previous step, and determines the radius value of the current iteration step based on the gradient difference between the current gradient value and the quantized gradient of the previous step; Judge the gradient difference, quantify the radius value and gradient information of this iteration step based on the quantization function, and transmit the quantization information to the central node, and the central node restores the gradient information.

2. The method according to claim 1, characterized in that The construction of the feasibility premise hypothesis includes: For \(i = 1,2,\ldots,N\), \(f i (\cdot)\) is a \(\mu\)-strongly convex function with a continuous \(L\)-Lipschitz gradient; \(i\) represents the node index and \(N\) is the total number of nodes; The distributed gradient descent algorithm satisfies σ-function linear convergence; The initialization and iterative calculation of the convex optimization model include: Set the initial value x of the variable 0 and the radius r 0 , and set the initial quantization gradient of each child node The initial quantization radius of each child node The execution of the gradient descent operation to update the decision variables includes: During the k-th step of the iteration process, use the x obtained from the previous iteration step k-1 and the quantized gradients collected from each child node At the central node, according to the gradient descent algorithm perform a gradient descent operation, where γ is the specified step size.

3. The method according to claim 2, wherein The calculation of the gradient information of the loss function includes: Traverse the 1st, .., Nth child nodes, and according to the data of the child nodes themselves, through the formula independently calculate the true gradient value of the i-th node at the current iteration step, where f i is the loss function of the i-th node, is the gradient information of f i .

4. The method according to any one of claims 1-3, characterized in that, The determination of the radius value of the current iteration step based on the gradient difference between the current gradient value and the quantized gradient of the previous step includes: Traverse the 1st,..., Nth child nodes. After calculating the gradient of the current iteration step, subtract it from the quantized gradient of the previous iteration step, and determine the radius value of the current iteration step based on the norm of the difference value; It is expressed as follows: Among them, represents the radius value of the previous iteration step, represents the quantization gradient value of the (k - 1)-th iteration step, represents the gradient value of the current k-th iteration step, ||·|| ∞ is the infinity norm of the vector.

5. The method according to claim 4, wherein The judgment of the gradient difference and the quantization of the radius value and gradient information of this iteration step based on the quantization function include: Define the radius quantization function quant. If the radius meets the set one-step judgment condition, then quantize it according to the radius quantization function to If it does not meet the condition, correct the radius until the condition is met, and then quantize it to Determine the quantization radius After that, use the gradient quantization function of the vector to quantize the gradient vector of the current iteration step.

6. The method according to claim 5, characterized in that Constructing a gradient quantization function includes: Construct a vector q. Establish a square with side length 2r centered at the vector q, and divide the square evenly into (2 b - 1) × (2 b - 1) square cells, where b is the given number of bits; Project the vector c that satisfies ||c - q|| ∞ ≤ r onto the nearest lattice point in the grid; Defining a gradient quantization function for vectors, which is defined component by component as follows: where represents the b-bit quantization of the j-th dimensional component of c, and δ = r / (2 b - 1); Define the radius quantization function r + = quant(Δ, r, r, b), to obtain the quantization radius expression as follows: Encode the quantized gradient information, quantization radius, and traversal multiple into binary and transmit them to the central node.

7. The method according to claim 6, wherein One-step judgment condition and correction are as follows: Let the number of bits for storing the quantization radius be b r , at each moment k, for each child node i, if then directly use the radius quantization function to quantize into to obtain the quantization result; If then perform a loop operation, traversing integers m starting from 1 until m satisfies stop, find the smallest m that satisfies the condition and correct the parameter of the quantization grid radius to Perform the quantization operation 8. An adaptive quantization radius distributed optimization device, characterized in that The device includes: A construction module for constructing the corresponding convex optimization model and constructing the feasibility premise hypothesis for the convex optimization problem under limited communication resources; An update module, which is used to initialize and iteratively calculate the decision variables of the convex optimization model, obtain the results of the previous iteration step and the gradient information of each current child node, perform a gradient descent operation in combination with the communication protocol, and update the decision variable x k ; A calculation module for each child node to calculate the gradient information of the loss function respectively according to the decision variables obtained in the previous step, and determine the radius value of the current iteration step based on the gradient difference between the current gradient value and the quantized gradient of the previous step; A quantization module for judging and correcting the gradient difference, quantifying the radius value and gradient information of this iteration step based on the quantization function, and transmitting the quantization information to the central node, and the central node restores the gradient information.

9. A computer device, characterized in that, The computer device includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the adaptive quantization radius distributed optimization method as described in any one of claims 6 to 8.

10. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set or an instruction set is stored in the readable storage medium. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the adaptive quantization radius distributed optimization method as described in any one of claims 6 to 8.