A Distributed Gradient Tracking Non-Convex Optimization Method Based on Zero-Order Gradient Technique

By introducing zero-step gradient technology and time-varying random graph model into the distributed optimization algorithm, a gradient estimator with variable sample capacity is designed, which solves the shortcomings of existing algorithms in non-convex scenarios and time-varying communication topology, and realizes a distributed optimization method with higher flexibility and adaptability.

CN116382087BActive Publication Date: 2025-05-27DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310418973.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2025-05-27
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

In actual applications, existing distributed optimization algorithms have limitations such as poor flexibility, inability to quickly converge in non-convex scenarios, and inability to adapt to time-varying communication topology.

Method used

A non-convex optimization method for distributed gradient tracking based on zero-step gradient technology is proposed. By establishing a networked distributed multi-subject optimization model, using time-varying random graphs to describe information exchange relationships, designing a zero-step gradient estimator with variable sample capacity, and optimizing using a strategy of adapting first and then fusion.

Benefits of technology

This method can realize the effective solution of non-convex optimization tasks without relying on the expression of the explicit objective function, with higher flexibility and adaptability, and can be effectively optimized in time-varying communication scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116382087B_ABST
    Figure CN116382087B_ABST
Patent Text Reader

Abstract

The present invention proposes a distributed gradient tracking non-convex optimization method (VS-ZOGT) based on the zero-gradient technology, which solves the non-convex optimization problem in the networked system based on the gradient-free technology. In particular, a class of zero-gradient estimator frameworks are designed under the variable sample size method to achieve almost deterministic convergence of the algorithm in the form of fixed-step updates in the case of biased gradient estimation. In the scenario of unstable network communication, a random network model is used to perform the information exchange task of the distributed system, and independent non-coordinated step sizes are used to execute the optimization objectives on each agent. The present invention can ensure a stable reduction in the variance of the gradient estimation of the objective function and eliminate the conflict relationship between the convergence speed and the function query complexity in high-dimensional optimization problems. Compared with the existing zero-order optimization methods, the method of the present invention is a more efficient non-convex problem optimization algorithm. In addition, the feasibility and effectiveness of the above technical solutions are proved through simulation experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of control and information technology, and in particular, to a distributed gradient tracking non-convex optimization method based on zero-gradient technology. Background Art

[0002] A multi-agent network is a network of intelligent agents tightly coupled by a group of agents according to a preset communication protocol, where each agent can be considered as a physical entity or an abstract individual. Each agent has simple information processing capabilities and communication capabilities between adjacent agents. In a multi-agent network, the attributes of a single agent are weak, and its perception, computing, communication, etc. capabilities are all very limited and are not sufficient to be used alone in complex task scenarios. However, when a group of agents form a tightly coupled network under a preset communication protocol, they can use communication to achieve the purpose of division of labor and cooperation, and then orderly allocate tasks to obtain solutions to complex task scenarios. The distributed optimization problem is a further exploration based on multi-agent networks, aiming to minimize the global optimization objective through the design of a class of distributed algorithms for the cooperative exchange of local information among multi-agent neighbors in a topological network. Benefiting from the advantages of high flexibility, high scalability, and strong adaptability of distributed optimization algorithms, the distributed optimization problem has been widely applied to many popular fields such as power grid scheduling, multi-robot trajectory planning and control, multi-sensor networks, and deep learning.

[0003] In recent years, with the continuous in-depth research, algorithms for solving distributed optimization problems have emerged continuously. Table 1 lists 5 classic optimization algorithms. According to the form of information communication in a multi-agent network, the modeled network topology models are mainly divided into two types: fixed topology model and time-varying topology model. According to the periodic behavior of topological evolution, the time-varying topology network is further divided into two categories: B-type connected network and time-varying random network. Among them, the B-type connected network means that the communication network corresponding to each moment k in the time interval (B, B + 1) may be a disconnected topological structure, but the aggregated topology of the communication network in the whole interval is a connected structure, while the time-varying random network has no such connectivity restriction, only requiring that the communication network is still a connected topological structure in the sense of expectation. In addition, according to the classification of the step size form of the distributed algorithm, there are mainly two choices: decaying step size and fixed step size. According to the applicability of the algorithm to the task cost function, there are mainly two types: convex optimization and non-convex optimization algorithms. According to the gradient requirement of the algorithm for the task objective function, there are mainly two categories: gradient-based optimization algorithms and gradient-free optimization algorithms. Existing distributed algorithms mainly use the exact gradient information of the function to update and estimate local variables, and rely heavily on the explicit expression of the function. Using a fixed topology model to describe the information exchange behavior limits the application of distributed algorithms in general scenarios. Specifically, in actual application requirements, the above algorithms mainly have the following defects:

[0004] First, in practical applications, all agents use the same fixed step size for local updates, resulting in poor flexibility.

[0005] Second, in practical applications, the algorithm cannot achieve fast convergence for black-box function form optimization tasks without the assumption of convexity conditions.

[0006] Third, in practical applications, the algorithm is only applicable to fixed-topology communication models and cannot be extended to scenarios with time-varying communication task requirements.

[0007] Table 1 Characteristics of Five Optimization Algorithms

[0008] Summary of the Invention

[0009] To address the limitations and accuracy defects existing in existing adversarial sample generation technologies, the present invention proposes a distributed zero-order gradient tracking optimization algorithm based on variable sampling capacity (VS-ZOGT), aiming to complete the decision-making scheme for task objectives while taking into account both the iteration speed and accuracy.

[0010] Technical Solution of the Present Invention:

[0011] A distributed gradient tracking non-convex optimization method based on zero-order gradient technology, the specific steps are as follows:

[0012] Step 1: Establish a networked distributed multi-agent optimization model, specifically to find the minimum solution of the function f(x), and the optimization model is represented by the following formula:

[0013]

[0014] where f i (x):R d →R is a non-convex local privacy function on node i, d is the dimension of the optimization variable, R represents the real number field, and n is the number of nodes in the network. x represents the decision estimation information; f(x) has a minimum value:

[0015] Step 2: Construct a strongly connected graph model in the expected sense, and use a time-varying random graph to characterize the complex information exchange relationship in the networked system. The time-varying random graph is represented as G(k):{V,E(k)}, where V:{1,2,…,n} is the node set, is the edge set where information exchange occurs at node k. At each k moment, node i performs local operations and exchanges non-private information with its neighbor node j, that is, (i,j)∈E(k), j∈N i (k). N i (k) represents the neighbor set of node i at time k, and i∈N always exists during the information exchange phasei (k) describes the self-loop situation on node i. At time k, the adjacency matrix corresponding to G(k) is denoted as A(k). When (i, j) ∈ E(k), a in A(k) ij > 0; otherwise, a ij = 0. A(k) satisfies the properties of a doubly stochastic matrix:

[0016]

[0017] where 1 n denotes an n×1 vector with all elements being 1, denotes the transpose vector of vector 1 n , with a dimension of 1×n. In addition, the neighbor matrix corresponding to the connected graph in the sense of expectation is denoted as

[0018] Step 3: Design a zero-gradient estimator using the variable sample size technique as follows:

[0019]

[0020] where x i (k) represents the local decision estimate on node i at time k; μ k denotes the smoothing parameter used at time k, satisfying μ k ≤ μ k-1 and N(k) represents the number of sampling samples at time k, satisfying N(k - 1) ≤ N(k) and u q denotes the random noise μ used in the q-th sampling at time k. The random noise μ is generated by the uniform distribution U sp on the unit sphere.

[0021] Step 4: Set the initial decision state x i (0) ∈ R d , estimate the gradient information at the initial time and initialize the gradient estimation information as y i (0) = g i (0). Select a suitable non-coordinated step size 0 < α i < (1 - ρ A ) 2 / ρ A L, and update the local estimate of the solution of the privacy function f i (x) at time k + 1:

[0022]

[0023] where ρ A = E(‖A(k) - C‖) represents the spectral radius of the matrix A(k) - C in the sense of expectation, Let \(L\) denote the Lipschitz constant; \(y\) j (k) represents the global gradient estimation information at node \(j\) at time \(k\). According to Step 2, \(\rho\) A \(\in(0,1)\).

[0024] Step 5: Using \(x\) obtained in Step 4 i (k + 1), perform the gradient estimation update of the privacy function \(f\) i (x) at time \(k + 1\). The gradient estimator uses the form given in Step 4, and then we get:

[0025]

[0026] Step 6: According to \(g\) obtained in Step 5 j (k + 1), use the gradient tracking technique to perform the update of the global gradient estimation at time \(k + 1\):

[0027]

[0028] Step 7: Determine whether the termination condition is satisfied

[0029] \(x\) i (k + 1)= \(x\) i (k + 2)=…= \(x\) i (k + l)

[0030] where \(x\) i (k + l) represents the \(l\)-th update of the state after time \(k\). If the termination condition is not satisfied, continue to perform the updates in Steps 5 - 7.

[0031] Advantages of the present invention:

[0032] The present invention uses multiple autonomous agent nodes to distributively cooperate to execute the target optimization task, getting rid of the defects that the storage, computing, and communication capabilities of a single node are limited and it cannot be applied in complex scenarios. In addition, compared with the centralized optimization algorithm, the decentralized task execution feature enhances the stability and robustness of the algorithm.

[0033] The present invention does not make strong convexity or convexity assumptions on the objective function, which means that the step size parameter \(\alpha\) used in the algorithm i is only related to the Lipschitz coefficient \(L\) of the function and the network communication topology. Therefore, the algorithm relaxes the convexity restriction on the objective function, is more universal and general, and the update forms of the variables in Steps 4, 5, 6, and 7 are also applicable to the local optimization task of non-convex functions.

[0034] The present invention uses a time-varying random graph to characterize the evolution behavior of the communication network and describes the information transmission process between any nodes. Compared with the deterministic neighbor node set form \(N\) under the fixed topology networki In a time-varying random topological network, the set of neighbor nodes N i (k) is updated as time k changes, which can more effectively simulate the information exchange process in general scenarios. Therefore, this communication model has more advantages in universality and flexibility.

[0035] The present invention does not require any prior knowledge of the objective function, relaxes the smoothness assumption of the function, and uses a gradient-free method to perform optimization updates. It breaks through the requirement for an explicit objective function in traditional methods and can be applied to more challenging black-box optimization models. Since the zero-order gradient optimization does not use first-order gradient information, the randomness of the estimated gradient helps the algorithm cross poor local minima or maxima, making it possible for the optimization scheme to seek the global optimization solution with a certain probability.

[0036] The present invention designs a general form of the gradient estimator, providing theoretical support for the gradient-free optimization method. In step 4, when N(k) is designed as an integer c (c > 1), the gradient estimator will degenerate into a gradient estimation method under small batch sample data. When N(k) ≡ 1, the gradient estimator will further degenerate into a central difference variant form of the two-point gradient estimation method. In addition, different from the unstable noise statistics generated by the Gaussian distribution, the spherical uniform distribution U sp generated in step 4 is used as the bounded noise statistic added to the sampling variable during gradient estimation, which is more suitable for popularization and use in general application scenarios.

[0037] The present invention uses a method with variable sample size to weaken the d-dimensional dependence limitation in the common two-point method gradient estimator, and then extends the gradient tracking technology to the gradient-free objective task optimization framework, and reduces the variance caused by the introduction of the gradient estimation technology. The variance expression is as follows:

[0038]

[0039] The present invention uses a strategy of adaptation-then-combination (ATC) to update the local estimate of the global solution. The two key steps performed by the ATC strategy are respectively called the local information combination link and the aggregation link. In step 5, node j first performs the combination of local momentum information {x j (k), y j (k)}, and then node i receives the combined information of all neighbor nodes and performs weighted aggregation on them. In the local information combination link, the present invention uses a non-coordinated fixed step size α j to enhance the distributed performance of the local estimate and also ensure the convergence speed of the algorithm's optimization update. In addition, the use of the gradient tracking technology in step 7 enables the algorithm to still accurately converge to the optimal solution instead of the neighborhood interval of the optimal solution under a fixed step size.

[0040] The algorithm designed by the present invention only utilizes its own state information and key information in the neighbor set to execute the optimization strategy during the information exchange process, does not directly transmit sensitive information, and does not use any global parameters, so it has a good role in privacy protection.

[0041] In summary, for the optimization problem of the black-box model, the present invention develops a distributed zero-order gradient tracking algorithm. A zero-order estimator framework based on the variable sample size technique is proposed, enabling the high-dimensional optimization problem to achieve a linear convergence rate under the condition of biased gradient estimation. This algorithm guarantees a steady reduction in the variance of the gradient estimation while solving the trade-off problem between the linear convergence rate and the number of function value queries. In the case where the time-varying random network is not always connected, the agents using non-consistent fixed step sizes will almost surely converge to the same optimal solution. Finally, through simulations using the black-box adversarial model, the practicality of the algorithm is verified. Brief Description of the Drawings

[0042] Figure 1 is the flowchart of the present invention.

[0043] Figure 2 is the topological change demonstration diagram of the communication model in the present invention.

[0044] Figures 3A to 3E are respectively the schematic diagrams of the black-box attack loss curves under the VS-ZOTD algorithm, ZO-SGD algorithm, 2-Point+DGD algorithm, ZONE-M algorithm, and ZO-SVRG-ave algorithm.

[0045] Figure 4 is the schematic diagram of the experimental results of generating adversarial samples by five algorithms. Detailed Description of the Specific Embodiment

[0046] The following combines the drawings to describe in detail a distributed gradient tracking non-convex optimization method based on zero-order gradient technology proposed by the present invention.

[0047] Images are an important form for humans to convey key information in modern society. Benefiting from the rapid development of computer technology, scholars have sought means to combine deep neural networks with traditional control methods to handle various complex problems in the field of image recognition. In recent years, the technology of generating adversarial examples has been a hot topic in the field of image recognition. It aims to cause the failure of the image classification task of deep neural networks by adding subtle changes to the original images. Specifically, the essence of the technology of generating adversarial examples is to create and generate tiny noise images and add them to the original images to obtain new composite image samples, that is, adversarial examples. Compared with the original images, adversarial examples usually cannot be perceived as having obvious differences by human visual perception, but they can effectively interfere with numerically sensitive deep neural networks and thus lead to errors in the execution of image classification tasks. In this embodiment, the effectiveness of the algorithm proposed in the present invention for generating adversarial examples to counter black-box deep neural network image classification is tested. It should be noted that for the deep neural network model, only input and output information is available, and this form can be regarded as a typical zero-order non-convex optimization problem. As Figure 1 shown, the specific implementation steps are as follows:

[0048] Step 1: Design a black-box attack loss function:

[0049]

[0050]

[0051] g(x) = 0.5tanh(tanh -1 2v i +x),

[0052] where f i (x):R d →R is a non-convex local privacy function at node i, d is the dimension of the optimization variable, R represents the real number field, and n is the number of nodes in the network. (v i ,y i ) represents the feature pair of the i-th original image, v i is the digital description of the i-th original image and satisfies v i ∈[-0.5,0.5] d ,y i is the original label of the i-th image. Taking g(x) as the input of the deep neural network, the scores on K image classes are output by the function F(·), where F(·) = [F 1 (·),F 2 (·),…F K (·)]. Under the constraint of the tanh operator, the generated adversarial example g(x) is constrained in [-0.5,0.5] dInterval. The constant c is a regulation coefficient used to balance the adversarial degree of the generated samples in the task and the degree of change ‖g(x) - v i ‖ 2 with respect to the original image. In particular, ‖g(x) - v i ‖ 2 is also called l 2 distortion.

[0053] Step 2: As Figure 2 shown, construct a communication topology graph containing 10 multi-agent nodes, and use a time-varying random graph to characterize the complex information exchange relationship in the networked system. At each k moment, the connection probability between any two nodes is 0.15, and the weight of the adjacency matrix is constructed based on the Metropolis rule:

[0054]

[0055] where |d i | represents the number of neighbors of agent i.

[0056] Step 3: Design a zero-order gradient estimator

[0057]

[0058] where represents rounding down. u q represents the random noise μ used in the q-th sampling at the k moment. The random noise μ is generated by the uniform distribution U sp on the unit sphere.

[0059] Step 4: Initialize the total number of iterations T = 5000, the state x i (0) = 0, estimate the gradient information at the initial moment average gradient information y i (0) = g i (0). Select the step size on the i-th picture as 0.75 + 0.01 * i, and update the local estimate of the solution of the privacy function f i (x):

[0060]

[0061] Step 5: Use x i (k + 1) obtained in Step 4, perform the gradient estimation update of the privacy function f i (x) at the k + 1 moment. The gradient estimator uses the form given in Step 4, and then obtain:

[0062]

[0063] Step 6: According to g j (k + 1) obtained in Step 5, perform the update of the global gradient estimation at time k + 1:

[0064]

[0065] Step 7: Repeat the iterative process of Steps 4 - 6, and determine whether the generated adversarial samples cause errors in the image classification task of the deep neural network when the iteration is terminated.

[0066] To further illustrate the effectiveness of the technical solution of the present invention, this example conducts a comparative experiment with classical optimization algorithms such as ZO - SGD, ZONE - M, and ZO - SVRG - ave under the same parameters. The experimental results are shown in Tables 2, 3 and Figure 4 as follows. Table 2 shows that under the same settings, the VS - ZOGT algorithm has a better improvement effect on l 2 distortion compared with other algorithms. Among them, the improvement relative to the algorithm ZO - SGD is the largest, with an improvement of 53.2%. The improvement relative to the algorithm ZO - SVRG - ave is the smallest, but it also improves by 39.9%. Figure 4 And Table 3 shows the attack effect diagrams and the results of re - identifying the handwritten set labels of the original image under the distributed attack task with the label of 1 for the ZO - SGD, 2 - Point + DGD, ZONE - M, ZO - SVRG - ave, and VS - ZOGT algorithms. The results show that under the premise of ensuring the success of the attack task, the VS - ZOGT algorithm makes the smallest changes to the original image. The black - box attack loss curve under the VS - ZOTD algorithm proposed by the present invention is as Figure 3A shown; the black - box attack loss curve under the ZO - SGD algorithm is as Figure 3B shown; the black - box attack loss curve under the 2 - Point + DGD algorithm is as Figure 3C shown; the black - box attack loss curve under the ZONE - M algorithm is as Figure 3D shown; the black - box attack loss curve under the ZO - SVRG - ave algorithm is as Figure 3E shown. From the comparison, it can be concluded that the VS - ZOGT algorithm provided by the present invention has a significant advantage in the optimization speed compared with the other four algorithms.

[0067] Table 2 Experimental results of l 2 distortion of five algorithms

[0068]

[0069] Table 3 Experimental results of generating adversarial samples of five algorithms

[0070]

Claims

1. A distributed gradient tracking non-convex optimization method based on the zero-gradient technique, characterized in that, the specific steps are as follows: Step 1: Establish a networked distributed multi-agent optimization model, specifically to find the minimum solution of the function f(x), and the optimization model is expressed by the following formula: where f i (x): R d → R is a non-convex local privacy function at node i, d is the dimension of the optimization variable, R represents the real number field, n is the number of nodes in the network; x represents the decision estimation information; f(x) has a minimum value, that is Step 2: Construct a graph model that is strongly connected in the desired sense, and use a time-varying random graph to characterize the complex information exchange relationships in the networked system; the time-varying random graph is represented as G(k): {V, E(k)}, where V: {1, 2, …, n} is the set of nodes, and E(k) is the edge set for information exchange among nodes at time k; at each time k, node i performs local operations and exchanges non-private information with its neighbor node j, i.e., (i, j) ∈ E(k), j ∈ N i (k); N i (k) represents the neighbor set of node i at time k, and there is always i ∈ N i (k) to describe the self-loop situation on node i; at time k, the adjacency matrix corresponding to G(k) is denoted as A(k), and when (i, j) ∈ E(k), a ij in A(k) is > 0, otherwise, a ij = 0; A(k) satisfies the properties of a doubly stochastic matrix: where 1 n represents an n×1 dimensional vector with all elements being 1, represents vector 1 n is the transposed vector of, with a dimension of 1×n; in addition, the neighbor matrix corresponding to the connected graph in the expected sense is denoted as Step 3: Design a zero-gradient estimator using the variable sample size technique as follows: where x i (k) represents the local decision estimate at node i at time k; μ k represents the smoothing parameter used at time k, satisfying μ k ≤ μ k-1 and N(k) represents the number of sampling samples at time k, satisfying N(k - 1) ≤ N(k) and u q represents the random noise μ used in the q-th sampling at time k, and the random noise μ is generated by the uniform distribution U sp on the unit sphere; Step 4: Set the initial decision state x of each node i (0) ∈ R d , estimate the gradient information at the initial moment and initialize the gradient estimation information as y i (0) = g i (0); select an appropriate non - coordinated step size 0 < α i <(1 - ρ A ) 2 / ρ A L, update the local estimate of the solution of the privacy function f i (x) at the (k + 1)-th moment: where ρ A = E(‖A(k)-C‖) represents the spectral radius of matrix A(k)-C in the sense of expectation, L represents the Lipschitz constant; y j (k) represents the global gradient estimation information at node j at time k; according to step 2, ρ A ∈(0,1); Step 5: Use the x obtained in Step 4 i (k + 1) to perform the gradient estimation update of the privacy function f i (x) at time k + 1. The gradient estimator uses the form given in Step 4, and then we get: Step 6: Based on g j (k + 1) obtained in Step 5, perform the update of the global gradient estimation at time k + 1 using the gradient tracking technique: Step 7: Determine whether the termination condition is satisfied x i (k + 1) = x i (k + 2) = … = x i (k + l) where x i (k + l) represents the l-th update of the state after the k-th moment; if the termination condition is not satisfied, continue to execute the updates in steps 5 - 7.

Citation Information

Patent Citations

  • Global gradient dual-tracking distributed method based on heterogeneous mixed data

    CN115906362A

  • Efficient and secure gradient-free black box optimization

    US20200279155A1