A bandit feedback based quantized distributed online proximal gradient descent optimization method

By introducing Bandit feedback and quantization strategies into wireless sensor networks, combined with proximal gradient descent techniques, the distributed online optimization problem under communication constraints and unbalanced topologies is solved, achieving sublinear convergence and low-overhead synergistic optimization.

CN121078462BActive Publication Date: 2026-02-10CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511606204.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-10
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Existing distributed optimization algorithms struggle to achieve efficient and reliable online collaborative optimization in scenarios with limited communication, unbalanced topologies, and bandit feedback, especially in sparse event detection tasks in wireless sensor networks, where suitable methods are lacking.

Method used

A quantized distributed online proximal gradient descent optimization method based on Bandit feedback is adopted. It integrates adaptive quantization communication strategy, proximal gradient descent technique and weight compensation strategy to construct a quantized distributed online proximal gradient descent algorithm suitable for non-equilibrium networks, and handles the problems of missing gradient information and limited communication.

Benefits of technology

Achieve sublinear growth in static regret convergence performance in communication-constrained and unbalanced network environments, reduce communication overhead, and adapt to distributed online optimization under complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121078462B_ABST
    Figure CN121078462B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of distributed online optimization and learning, and discloses a quantized distributed online proximal gradient descent optimization method based on Bandit feedback. The method aims to solve the distributed online composite optimization problem under the condition that the communication resource is limited, the network structure is unbalanced, and the gradient information of the loss function is difficult to obtain or even cannot be obtained. The technical scheme comprises the following steps: constructing an unbalanced network graph and determining the node neighborhood relationship, modeling a distributed online composite optimization problem based on Bandit feedback, designing a quantized distributed online proximal gradient descent algorithm combined with an adaptive uniform quantization strategy and a proximal gradient descent technology, and performing static regret analysis on the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distributed optimization and online learning technology, specifically relating to a quantitative distributed online proximal gradient descent optimization method based on Bandit feedback. Background Technology

[0002] In recent years, with the rapid development of applications such as edge computing, industrial IoT, and smart cities, distributed online optimization technology has played a crucial supporting role in a wide range of engineering scenarios. Traditional distributed optimization algorithms mostly rely on ideal communication conditions and complete gradient information, requiring a balanced network topology and unconstrained exchange of high-precision data between nodes. However, in real-world engineering scenarios, these assumptions are often difficult to meet, specifically in the following two aspects:

[0003] First, due to bandwidth limitations, energy constraints, and privacy protection requirements, nodes often struggle to obtain, or are unable to obtain, the precise gradient of the objective function, and can only learn and make decisions through partial function value feedback (i.e., bandit feedback). This type of limited information feedback is common in scenarios such as online control, ad delivery, and black-box optimization. Bandit feedback, as a representative strategy of limited information learning, provides a new approach to solving such problems. Its core lies in estimating gradient information based on the limited function values ​​and guiding the optimization process. However, its expansion into distributed environments, especially in non-equilibrium networks, still faces challenges, including how to design low-variance gradient estimation strategies and balance the contradiction between exploration and utilization.

[0004] Secondly, real-world communication networks often exhibit asymmetric topologies, such as unidirectional links and differences in communication capabilities among heterogeneous nodes in wireless sensor networks, rendering traditional consensus algorithms based on double random matrices unsuitable for direct application. Furthermore, to alleviate communication pressure, the exchanged information needs to be quantized and compressed; however, existing quantitative distributed optimization methods mostly focus on static optimization or full-information feedback scenarios, and are still relatively lacking in addressing online learning, non-smooth regularization terms, and finite feedback fusion problems.

[0005] Existing research has introduced methods such as gradient quantization, stochastic compression, and sparse aggregation to reduce communication overhead. However, these methods typically do not comprehensively consider consensus mechanisms on unbalanced graphs, gradient estimation bias under Bandit feedback, and the sequential optimization requirements in dynamic environments. Therefore, how to construct a distributed online optimization algorithm that is convergent, communication-efficient, and suitable for complex optimization problems under complex conditions of limited communication, incomplete feedback, and unbalanced topology remains a critical problem that urgently needs to be solved. Summary of the Invention

[0006] To address the technical problems existing in the background art, this invention proposes a quantized distributed online proximal gradient descent optimization method based on Bandit feedback. This method is applicable to communication-constrained non-equilibrium network environments and does not rely on the precise gradient information of the loss function. It can achieve efficient collaborative optimization in scenarios such as wireless sensor networks, for example, sparse event detection tasks, and possesses sublinear static regret convergence performance.

[0007] The core of this invention lies in integrating an incomplete information feedback mechanism, an adaptive quantization communication strategy, a proximal gradient descent technique, and a weight compensation strategy to overcome multiple challenges in distributed online composite optimization, such as limited communication resources, network graph imbalance, and missing gradient information. Specifically, the Bandit feedback mechanism addresses the difficulty or complete inability to obtain the gradient of the loss function; an adaptive uniform quantization strategy alleviates communication constraints; proximal gradient descent handles the non-smooth regularization term in the objective function; and a weight compensation strategy eliminates the dependence on network graph balance. Theoretical proof and experimental verification demonstrate that this method can still achieve sublinear growth in static regrets even in communication-constrained and imbalanced network environments, exhibiting good convergence performance and robustness.

[0008] To achieve the above-mentioned objectives, this invention employs the following technical method: a quantized distributed online near-end gradient descent optimization method based on Bandit feedback, comprising the following steps: S1, Unbalanced network graph construction: Based on the connection relationships of communication nodes in a wireless sensor network, a network graph model representing the information interaction between nodes is constructed, and the in-neighbor set and out-neighbor set of each node are determined; S2, Distributed online composite optimization modeling under Bandit feedback: The sparse event detection problem in a wireless sensor network is expressed as a distributed online composite optimization model under Bandit feedback, defining its decision variables, local time-varying loss function, non-smooth regularization term, and feasible constraint set; S3, Quantized distributed online near-end gradient descent algorithm design: Based on the optimization model, an unbalanced network graph, uniform quantization strategy, and near-end gradient descent method are integrated to construct a quantized distributed online near-end gradient descent algorithm suitable for bandwidth-constrained unbalanced networks; S4, Algorithm convergence analysis: Using static regret as a performance evaluation index, the convergence behavior of the algorithm is theoretically analyzed, and its upper bound on static regret is established.

[0009] Further, S1 specifically includes the following steps: S11, acquiring each communication node in the wireless sensor network; S12, constructing an unbalanced connected network graph. ,in Represents a network diagram. For a set of nodes, Represents the total number of nodes. S13, Define the network graph. The weight matrix in is ,in Represent a OK A matrix of real numbers in columns, using a matrix The Line number Column elements Representation Nodes To node The communication weight, whose value is between 0 and 1, when Sometimes, Otherwise there are ,in Represents a node Able to direct nodes Sending information; S14, Network diagram weight matrix It is a row random matrix, meaning that the sum of the elements in each row is 1; S15, based on the network graph ,definition and They are nodes The set of incoming and outgoing neighbors, that is, the set of neighbors that can be sent to the node. The set of nodes that send information, and the nodes that can receive it. The set of nodes that sent the information.

[0010] Furthermore, S2 specifically includes the following steps: S21, defining the decision variables of the node as... ,in express 3D real vector space; S22, definition The constraints are non-empty convex sets that are shared and known by all communication nodes, where express 2D real vector space; S23, defined in the 2D real vector space; Only nodes during round iteration The known local convex loss function is: And define the non-smooth convex regularization term known to all nodes as: ,in S24. Based on the aforementioned unbalanced network graph, the sparse event detection problem in wireless sensor networks is modeled as a distributed online composite optimization problem under Bandit feedback, as follows: ,in, This represents the total number of iterations of the algorithm. This represents the total number of sensor nodes.

[0011] Furthermore, S3 specifically includes the following steps: S31, parameter initialization: setting parameters. , , , , ,in It is the total number of iterations of the algorithm. yes The iteration step size at time is always positive and increases with time. The increase shows a monotonic, non-increasing trend. yes The quantization parameter at time has a value between 0 and 1, and it increases with... The increase shows a monotonic, non-increasing trend. For the iteration time, It is a quantification level parameter. This is the weight matrix of the network graph; S32, Variable Initialization: Initialize the decision variables of each node. Quantization intermediate value vector and weight compensation vector ,in Representative node The decision vector at the initial iteration. Representative node The quantization intermediate value vector during the initial iteration Representative node The weight compensation vector at the initial iteration. represent The first order identity matrix Column elements; S33, Quantization interval size setting: in the first... In each iteration, the quantization interval size vector of the quantizer is set as... ,in It is a constant used to adjust the size of the quantization interval. To represent multiplication, It is a regularization term The upper bound of the gradient, yes Iteration step size at time, yes Quantization parameters at time, For the iteration time, It is a set of elements that are all 1 3D column vector; S34, for each iteration S35. Each node executes the iterative update process of the algorithm; S36. Output the decision vector sequence of all nodes.

[0012] Further, in step S34, the iterative update process of the algorithm is specifically as follows: S34.1, Decision and Feedback: In the first step... In the round of iteration, nodes Make a decision Received Bandit loss feedback value ,in These are exploration parameters. It is a node From the Euclidean unit sphere The random vector obtained above, Representative vector Euclidean norm; S34.2, gradient estimation: node The gradient estimate at the decision point is calculated based on the single-point gradient estimation method, as shown below: ,in, Represents a node In the The decision vector during each iteration. Indicates in The estimated gradient at that point, It is the dimension of the vector. These are exploration parameters. It is a node From the Euclidean unit sphere The random vector obtained above, Representative vector The Euclidean norm, It is a node In the The local convex loss function during round iteration.

[0013] S34.3, Quantitative Decision Making: Nodes Quantify the decision-making process to obtain quantitative decision-making results. ,in Represents a uniform quantization function. It is a node In the The quantization intermediate value vector in the round of iteration, It is the first The quantization interval size vector in the round iteration.

[0014] S34.4 Consistent Update and Gradient Descent: Nodes From node Receive quantitative decision and its own quantitative decision-making By comparing and combining the weight matrix and weight compensation vector, consistency and gradient descent updates are performed to obtain intermediate variables. The details are as follows:

[0015]

[0016] in, It is a node In the Intermediate variables in round iteration, and These are nodes and In the The decision vector in the round of iteration, It is the total number of communication nodes. It is a weight matrix The Line number Column elements, It is a node In the Quantization decision vector in round iteration, It is a node In the Quantization decision vector in round iteration, and These are nodes and In the The quantization intermediate value vector in the round of iteration, It is the first The quantization interval size vector in the round iteration, It is the first The iteration step size in a round of iteration, It is a node In the Weight compensation vector in round iteration The first in One element, It is a node exist The gradient estimate at point .

[0017] S34.5, Near-end projection update: Node Using the near-end projection operator on intermediate variables Perform a projection update to obtain the decision variables for the next iteration. and quantization intermediate value vector Specifically:

[0018]

[0019] in, and These are nodes In the Decision variables and quantized intermediate value vectors in round iterations It's about reducing parameters. It is a regularization term. and They are the first The iteration step size and quantization level parameters in the round of iteration, It is a node In the Intermediate variables in round iteration, Representative vector The Euclidean norm, Indicates in the constraint set In, make the function When the minimum value is obtained Value, of which It is about The function.

[0020] S34.6, Weight Compensation Update: Node The weight compensation vector is updated based on the weight matrix to obtain a new round of weight compensation vector. The details are as follows: ,in, For nodes In the The weight compensation vector in the round of iteration, It is the total number of communication nodes. It is a weight matrix The Line number Column elements, It is a node In the The weight compensation vector in the round of iteration.

[0021] Further, in step S34.3, the uniform quantization function is constructed as follows: Let the vector to be quantized be... The quantification level parameter is The quantization intermediate value vector is The quantization interval size vector is ,in represent 3D real vector space, If the set represents positive integers, then the uniform quantization function... The Each component The definition is as follows:

[0022]

[0023] in, ; , , and They represent quantization vectors respectively Quantization vector Quantization intermediate value vector and quantization interval size vector The Each element.

[0024] According to the definition of the quantization function, when When the quantization error satisfies the following inequality:

[0025]

[0026] in, It is the dimension of the vector. Representative vector The Euclidean norm, Representative vector The infinite norm of .

[0027] Further, S4 specifically includes the following steps: S41, using an individual static regret evaluation algorithm to assess the performance of nodes. The static drawbacks are as follows:

[0028]

[0029] in, It is the total number of iterations of the algorithm. It is a node Total number of iterations Static regrets within It is the total number of communication nodes. Indicates the first Nodes generated during round iteration The decision vector, Indicates the first Only nodes during round iteration Known local convex loss function, This represents a regularization term that is known to all nodes. Indicates in the constraint set inside, when The optimal solution obtained when the value is minimized.

[0030] S42. Prove, based on convex optimization theory, that individual static regret is related to the total number of iterations. It exhibits sublinear growth, that is, when When it approaches positive infinity, The limit is 0.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] 1. Overcoming the limitations of communication graph symmetry and applicable to unbalanced network environments: This invention does not rely on the symmetry or topological balance assumptions of the communication graph, and can achieve effective information transmission and collaborative optimization in unbalanced network graphs with only row random weight matrices, significantly expanding the applicability of distributed online composite optimization algorithms in real-world network environments.

[0033] 2. Reduce communication overhead and adapt to communication-constrained scenarios: By introducing an adaptive uniform quantization mechanism, the decision vectors to be exchanged in each iteration are compressed with limited bits, which effectively reduces the communication load. At the same time, while ensuring optimization accuracy, the practicality and deployability of the algorithm in low-bandwidth and resource-constrained environments are improved.

[0034] 3. Integrating Bandit feedback and proximal gradient optimization to ensure static regret convergence performance: To address the practical challenge of obtaining or not obtaining gradient information at all, a Bandit feedback mechanism is adopted to construct gradient estimates under limited information conditions. Combined with proximal gradient descent techniques to handle non-smooth regularization terms, even under complex conditions of limited communication and network imbalance, sublinearly growing individual static regret can still be achieved, thereby ensuring the convergence and long-term optimization performance of the algorithm. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a schematic diagram of the implementation steps of the method of the present invention;

[0037] Figure 2 A random network diagram containing 20 nodes is provided for an embodiment of the present invention;

[0038] Figure 3 This is an example of the average dynamic regret effect of the algorithm under different quantization levels provided in the embodiments of the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] like Figure 1As shown, this invention provides a quantized distributed online proximal gradient descent optimization method based on Bandit feedback. This method considers the communication limitations of unbalanced networks, the difficulty or complete inability to obtain gradient information of the loss function, and the non-smooth regularization function in the objective function. It integrates an adaptive quantization communication strategy, a Bandit feedback mechanism, proximal gradient descent technology, and a weight compensation strategy to construct a quantized distributed online composite optimization framework. This enables online collaborative solving and dynamic optimization of a global composite objective function containing a common regularization term by multiple nodes, thereby reducing communication overhead while ensuring algorithm convergence performance. The method specifically includes the following steps:

[0041] Step 1: Unbalanced network graph construction. This involves constructing a network graph model representing the information interaction between nodes based on the connection relationships of communication nodes in the wireless sensor network, and determining the in-neighbor set and out-neighbor set of each node. Specifically: S11, Obtain the communication nodes in the wireless sensor network; S12, Construct the unbalanced connected network graph. ,in Represents a network diagram. For a set of nodes, Represents the total number of nodes. S13, Define the network graph. The weight matrix in is ,in Represent a OK A matrix of real numbers in columns, using a matrix The Line number Column elements Representation Nodes To node The communication weight, whose value is between 0 and 1, when Sometimes, Otherwise there are ,in Represents a node Able to direct nodes Sending information; S14, Network diagram weight matrix It is a row random matrix, meaning that the sum of the elements in each row is 1; S15, based on the network graph ,definition and They are nodes The set of incoming and outgoing neighbors, that is, the set of neighbors that can be sent to the node. The set of nodes that send information, and the nodes that can receive it. The set of nodes that sent the information.

[0042] Step 2: Distributed online composite optimization modeling under Bandit feedback. This involves formulating the sparse event detection problem in wireless sensor networks as a distributed online composite optimization model under Bandit feedback, defining its decision variables, local time-varying loss function, non-smooth regularization term, and feasible constraint set. Specifically: S21, Define the decision variables of the nodes as... ,in express 3D real vector space; S22, definition The constraints are non-empty convex sets that are shared and known by all communication nodes, where express 2D real vector space; S23, defined in the 2D real vector space; Only nodes during round iteration The known local convex loss function is: And define the non-smooth convex regularization term known to all nodes as: ,in S24. Based on the aforementioned unbalanced network graph, the sparse event detection problem in wireless sensor networks is modeled as a distributed online composite optimization problem under Bandit feedback, as follows: ,in, This represents the total number of iterations of the algorithm. This represents the total number of sensor nodes.

[0043] Step 3: Design of a quantized distributed online proximal gradient descent algorithm. Based on the optimization model, this involves integrating the non-equilibrium network graph, uniform quantization strategy, and proximal gradient descent method to construct a quantized distributed online proximal gradient descent algorithm suitable for bandwidth-constrained non-equilibrium networks. Specifically: S31, Parameter initialization: Set parameters. , , , , ,in It is the total number of iterations of the algorithm. yes The iteration step size at time is always positive and increases with time. The increase shows a monotonic, non-increasing trend. yes The quantization parameter at time has a value between 0 and 1, and it increases with... The increase shows a monotonic, non-increasing trend. For the iteration time, It is a quantification level parameter. It is the weight matrix of the network graph.

[0044] S32. Variable Initialization: Initialize the decision variables for each node. Quantization intermediate value vector and weight compensation vector ,in Representative node The decision vector at the initial iteration. Representative node The quantization intermediate value vector during the initial iteration Representative node The weight compensation vector at the initial iteration. represent The first order identity matrix Column elements.

[0045] S33, Quantization interval size setting: In the... In each iteration, the quantization interval size vector of the quantizer is set as... ,in It is a constant used to adjust the size of the quantization interval. To represent multiplication, It is a regularization term The upper bound of the gradient, yes Iteration step size at time, yes Quantization parameters at time, For the iteration time, It is a set of elements that are all 1 Dimensional column vector.

[0046] Step 34: For each iteration Each node performs the following steps: Step 34.1, Decision and Feedback: In the... In the round of iteration, nodes Make a decision Received Bandit loss feedback value ,in These are exploration parameters. It is a node From the Euclidean unit sphere The random vector obtained above, Representative vector The Euclidean norm.

[0047] Step 34.2, Gradient Estimation: Nodes The gradient estimate at the decision point is calculated based on the single-point gradient estimation method, as shown below: ,in, Represents a node In the The decision vector during each iteration. Indicates in The estimated gradient at that point, It is the dimension of the vector. These are exploration parameters. It is a node From the Euclidean unit sphere The random vector obtained above, Representative vector The Euclidean norm, It is a node In the The local convex loss function during round iteration.

[0048] Step 34.3, Quantitative Decision Making: Nodes Quantify the decision-making process to obtain quantitative decision-making results. ,in Represents a uniform quantization function. It is a node In the The quantization intermediate value vector in the round of iteration, It is the first The quantization interval size vector in the round of iteration; where the uniform quantization function is constructed as follows: Let the vector to be quantized be... The quantification level parameter is The quantization intermediate value vector is The quantization interval size vector is ,in represent 3D real vector space, If the set represents positive integers, then the uniform quantization function... The Each component The definition is as follows:

[0049]

[0050] in, ; , , and They represent quantization vectors respectively Quantization vector Quantization intermediate value vector and quantization interval size vector The Each element.

[0051] According to the definition of the quantization function, when When the quantization error satisfies the following inequality:

[0052]

[0053] in, It is the dimension of the vector. Representative vector The Euclidean norm, Representative vector The infinite norm of .

[0054] Step 34.4, Consistent Update and Gradient Descent: Node From node Receive quantitative decision and its own quantitative decision-making By comparing and combining the weight matrix and weight compensation vector, consistency and gradient descent updates are performed to obtain intermediate variables. The details are as follows:

[0055]

[0056] in, It is a node In the Intermediate variables in round iteration, and These are nodes and In the The decision vector in the round of iteration, It is the total number of communication nodes. It is a weight matrix The Line number Column elements, It is a node In the Quantization decision vector in round iteration, It is a node In the Quantization decision vector in round iteration, and These are nodes and In the The quantization intermediate value vector in the round of iteration, It is the first The quantization interval size vector in the round iteration, It is the first The iteration step size in a round of iteration, It is a node In the Weight compensation vector in round iteration The first in One element, It is a node exist The gradient estimate at point .

[0057] Step 34.5, Near-end projection update: Node Using the near-end projection operator on intermediate variables Perform a projection update to obtain the decision variables for the next iteration. and quantization intermediate value vector Specifically:

[0058]

[0059] in, and These are nodes In the Decision variables and quantized intermediate value vectors in round iterations It's about reducing parameters. It is a regularization term. and They are the first The iteration step size and quantization level parameters in the round of iteration, It is a node In the Intermediate variables in round iteration, Representative vector The Euclidean norm, Indicates in the constraint set In, make the function When the minimum value is obtained Value, of which It is about The function.

[0060] Step 34.6, Weight Compensation Update: Node The weight compensation vector is updated based on the weight matrix to obtain a new round of weight compensation vector. The details are as follows: ,in, For nodes In the The weight compensation vector in the round of iteration, It is the total number of communication nodes. It is a weight matrix The Line number Column elements, It is a node In the The weight compensation vector in the round of iteration.

[0061] Step 35: Output the decision vector sequence of all nodes.

[0062] Step 4: Algorithm convergence analysis, which uses static regret as a performance evaluation metric to theoretically analyze the convergence behavior of the algorithm and establish its upper bound on static regret. Specifically, Step 41: Evaluate algorithm performance using individual static regret, nodes... The static drawbacks are as follows:

[0063]

[0064] in, It is the total number of iterations of the algorithm. It is a node Total number of iterations Static regrets within It is the total number of communication nodes. Indicates the first Nodes generated during round iteration The decision vector, Indicates the first Only nodes during round iteration Known local convex loss function, This represents a regularization term that is known to all nodes. Indicates in the constraint set inside, when The optimal solution obtained when the value is minimized.

[0065] Step 42: Prove, based on convex optimization theory, that individual static regret is related to the total number of iterations. It exhibits sublinear growth, that is, when When it approaches positive infinity, The limit is 0, which shows that the algorithm framework described in this method is effective.

[0066] In this embodiment, the present invention applies a quantized distributed online proximal gradient descent algorithm based on Bandit feedback to the online distributed LASSO problem. In this problem, all nodes need to collaborate to minimize the following function:

[0067]

[0068] in It is a node exist The feature matrix at time step, It is a node exist The response vector at time t is defined as follows: ,in It is a predefined vector. It is a node exist The noise that is constantly subjected to It is a regularization parameter. Representative vector The Euclidean norm, Representative vector The 1 norm, represent OK A real matrix of columns, represent A real vector space of dimension , where It is a positive integer, and The value ranges from 1 to between.

[0069] Figure 2 This is a schematic diagram of the communication topology of the unbalanced network used in the embodiments of the present invention. The diagram abstractly describes the underlying communication environment in which the algorithm of the present invention operates, and its main features are as follows: Multiple dots in the diagram represent sensor nodes participating in distributed communication and computation, totaling 20 nodes, each with independent computation, storage, and communication capabilities; the straight lines connecting these nodes represent communication links between nodes, forming channels for information exchange; the network presents a complex, irregular topology, where the connections between nodes are not fully connected or simple star or ring topologies, but rather consist of some central nodes with high connectivity and some peripheral nodes. Based on these features, it can be seen that the algorithm of the present invention is applicable to unbalanced communication scenarios. This asymmetric characteristic between nodes is mathematically represented by a row random weight matrix, which corresponds completely to the description in the claims.

[0070] In such Figure 2 The network graph shown is used to execute a quantized distributed online proximal gradient descent algorithm based on Bandit feedback. The algorithm parameters are set as follows: , , , Based on this, we obtain the following: Figure 3 The average static regret performance chart is shown. Figure 3 Plotting the iteration time on the x-axis and the average cumulative static regret of all nodes on the y-axis, the average static regret performance of the algorithm under different quantization levels in the test problem is characterized. Experimental results show that the method of the present invention can still maintain good performance under the conditions of difficulty in obtaining gradient information of loss function and limited communication bandwidth. Moreover, the convergence performance of the algorithm is further improved with the increase of quantization level parameter, thus verifying its effectiveness and robustness.

[0071] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.

Claims

1. A quantitative distributed online proximal gradient descent optimization method based on Bandit feedback, characterized in that, Includes the following steps: S1. Unbalanced network graph construction: Based on the connection relationship of communication nodes in a wireless sensor network, construct a network graph model that represents the information interaction between nodes, and determine the in-neighbor set and out-neighbor set of each node. S2. Distributed online composite optimization modeling under Bandit feedback: The sparse event detection problem in wireless sensor networks is expressed as a distributed online composite optimization model under Bandit feedback, defining its decision variables, local time-varying loss function, non-smooth regularization term and feasible constraint set. S3. Design of Quantized Distributed Online Proximal Gradient Descent Algorithm: Based on the optimization model, the quantized distributed online proximal gradient descent algorithm suitable for bandwidth-constrained non-equilibrium networks is constructed by integrating the non-equilibrium network graph, uniform quantization strategy and proximal gradient descent method. S4. Algorithm Convergence Analysis: Using static regret as a performance evaluation index, the convergence behavior of the algorithm is theoretically analyzed, and its upper bound on static regret is established. S2 specifically includes the following steps: S21. Define the decision variables for the nodes as follows: ,in express 3D real vector space; S22, Definition The constraints are non-empty convex sets that are shared and known by all communication nodes, where express 3D real vector space; S23, defined in the section Only nodes during round iteration The known local convex loss function is: And define the non-smooth convex regularization term known to all nodes as: ,in Represents the set of real numbers; S24. Based on the aforementioned unbalanced network graph, the sparse event detection problem in wireless sensor networks is modeled as a distributed online composite optimization problem under Bandit feedback. The objective of this optimization problem is to minimize the global cumulative loss of all nodes across all iterations. Its mathematical model is expressed as follows: ,in, This represents the total number of iterations of the algorithm. This represents the total number of sensor nodes. S3 specifically includes the following steps: S31. Parameter Initialization: Set parameters , , , , ,in It is the total number of iterations of the algorithm. yes The iteration step size at time is always positive and increases with time. The increase shows a monotonic, non-increasing trend. yes The quantization parameter at time has a value between 0 and 1, and it increases with... The increase shows a monotonic, non-increasing trend. For the iteration time, It is a quantification level parameter. It is the weight matrix of the network graph; S32. Variable Initialization: Initialize the decision variables for each node. Quantization intermediate value vector and weight compensation vector ,in Representative node The decision vector at the initial iteration. Representative node The quantization intermediate value vector during the initial iteration Representative node The weight compensation vector at the initial iteration. represent The first order identity matrix Column elements; S33, Quantization interval size setting: In the... In each iteration, the quantization interval size vector of the quantizer is set as... ,in It is a constant used to adjust the size of the quantization interval. To represent multiplication, It is a regularization term The upper bound of the gradient, yes Iteration step size at time, yes Quantization parameters at time, For the iteration time, It is a set of elements that are all 1 3D column vector; S34. For each iteration Each node executes the iterative update process of the algorithm; S35. Output the decision vector sequence of all nodes; In step S34, the algorithm iterative update process is specifically as follows: S34.1, Decision-making and Feedback: In the first section In the round of iteration, nodes Make a decision Received Bandit loss feedback value ,in These are exploration parameters. It is a node From the Euclidean unit sphere The random vector obtained above, Representative vector The Euclidean norm; S34.2 Gradient Estimation: Nodes The gradient estimate at the decision point is calculated based on the single-point gradient estimation method, as shown below: , in, Represents a node In the The decision vector during each iteration. Indicates in The estimated gradient at that point, It is the dimension of the vector. These are exploration parameters. It is a node From the Euclidean unit sphere The random vector obtained above, Representative vector The Euclidean norm, It is a node In the Local convex loss function during round iteration; S34.3, Quantitative Decision Making: Nodes Quantify the decision-making process to obtain quantitative decision-making results. ,in Represents a uniform quantization function. It is a node In the The quantization intermediate value vector in the round of iteration, It is the first The quantization interval size vector in the round iteration; S34.4 Consistent Update and Gradient Descent: Nodes From node Receive quantitative decision and its own quantitative decision-making By comparing and combining the weight matrix and weight compensation vector, consistency and gradient descent updates are performed to obtain intermediate variables. The details are as follows: , in, It is a node In the Intermediate variables in round iteration, and These are nodes and In the The decision vector in the round of iteration, It is the total number of communication nodes. It is a weight matrix The Line number Column elements, It is a node In the Quantization decision vector in round iteration, It is a node In the Quantization decision vector in round iteration, and These are nodes and In the The quantization intermediate value vector in the round of iteration, It is the first The quantization interval size vector in the round iteration, It is the first The iteration step size in a round of iteration, It is a node In the Weight compensation vector in round iteration The first in One element, It is a node exist The gradient estimate at point ; S34.5, Near-end projection update: Node Using the near-end projection operator on intermediate variables Perform a projection update to obtain the decision variables for the next iteration. and quantization intermediate value vector Specifically: , in, and These are nodes In the Decision variables and quantized intermediate value vectors in round iterations It's about reducing parameters. It is a regularization term. and They are the first The iteration step size and quantization level parameters in the round of iteration, It is a node In the Intermediate variables in round iteration, Representative vector The Euclidean norm, Indicates in the constraint set In, make the function When the minimum value is obtained Value, of which It is about The function; S34.6, Weight Compensation Update: Node The weight compensation vector is updated based on the weight matrix to obtain a new round of weight compensation vector. The details are as follows: , in, For nodes In the The weight compensation vector in the round of iteration, It is the total number of communication nodes. It is a weight matrix The Line number Column elements, It is a node In the The weight compensation vector in the round of iteration.

2. The quantitative distributed online proximal gradient descent optimization method based on Bandit feedback according to claim 1, characterized in that, S1 specifically includes the following steps: S11. Obtain the communication nodes in the wireless sensor network; S12. Construct an unbalanced connected network graph. ,in Represents a network diagram. For a set of nodes, Represents the total number of nodes. Represents the set of links in a network graph; S13. Define the network diagram The weight matrix in is ,in Represent a OK A matrix of real numbers in columns, using a matrix The Line number Column elements Representation Nodes To node The communication weight, whose value is between 0 and 1, when Sometimes, Otherwise there are ,in Represents a node Able to direct nodes Send a message; S14, Network Diagram weight matrix It is a row random matrix, meaning that the sum of the elements in each row is 1; S15. Based on the network graph ,definition and They are nodes The set of incoming and outgoing neighbors, that is, the set of neighbors that can be sent to the node. The set of nodes that send information, and the nodes that can receive it. The set of nodes that sent the information.

3. The quantitative distributed online proximal gradient descent optimization method based on Bandit feedback according to claim 1, characterized in that, In step S34.3, the uniform quantization function is constructed as follows: Let the vector to be quantized be The quantification level parameter is The quantization intermediate value vector is The quantization interval size vector is ,in represent 3D real vector space, If the set represents positive integers, then the uniform quantization function... The Each component The definition is as follows: , in, ; , , and They represent quantization vectors respectively Quantization vector Quantization intermediate value vector and quantization interval size vector The One element; According to the definition of the quantization function, when When the quantization error satisfies the following inequality: , in, It is the dimension of the vector. Representative vector The Euclidean norm, Representative vector The infinite norm of .

4. The quantitative distributed online proximal gradient descent optimization method based on Bandit feedback according to claim 1, characterized in that, S4 specifically includes the following steps: S41. The performance of the individual static regret evaluation algorithm is adopted, and the nodes... The static drawbacks are as follows: , in, It is the total number of iterations of the algorithm. It is a node Total number of iterations Static regrets within It is the total number of communication nodes. Indicates the first Nodes generated during round iteration The decision vector, Indicates the first Only nodes during round iteration Known local convex loss function, This represents a regularization term that is known to all nodes. Indicates in the constraint set inside, when The optimal solution obtained when the value is minimized; S42. Prove, based on convex optimization theory, that individual static regret is related to the total number of iterations. It exhibits sublinear growth, that is, when When it approaches positive infinity, The limit is 0.

Citation Information

Patent Citations

  • Wireless network resource management optimization method and device, electronic equipment and storage medium

    CN116647873A

  • Quantitative distributed self-adaptive online optimization method and system under bandwidth limited scene

    CN117812616A