Complex industrial system value self-organizing network optimization method based on dynamic game

By constructing a value self-organizing network optimization method for complex industrial systems based on dynamic game theory, this method solves the problem that the value interaction and value-added mechanism of nodes is difficult to characterize in existing technologies. It achieves efficient collaborative optimization under complex, dynamic, and privacy-preserving constraints, thereby improving the security and robustness of the system.

CN122222112APending Publication Date: 2026-06-16TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-03-12
Publication Date
2026-06-16

Smart Images

  • Figure CN122222112A_ABST
    Figure CN122222112A_ABST
Patent Text Reader

Abstract

The application discloses a complex industrial system value self-organizing network collaborative optimization method based on dynamic game, and relates to the technical field of industrial value collaborative optimization. The method comprises the following steps: abstracting a complex industrial system into a three-layer hierarchical heterogeneous network model, constructing node state evolution rules and a value function, and describing the generation, transmission and evolution mechanism of the value in the network; through a rolling time domain multi-stage dynamic game model, the conflict and cooperation between individual rational decision and system overall performance are explicitly described. The node performs local optimization based on a local information set, and updates the value weight online, so as to effectively adapt to environmental uncertainty and system dynamic evolution. Under the constraints of privacy protection and local information, theoretically guaranteed solving paths such as potential game + ADMM, distributed MPC or MARL are adaptively selected, efficient, safe and robust collaborative optimization is realized, and the comprehensive performance of the system is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial value collaborative optimization technology, and in particular to a method for optimizing the value of complex industrial systems based on dynamic game theory using self-organizing networks. Background Technology

[0002] As industrial systems continue to expand in scale, become increasingly complex in structure, and operate in highly dynamic environments, modern complex industrial systems typically exhibit characteristics such as multi-level, multi-entity, strong coupling, nonlinearity, and time-varying nature. Different nodes within the system (such as equipment, processes, units, or subsystems) not only generate their own value independently during operation but also interact with each other through various means, including material flow, energy flow, and information flow, forming complex value transfer and collaborative relationships.

[0003] Existing industrial system optimization and scheduling methods mainly focus on centralized optimization, static programming, or single-objective control frameworks. These methods typically assume a fixed system structure, complete information availability, or unified decision-making on the global state through a centralized controller. However, in real-world complex industrial systems, due to the large number of nodes, strong structural heterogeneity, time-varying operating states, and data privacy and security constraints, centralized methods often face problems such as high computational complexity, poor real-time performance, insufficient scalability, and privacy leakage risks, making it difficult to meet the needs of engineering applications.

[0004] In existing technologies, distributed optimization and multi-agent methods mainly focus on state consistency, resource allocation, or cost minimization, and have the following shortcomings: The lack of a systematic characterization of node value and its interactive value-added mechanism makes it difficult to reflect the generation, evolution and transfer of value in the network structure of complex industrial systems; Most methods do not explicitly consider the strategic game relationship between nodes, and cannot effectively describe the inherent conflict and coordination mechanism between individual rational decision-making and the overall performance of the system; Insufficient consideration of the dynamic nature of system operation, environmental uncertainty, and online evolution adaptability makes continuous optimization difficult; Under the conditions of privacy protection and local information, there is a lack of theoretically supported collaborative optimization mechanisms. Summary of the Invention

[0005] The purpose of this invention is to provide a value self-organizing network optimization method for complex industrial systems based on dynamic game theory, aiming to solve or improve at least one of the above-mentioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following solution: Value self-organizing network optimization methods for complex industrial systems based on dynamic game theory include: Complex industrial systems are abstracted into a three-layer hierarchical heterogeneous network model; each node includes an inherent attribute vector. State vector Environment vectors and decision variables ; Construct the state variables of the nodes and set the dynamic evolution rules for the node state variables; Construct a value function for each node in the model to calculate the instantaneous revenue of each node; Construct a rolling time-domain multi-stage dynamic game model, so that the optimization objective of the nodes in the model is to maximize the cumulative discounted total revenue; Obtain the information set of the model nodes and update the value weight of the information set online; Configure privacy-preserving communication between nodes; Local optimization is performed based on the local information set of the node to generate the node's decision variables; At the start of each control cycle, the node adaptively selects the solution path based on the industrial scenario.

[0007] Furthermore, the state variables of the nodes are constructed, including: The observable dynamic characteristics of a node use state variables Representation is performed based on the state vector. Mapped to obtain; If the state variable Unable to be measured directly, it is estimated through sensor observation sequences, and the expression is: In the formula, Let be the sensor observation value of node i at time t; It is a nonlinear observation function; To observe noise; Let i be the state variable of node i at time t; The state estimate is generated by the node's local state estimation algorithm, and the expression is: In the formula, This is the state estimate of node i; This is a state estimation algorithm; For historical observation; These are historical decision variables.

[0008] Furthermore, the dynamic evolution rules for node state variables include: For a linear dynamic system, the expression for the node state transition is: In the formula, for; This is the state transition matrix; For input control matrix; For disturbance terms; For a nonlinear dynamic system, the expression for the node state transition is: In the formula, For disturbance terms; It is a nonlinear state transition function.

[0009] Furthermore, the value function includes: The instantaneous value function of node i at time t is expressed as: In the formula, Let i be the value function of node i; , These are the state variables and decision variables for all other nodes, respectively; It is its own value function; For interactive value items; Let the decision cost function be used. As its own value weight; Weighting based on interaction value; For nodes The neighborhood set; The self-value function is expressed as: In the formula, This is the state value weight vector; To control the payoff matrix; The linear marginal benefit vector for the control action; The interactive value item, expressed as: In the formula, This is the state-state coupling matrix; For control-control coupling matrix; This is the state-control coupling matrix; The decision cost function is expressed as follows: In the formula, This is the secondary penalty coefficient; The linear cost vector for control.

[0010] Furthermore, the rolling time-domain multi-stage dynamic game model includes: The optimization objective of node i is to maximize the cumulative discounted total revenue, expressed as: In the formula, This represents the total cumulative discount revenue; Discount factor; It is a value function; The periodic step size; This is the value decay time constant.

[0011] Furthermore, the information set includes: Node i observes the information set at time t Including: state variables Self-value weight Interaction value weight Sharing intermediate variables with its neighbors, the expression is: In the formula, This is a state estimate; As its own value weight; Interaction value weight; Let be the dual variable between node i and its neighbor j; For interactive value basis functions Decision variables for node i The partial derivatives; It is a neighborhood set.

[0012] Furthermore, the value weights of the online updated information set include: Self-value weight and interaction value weight The expression for online updating via exponentially weighted moving average is: In the formula, Forgetting factor; This is an instantaneous self-value weight estimate; For instantaneous interaction weights.

[0013] Furthermore, privacy-preserving communications include: node Only to nodes Send the local gradient intermediates and dual variables, expressed as follows: In the formula, The intermediate quantity of the interaction gradient provided by node j to node i. Let be the dual variable between node i and its neighbor j; and The estimated state values ​​of nodes i and j at time t; and The control decisions made by nodes i and j at the previous time step; This is the penalty parameter.

[0014] Furthermore, local optimization is performed based on the node's local information set to generate the node's decision variables, including: node Local information set Construct a local approximate objective function, expressed as: In the formula, Let be the local approximate utility function for node i; Node i calculates the first derivative of the local approximate objective function with respect to its own control, expressed as: In the formula: The basis function of its own value with respect to decision variables The gradient; For interactive gradients; Unconstrained policy gradient ascent update is performed based on the first derivative, expressed as follows: In the formula, For the updated temporary decision variables; The learning rate; The temporary decision variables are projected onto the feasible region to obtain the final decision variables, expressed as follows: In the formula, Let i be the locally feasible control set of node i; This is the Euclidean projection operator.

[0015] Furthermore, at the beginning of each control cycle, the node adaptively selects the solution path based on the industrial scenario, including: If the model is available and the minimum eigenvalue of the Hessian matrix of the local optimization problem is not less than the preset non-negative tolerance threshold, then the potential game and alternating direction multiplier method (ADMM) are used to solve the problem. If there are hard dynamic constraints or a strong forward prediction capability is required, then distributed MPC is used to solve the problem. If the model is unknown, the target value is obtained through simulation. If the minimum eigenvalue of the Hessian matrix in the local optimization problem is less than the preset non-negative tolerance threshold, then distributed gradient estimation or decentralized multi-agent reinforcement learning (MARL) is used to solve the problem.

[0016] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention discloses a value self-organizing network optimization method for complex industrial systems based on dynamic game theory. The method constructs a three-layer hierarchical heterogeneous network model to systematically characterize the generation, evolution, and transmission mechanism of node value within the network. It introduces a rolling time-domain multi-stage dynamic game framework to explicitly model the conflict and synergy between individual rational decision-making and overall system performance. By updating the value weights of the information set online, it effectively addresses environmental uncertainty and system dynamic evolution, achieving continuous adaptive optimization. Furthermore, under privacy protection and local information constraints, it combines theoretically convergent solution paths such as latent game theory, distributed MPC, and MARL to construct a provably effective collaborative optimization mechanism, significantly improving the overall performance of complex industrial systems in terms of security, efficiency, and robustness. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 The decision variables in this embodiment A flowchart illustrating the acquisition process. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The purpose of this invention is to provide a value self-organizing network optimization method for complex industrial systems based on dynamic game theory, aiming to solve or improve at least one of the above-mentioned technical problems.

[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] like Figure 1 As shown, this invention provides a method for optimizing the value self-organizing network of complex industrial systems based on dynamic game theory, including: Step 1: Abstract the complex industrial system into a three-layer hierarchical heterogeneous network model, including: Construct a graph model G=(N,E), where N is the set of nodes and E is the set of directed or undirected edges connecting the nodes.

[0023] The set of nodes N is divided into disjoint subsets, including: The resource layer includes physical or human resources, such as equipment, energy units, and operators. The process layer represents the production or service execution unit, such as process steps, work tasks, production line modules, etc. The value layer represents the abstract entities that connect the system's output with external demands, such as products, customer orders, and service commitments. The expression for the node set N is: In the formula, This is a set of nodes in the resource layer. This is the set of nodes in the process layer; The set of nodes for the value layer; Let be the i-th node; m is the total number of nodes; Each node includes an intrinsic attribute vector. State vector Environment vectors and decision variables ; Inherent property vector This describes the static capability characteristics of a node, which remain unchanged or change slowly during system operation. It is used to characterize its capability boundaries, cost structure, and resource consumption characteristics, such as maximum equipment capacity, unit energy consumption, human skill level, and process precision. The expression is: In the formula, For the i-th node One attribute; Let i be the attribute dimension of the i-th node; State vector This represents the dynamic operating state of a node at time t, which evolves over time and reflects its current performance and health status, such as device load rate, task queue length, failure probability, and order fulfillment progress. The expression is: In the formula, For the first One state component; For the state dimension; Environment Vector , representing the external dynamic disturbances or driving signals acting on the nodes, originating from outside the system boundary, serving as the decision context, such as market demand fluctuations, new order arrivals, energy price changes, etc., and is expressed as; In the formula, This is the nth environmental component; For the environmental dimension.

[0024] Decision variables The behavior of the characterizing nodes that can be actively controlled, such as device start-up and shutdown, task rescheduling and resource reallocation, is estimated by the distributed collaborative optimization mechanism based on the current state. Adaptive generation of inherent attributes and environmental information.

[0025] Step 2, construct the node's state variables and set the dynamic evolution rules for the node's state variables, including: The observable dynamic characteristics of a node use state variables Representation is performed based on the state vector. Mapped to obtain; In this embodiment, the mapping method is a linear transformation, and the expression is: In the formula, This is a preset dimensionality reduction matrix.

[0026] If the state variable Unable to be measured directly, it is estimated through sensor observation sequences, and the expression is: In the formula, Let be the sensor observation value of node i at time t; It is a nonlinear observation function; The observed noise follows a zero-mean Gaussian distribution; Let i be the state variable of node i at time t; The state estimate is generated by a node-local state estimation algorithm, such as Kalman filtering, extended Kalman filtering, or particle filtering, and is expressed as: In the formula, This is the state estimate of node i; This is a state estimation algorithm; For historical observation; These are historical decision variables.

[0027] For a linear dynamic system, the expression for the node state transition is: In the formula, for; Let be the state transition matrix, describing the internal dynamics of the linear system; The input control matrix describes the effect of the control input on the state; This is a disturbance term.

[0028] For a nonlinear dynamic system, the expression for the node state transition is: In the formula, For disturbance terms; It is a nonlinear state transition function.

[0029] In implementation, the disturbance term in the nonlinear system The present invention employs the following process: First, a suitable disturbance model is selected based on the engineering scenario to estimate the state. The perturbation statistics are then considered in the local control optimization as expected, probabilistic constraints, or worst-case (min–max) scenarios. Preferred implementations include stochastic MPC with sample mean approximation (SAA), probabilistic constraints, etc. Furthermore, a forgetting factor or recursive least squares is used to update the perturbation statistics online to adapt to time-varying environments. The specification provides parameter ranges, sample sizes, window lengths, and verification methods in the embodiments to ensure repeatability and robustness of the implementation.

[0030] Step 3, construct the value function of the nodes in the model, the expression is: The instantaneous value function of node i at time t is expressed as: In the formula, Let i be the value function of node i, serving as the optimization objective in the dynamic game; , These are the state variables and decision variables for all other nodes, respectively; It is its own value function; For interactive value items; Let the decision cost function be used. The self-value weight represents the degree of importance a node attaches to its own performance; The interaction value weight represents the degree of importance attached to the collaborative effect between node i and its neighbor j. For nodes The neighborhood set is determined according to the graph model G=(N,E); The intrinsic value function reflects the marginal benefits of state performance and proactive scheduling, such as output, efficiency, and service level, and is expressed as: In the formula, This is the state value weight vector; To control the payoff matrix, it is a positive semi-definite matrix in this embodiment; The linear marginal benefit vector for the control action; Interactive value terms, representing material / energy transfer efficiency, synergistic gain, and conflict coupling, are expressed as follows: In the formula, Let be the state-state coupling matrix, representing the state matching degree between node i and node j; The control-control coupling matrix represents the synergistic benefits of joint scheduling. Let be the state-control coupling matrix, representing the value of the influence of neighbor j's control on the state of node i; The decision cost function, which includes energy consumption, maintenance costs, and delay penalties, is expressed as follows: In the formula, A secondary penalty coefficient for controlling energy consumption, equipment wear, or operational complexity; A linear cost vector for control, such as unit resource scheduling cost, delay penalty, etc.

[0031] Step 4: Construct a rolling time-domain multi-stage dynamic game model, where the optimization objective of each node in the model is to maximize the cumulative discounted total revenue, including: The rolling time-domain mechanism is as follows: at each decision time t, based on the current observation, a finite time-domain dynamic game of length T is constructed, and only the optimal strategy at the current time is executed. Subsequent strategies are recalculated based on the latest observation at the next time.

[0032] The optimization objective of node i is to maximize the cumulative discounted total revenue, expressed as: In the formula, This represents the total cumulative discount revenue; The discount factor is used to measure the present value of future returns. This indicates that there is no discount. This indicates a preference for recent returns; It is a value function; The periodic step size; This is the time constant for value decay, reflecting the degree to which decision-makers value future returns. The larger, As the month approaches 1, the more emphasis is placed on long-term returns.

[0033] Step 5: Obtain the information set of the model nodes and update the value weights of the information set online, including: like Figure 2 As shown, node i observes the information set at time t. Including: state variables Self-value weight Interaction value weight Sharing intermediate variables with its neighbors, the expression is: In the formula, This is a state estimate; As its own value weight; Interaction value weight; Let be the dual variable between node i and its neighbor j; For interactive value basis functions Decision variables for node i The partial derivatives; It is a neighborhood set.

[0034] The instantaneous self-value weight estimation process is as follows: Set nodes At any moment The actual performance observation values ​​are: In the formula, For nodes At any moment Observable performance indicators; For nodes At any moment Profits; For nodes At any moment The output; Assuming that the actual performance observations are linearly related to the node's own value function, the expression is: In the formula, As its instantaneous self-value weight; It is a value function; To observe noise; The gradient regression method is used to estimate the instantaneous self-value weights online. The expression is: In the formula, For a moment An estimate of its own value weight; It is a non-negative constant to avoid the denominator being zero.

[0035] The process for estimating the instantaneous interaction value weight is as follows: Define nodes with neighbors The collaborative observation index is expressed as: In the formula, For synergistic gain; For nodes and Joint efficiency during collaborative work (such as combined production line throughput, total revenue under shared energy). and They are nodes and Efficiency baseline when running alone; Assuming that the cooperative gain is linearly related to the basis function of the interaction value, the expression is: In the formula, Instantaneous interaction weight; As the basis function of interaction value, it characterizes the coupling effect between nodes (such as information sharing gain and load balancing benefit). To reduce noise during collaborative observation; The instantaneous interaction weights are then estimated as follows: In the formula, For instantaneous interaction weights.

[0036] Self-value weight and interaction value weight Updated online via Exponentially Weighted Moving Average (EWMA), the expression is: In the formula, The forgetting factor is such that the closer it is to 1, the more dependent it is on history, and the closer it is to 0, the more it responds to the latest observations. The instantaneous self-value weight estimate represents the node's value. At any moment The marginal contribution of its own operational performance to the overall revenue is used to characterize whether "at the current stage, we should focus more on its own efficiency or on collaborative relationships." It can be obtained based on the actual performance in the current period (such as profit and delivery rate) through regression, gradient estimation, or reinforcement learning. Instantaneous interaction weights are used to characterize nodes. With nodes The effectiveness or coupling strength of collaboration in the current operational phase reflects whether collaboration brings additional value gains. It is estimated based on the collaborative effects (such as material flow efficiency and joint energy efficiency) between node i and node j.

[0037] Step 6, configure privacy-preserving communication between nodes, including: To avoid directly sharing private information, nodes Only to nodes Send the local gradient intermediates and dual variables, expressed as follows: In the formula, The intermediate quantity of the interaction gradient provided by node j to node i. Let be the dual variable between node i and its neighbor j; and The estimated state values ​​of nodes i and j at time t; and The control decisions made by nodes i and j at the previous time step; This is a penalty parameter used to adjust the dual update step size; These quantities do not include the original state or value function parameters to meet privacy protection requirements.

[0038] Step 7: Perform local optimization based on the node's local information set to generate the node's decision variables, including: node Local information set Construct a local approximate objective function, expressed as: In the formula, Let be the local approximate utility function for node i; Node i calculates the first derivative of the local approximate objective function with respect to its own control, expressed as: In the formula: The basis function of its own value with respect to decision variables The gradient is computed locally at node i; The interaction gradient is provided by neighbor j; The above calculations for local areas only require first-order information, have low computational complexity, and are suitable for real-time industrial scenarios.

[0039] Unconstrained policy gradient ascent update is performed based on the first derivative, expressed as follows: In the formula, For the updated temporary decision variables; The learning rate can be reduced using a decay strategy to ensure convergence. The temporary decision variables are projected onto the feasible region to obtain the final decision variables, expressed as follows: In the formula, For node i, there is a locally feasible control set defined by physical, technological, or safety constraints (such as power limits or valve opening ranges). It is a Euclidean projection operator, ensuring that the output satisfies all hard constraints.

[0040] Step 8: At the beginning of each control cycle, the node adaptively selects the solution path based on the industrial scenario, including: Prioritize dynamic constraints / predictive requirements: If there are hard dynamic constraints related to future states (such as...) , (etc.) or requires a length of To ensure recursive feasibility and closed-loop stability, forward prediction is used, so path two: distributed MPC is adopted for the solution. Next, determine the applicability of the latent game + ADMM: Assuming the model is available, if the node objective function satisfies the latent game structural conditions (see the conditions described in Path 1, such as symmetric interaction weights and deriveable symmetric potential from the coupling kernel), and the local subproblems of the nodes corresponding to ADMM are related to... To approximate convexity, we adopt Path 1: Latent Game Theory + ADMM for solution; where the criterion for approximate convexity is... In the formula, The tolerance threshold for Hessian spectrum discrimination is used to absorb numerical errors and characterize "approximate convexity". For nodes At any moment The objective function of the local subproblem (consistent with the ADMM subproblem of path one) with respect to the decision variables The Hessian matrix, Its minimum eigenvalue; "model can be obtained" means that the nodes are known or can be identified online to obtain the local state transition model. Constraint Sets And computable (At least the gradient can be calculated or a stable numerical approximation can be obtained); "Model unknown" means that the above model and gradient are unavailable, and the objective function value can only be obtained through simulation or interactive sampling.

[0041] If the model is unknown (the target value can only be obtained through simulation sampling, and it is difficult to obtain the analytical gradient or state transition model), or does not meet the above potential game structure conditions and (approximate) convexity conditions, then path three is adopted, and distributed gradient estimation or decentralized multi-agent reinforcement learning (MARL) is used to solve the problem.

[0042] The above selection is constrained by real-time latency budget, communication bandwidth, and privacy requirements, and supports runtime monitoring and dynamic switching. The path selection logic can be implemented using pseudocode, and default thresholds are given in the embodiments for engineering implementation.

[0043] Path 1: When the node objective function satisfies the latent game conditions, the global Nash equilibrium is equivalent to maximizing the scalar latent function. This reduces a multi-objective game to a single-objective optimization problem, including: A set of sufficient structural conditions for the existence of a latent function is: If a node objective can be decomposed into self-objects and interactive coupling objects, the expression is: When the interaction weights are symmetrical Furthermore, the coupling kernel can derive a symmetric coupling potential (e.g., bilinear interaction and the coupling matrix satisfies...). ).

[0044] At this point, a latent function can be constructed, with the expression: satisfy This holds true for all nodes i.

[0045] Introducing local copy variables into the latent function and consistency constraints An augmented Lagrangian function is constructed, and the Alternating Direction Method of Multipliers (ADMM) is used for decomposition and solution. In the k-th iteration, node i solves the following local self-problem: In the formula, Let i be the local cumulative profit of node i; It is a node The feasible set represents the set of local constraints; For penalty parameters; Let be the Lagrange multiplier corresponding to node i; This is a consistency variable for the current iteration.

[0046] When transforming a game into a latent function, local subproblems are typically employed. To minimize. Within the framework of potential games, Can be regarded as a local pair This approximation ensures that the update direction of each node is consistent with the global goal.

[0047] When using ADMM to implement distributed solutions, consistency variables... The updates can employ various aggregation strategies to adapt to different engineering constraints.

[0048] If a central coordinator exists or global communication is allowed, the consistency variables behave differently: If there exists an additional set of global coupling constraints In this case, updates are performed using projection: In the formula, To reach the set Euclidean projection operator.

[0049] When the system encounters situations where centralized aggregation is unsuitable, such as privacy protection or communication restrictions, a decentralized neighborhood consensus strategy should be adopted: Each node Maintaining local aggregation variables Only with neighboring areas Exchange quantity And updated: Accordingly, the Lagrange multipliers are updated as follows: This strategy requires only neighborhood communication, does not disclose the original state or control commands, and meets the requirements of privacy and scalability.

[0050] In engineering implementation, to improve the convergence, stability, and efficiency of the algorithm in real industrial systems, the following practical strategies are adopted: For decision variables Normalize according to physical dimensions (e.g., power divided by rated value) to ensure consistent scale across dimensions; set initial values ​​for penalty parameters. Use hot start. , , .

[0051] Define the original residual With dual residual .

[0052] like This strengthens the consistency constraint: , ; like Then the constraints are weakened: , ; In this embodiment, .

[0053] For non-convex local subproblems, a proximal term is introduced to enhance stability: Proximal ADMM is formed to suppress excessively large update step sizes and improve convergence behavior.

[0054] The algorithm terminates when any condition is met: The change in the decision variable is less than the threshold: ; Both the original residual and the dual residual are less than the preset tolerance: , ; In this embodiment, Smaller values ​​are used for high-precision offline optimization, while larger values ​​are used for real-time online control.

[0055] Path 2, based on current state estimation Intermediate values ​​predicted with the neighborhood, for each node At any moment Build length is Predicted trajectory: The prediction model is linear, and its expression is: Or the prediction model is nonlinear, and the expression is: Node i solves a finite-time optimization problem with coupling terms at time t: In the formula, The forecast window length represents the time required for MPC to perform rolling forecasts and optimizations; it is an integer and reflects the forward-looking nature of the planning. A larger value indicates a more comprehensive and global prediction, but also increases computational cost; conversely, a smaller value results in a faster response. In this embodiment, the value range is [value range missing]. ; Discount factor; The instantaneous value function of a node; The following forms are available: decision variables are bilinearly coupled. Coupled with state consistency , or a combination of the two.

[0056] The coupling terms mentioned above are used to reflect the resource sharing, process connection, or collaborative operation relationships between adjacent nodes in a complex industrial system.

[0057] To ensure the recursive feasibility and asymptotic stability of the closed-loop system, a terminal cost is added to the end of the objective function, and terminal constraints are applied.

[0058] For nodes Cost of using linearized LQR terminals: In the formula, 0 represents a solution to the discrete algebraic Riccati equation, corresponding to the local linear feedback gain; Terminal invariant set In order to provide local feedback Invariant sets under: If the coupling between nodes is strong, tube-MPC can be used or a cluster-level collaborative terminal set can be designed.

[0059] The aforementioned terminal design theoretically guarantees the recursive feasibility of MPC and ensures closed-loop stability based on the Lyapunov reduction principle. In engineering implementation, It can be calculated using the standard LQR. The computational complexity can be reduced by using iterative approximation methods to numerically construct the model or by using conservative polyhedra as approximations.

[0060] Due to the coupling term's dependence on the domain control sequence Parallel optimization across nodes requires achieving approximate consistency through inner-layer iterations. Within control time t, the maximum number of executions... Wheel ADMM-style coordination: The local subproblem of the k-th round is expressed as: In the formula, The objective function is the finite-time optimization problem. The expression for consistent variable updates is: The dual variable update expression is: The inner iteration terminates when any of the following conditions are met: Original residual Dual residuals ; Reaching the maximum number of iterations .

[0061] Final implementation .

[0062] For scenarios sensitive to communication latency, the allowable number of inner iterations should be determined by quantifying network and computing performance.

[0063] Let the control period be Reserve a safety margin The estimated time for a single inner layer iteration is: The maximum allowed number of iterations is expressed as: In the formula, The time (ms) for solving a single-node local subproblem in a single run includes linearization, QP solution, and one gradient step. It is an aggregator calculation Time (ms); Use a 95%-percentile value for the round-trip time (ms) of the relevant network to ensure real-time performance; The communication transmission time is in milliseconds. This represents the number of round trips required for each inner iteration.

[0064] Path 3 involves obtaining necessary terms through numerical differencing, enabling each node to estimate its local gradient and perform a projection update. The specific steps are as follows: Step 1: Initialization Current control point The default value is used to avoid frequent communication during the estimation period.

[0065] The neighborhood value is fixed as And set the sample size Disturbance scale Step length Gradient smoothing factor Gradient clipping threshold ,node Control dimensionality .

[0066] Step 2: Disturbance Generation and Target Assessment right Generate random perturbation vector (Each component is independent, with a probability of 0.5).

[0067] Construct test points on both sides: When evaluating local targets, neighborhood values ​​are used in the coupling term. Or obtain according to the agreement: Step 3: Single-sample gradient estimation In the formula, It is a symbol vector.

[0068] Step 4: Sample averaging and smoothing Step 5: Gradient clipping and normalization like but .

[0069] For numerical stability, it is necessary to first... Normalize by dividing each component by the range of the corresponding variable.

[0070] Step 6: Projected gradient update In the formula, For feasible region projection operators.

[0071] Decentralized multi-agent reinforcement learning (MARL) includes: The node is set as an intelligent agent, with the state being... The action is The reward is .

[0072] To induce cooperative behavior, the original reward is enhanced with neighborhood incentives, and the locally shaped reward is defined as: In the formula, Let i be the communication neighborhood of node i; This is a cooperation strength coefficient, used to adjust the trade-off between individual and group interests. The larger the value, the stronger the incentive for a node to make concessions for the overall value of its neighborhood; The smaller the node, the more it focuses on its own value. In this embodiment, This range is applicable to most industrial collaboration scenarios. The lower end of the range is used for scenarios with weak coupling or where security is the priority, while the higher end is used for scenarios where cooperation needs to be clearly encouraged.

[0073] Decentralized Actor-Critic is adopted: each node maintains local policy network Actor parameters. and Critic parameters The strategy is updated using experience replay and neighborhood information.

[0074] To address the non-stationarity of multi-agent environments caused by simultaneous policy updates, a fingerprinting mechanism, V-trace importance weight correction, and priority experience replay are employed, including: Each transition in the experience replay is appended with fingerprint information, for example: In the formula, k is the number of training steps; This is the normalization constant; This represents the current exploration rate.

[0075] During training As an additional input to the Actor and Critic networks, it explicitly models the policy evolution process and compensates for distribution drift caused by changes in the policies of other agents.

[0076] For off-policy replay samples, V-trace type truncation importance sampling is used, with the expression as follows: In the formula, For behavioral strategies; For strategic objectives; and This is the truncation threshold.

[0077] It effectively reduces gradient bias caused by strategy differences and improves sample utilization.

[0078] A sliding window is used to limit the size of the replay buffer; priority sampling based on sample age gives fresh data a higher sampling probability, which accelerates strategy convergence and adapts to dynamic environments.

[0079] Each node updates the Actor parameters based on local experience replay and neighborhood information. and Critic parameters .

[0080] The learning rate uses a decay strategy: In the formula, For the first The learning rate for each iteration; The initial learning rate can typically range from [value range missing]. ; Index for iteration count; This is the decay index.

[0081] Maximum inner iteration Determined by the real-time nature of the observed system, including: If real-time pressure is high, a single communication update will be used, i.e. This is used in conjunction with MPC rolling to compensate; In typical online collaborative optimization scenarios Desirable This allows for obtaining a better approximate consistent solution under limited communication and computing resources.

[0082] In offline simulations, parameter tuning, or when computational resources are sufficient, it is necessary to appropriately increase the [computing capacity / resources]. To ensure the convergence properties of the algorithm, at this time Pick .

[0083] Each iteration terminates when any of the following conditions are met: The change in the decision variable is less than the threshold: Both the original residual and the dual residual are below the preset tolerance.

[0084] In this embodiment, The range of values ​​is The smaller threshold corresponds to higher inner-layer convergence accuracy and is suitable for scenarios with sufficient computing resources or high consistency requirements; the larger threshold is suitable for online control or fast response scenarios with strict real-time requirements.

[0085] To handle non-convex or ill-conditioned optimization problems, a proximal regularization term is introduced, expressed as: In the formula, Represents a node The vector of decision variables to be solved in the current inner iteration; The decision variables from the previous iteration are assigned values. is the damping coefficient, used to constrain the change range of solutions between two adjacent iterations.

[0086] The above approach improves the numerical stability of the subproblem by suppressing excessively large update step sizes and helps to achieve smooth iterative convergence behavior under non-convex optimization conditions.

[0087] Within the MARL framework, at the node The strategy objective function adds an entropy incentive term to the original value function to avoid getting trapped in local suboptimal conditions, expressed as: In the formula: The entropy of the policy distribution; This is the entropy weighting coefficient, used to balance maximizing returns with exploratory behavior.

[0088] The Actor parameter update uses a policy gradient with an entropy term, expressed as: The final control update still uses a partial update strategy, expressed as: In the formula, For the gradient approximation of the local objective, entropy regularization only applies to the calculation process of the gradient or policy parameters and does not change the constraint form of the decision variables.

[0089] Define a local feasible set for constraint processing: The corresponding projection operator: .

[0090] In the distributed implementation of this invention, certain local information is defined as protected data objects and should be protected. Protected data objects include, but are not limited to: local sensing and measurement values, local system identification parameters, local value and cost functions, policy parameters and historical decisions, precise constraint parameters, and any node privacy-related data.

[0091] In the default implementation, all of the above objects are stored locally on the nodes and are not broadcast in plaintext. To achieve collaborative optimization, nodes only exchange necessary aggregates or intermediate quantities, such as low-dimensional summaries of neighborhood actions and encrypted intermediate variables. If more granular information must be shared, differential privacy and other protective measures can be used before transmission, and all network communication is conducted through encrypted channels. This design ensures the performance of collaborative optimization while protecting the commercially sensitive information and security privacy of the participating parties.

[0092] Differential Privacy (DP) Implementation Details When nodes need to share gradients / intermediates Protection is achieved through noise mechanisms: Calculation of the standard deviation of injected Gaussian noise based on the Gaussian mechanism .

[0093] In the formula: This is the amplification factor determined by the failure probability parameter, and its specific value is related to the required differential privacy failure probability. (Also known as the confidence violation probability) is related. According to the non-strict form of the standard Gaussian mechanism, taking... . and To replace the neighboring datasets in the neighborhood. It is a sensitivity estimate when there is a different sample in adjacent datasets, reflecting the global sensitivity. It is the overall privacy budget estimated using a combination of parallel and serial rules during multiple communications, and it is relatively small. It indicates stronger privacy protection but usually results in a greater loss of utility.

[0094] Implementation example for path three: Specific update rules for decentralized Actor-Critic In this embodiment, each decision node Maintain a pair of neural network parameters: Actor parameters (Policy Network) ) and Critic parameters (Value function / action value network) To ensure training stability and retain neighborhood interaction information during experience replay, replay entries are saved as follows: In the formula, This is a summary of neighborhood action information (which can be a neighborhood action vector, mean, or low-dimensional embedding). Record the probability of the behavior strategy (record if the behavior is random, so that necessary adjustments can be made in the future). For timestamps (used for priority / freshness sampling).

[0095] By sampling a small batch from experience replay and updating the parameters for each training iteration, the following is a specific implementable formula.

[0096] 1. Sampling Sample a batch from the playback buffer using a priority / sliding window strategy. Sample .

[0097] 2. Critic value function objective Use target network and target actor Constructing TD targets: In the formula, is the discount factor. Critic's least squares loss is... Update Critic using gradient descent: In the formula, is the learning rate for Critic.

[0098] 3. Actor Strategy Update Adopt a deterministic strategy To simplify implementation. In the deterministic case, the Actor objective is to maximize the Critic estimate: Its parameters are updated using chain rule differentiation: In the formula, The learning rate is the Actor's rate.

[0099] 4. Target network soft update To avoid training oscillations, improve numerical stability, and maintain the target network And perform soft updates periodically: In the formula, It is a soft update coefficient (typical) ).

[0100] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0101] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A self-organizing collaborative optimization method for industrial value based on dynamic game theory, characterized in that, include: Complex industrial systems are abstracted into a three-layer hierarchical heterogeneous network model; each node includes an inherent attribute vector. State vector Environment vectors and decision variables ; Construct the state variables of the nodes and set the dynamic evolution rules for the node state variables; Construct a value function for each node in the model to calculate the instantaneous revenue of each node; Construct a rolling time-domain multi-stage dynamic game model, so that the optimization objective of the nodes in the model is to maximize the cumulative discounted total revenue; Obtain the information set of the model nodes and update the value weight of the information set online; Configure privacy-preserving communication between nodes; Local optimization is performed based on the local information set of the node to generate the node's decision variables; At the start of each control cycle, the node adaptively selects the solution path based on the industrial scenario.

2. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The state variables of the constructed node include: The observable dynamic characteristics of a node use state variables Representation is performed based on the state vector. Mapped to obtain; If the state variable Unable to be measured directly, it is estimated through sensor observation sequences, and the expression is: In the formula, Let be the sensor observation value of node i at time t; It is a nonlinear observation function; To observe noise; Let i be the state variable of node i at time t; The state estimate is generated by the node's local state estimation algorithm, and the expression is: In the formula, This is the state estimate of node i; This is a state estimation algorithm; For historical observation; These are historical decision variables.

3. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The dynamic evolution rules of the node state variables include: For a linear dynamic system, the expression for the node state transition is: In the formula, for; This is the state transition matrix; For input control matrix; For disturbance terms; For a nonlinear dynamic system, the expression for the node state transition is: In the formula, For disturbance terms; It is a nonlinear state transition function.

4. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The value function includes: The instantaneous value function of node i at time t is expressed as: In the formula, Let i be the value function of node i; , These are the state variables and decision variables for all other nodes, respectively; It is its own value function; As an interactive value item; Let the decision cost function be used. As its own value weight; Weighting based on interaction value; For nodes The neighborhood set; The self-value function is expressed as: In the formula, This is the state value weight vector; To control the payoff matrix; The linear marginal benefit vector for the control action; The interactive value item, expressed as: In the formula, This is the state-state coupling matrix; For control-control coupling matrix; This is the state-control coupling matrix; The decision cost function is expressed as follows: In the formula, This is the secondary penalty coefficient; The linear cost vector for control.

5. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The rolling time-domain multi-stage dynamic game model includes: The optimization objective of node i is to maximize the cumulative discounted total revenue, expressed as: In the formula, This represents the total cumulative discount revenue; Discount factor; Value function; The periodic step size; This is the value decay time constant.

6. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The information set includes: Node i observes the information set at time t Including: state variables Self-value weight Interaction value weight Sharing intermediate variables with its neighbors, the expression is: In the formula, This is a state estimate; As its own value weight; Interaction value weight; Let be the dual variable between node i and its neighbor j; For interactive value basis functions Decision variables for node i The partial derivatives; It is a neighborhood set.

7. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The value weights of the online updated information set include: Self-value weight and interaction value weight The expression for online updating via exponentially weighted moving average is: In the formula, Forgetting factor; This is an instantaneous self-value weight estimate; For instantaneous interaction weights.

8. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The privacy-preserving communication includes: node Only to nodes Send the local gradient intermediates and dual variables, expressed as follows: In the formula, The intermediate quantity of the interaction gradient provided by node j to node i. Let be the dual variable between node i and its neighbor j; and The estimated state values ​​of nodes i and j at time t; and The control decisions made by nodes i and j at the previous time step; This is the penalty parameter.

9. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, The step of performing local optimization based on the node's local information set to generate the node's decision variables includes: node Local information set Construct a local approximate objective function, expressed as: In the formula, Let be the local approximate utility function for node i; Node i calculates the first derivative of the local approximate objective function with respect to its own control, expressed as: In the formula: The basis function of its own value with respect to decision variables The gradient; For interactive gradients; Unconstrained policy gradient ascent update is performed based on the first derivative, expressed as follows: In the formula, For the updated temporary decision variables; The learning rate; The temporary decision variables are projected onto the feasible region to obtain the final decision variables, expressed as follows: In the formula, Let i be the locally feasible control set of node i; This is the Euclidean projection operator.

10. The industrial value self-organizing collaborative optimization method based on dynamic game theory according to claim 1, characterized in that, At the beginning of each control cycle, the node adaptively selects a solution path based on the industrial scenario, including: If the model is available and the minimum eigenvalue of the Hessian matrix of the local optimization problem is not less than the preset non-negative tolerance threshold, then the potential game and alternating direction multiplier method (ADMM) are used to solve the problem. If there are hard dynamic constraints or a strong forward prediction capability is required, then distributed MPC is used to solve the problem. If the model is unknown, the target value is obtained through simulation. If the minimum eigenvalue of the Hessian matrix in the local optimization problem is less than the preset non-negative tolerance threshold, then distributed gradient estimation or decentralized multi-agent reinforcement learning (MARL) is used to solve the problem.