Networking control method and system under strong constraint

By combining a two-stage strategy design with a neural network optimization solution layer, the problem of strong constraint feasibility under high-dimensional cross-node coupling conditions is solved, enabling efficient training and multi-scenario applicability of large-scale networked control, and providing stable resource scheduling and control decisions.

CN121806501APending Publication Date: 2026-04-07SHANGHAI JIAOTONG UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to guarantee strong constraint feasibility under high-dimensional, cross-node coupling conditions, resulting in high training costs, low exploration efficiency, and insufficient cross-scenario adaptability, making it difficult to achieve efficient training and online deployment for large-scale networked control.

Method used

A two-stage strategy design is adopted, combining a neural network model with an optimization solution layer. Through online control and offline training stages, Lagrange iteration algorithm and focus loss function are used to explicitly apply constraints, provide stable update signals and unified mathematical modeling, and realize networked control applicable to multiple scenarios.

Benefits of technology

It achieves stable satisfaction of strong constraints in large-scale network systems, improves training efficiency and adaptability, and ensures the high efficiency, stability and applicability of the control strategy in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121806501A_ABST
    Figure CN121806501A_ABST
Patent Text Reader

Abstract

The invention provides a networked control method and system under strong constraint, and the method comprises the steps: S1, providing a dual-stage strategy framework combining a neural network model and a constraint optimization problem, and guaranteeing the satisfaction of a strong constraint condition; s2, proposing an expert decision obtaining method, and supporting the balance of strategy performance and calculation complexity; s3, providing an iterative weight truth value acquisition method which has the characteristics of high efficiency, stability, parameter insensitivity and the like; and S4, proposing a decision-focused imitation learning loss and training algorithm, and realizing efficient and extensible training. The control method disclosed by the invention can be expanded to a large-scale actual networked system problem, strong constraint satisfaction of a decision is realized by combining constraint optimization problem solution, and efficient and low-cost training is realized by adopting an imitation learning algorithm; and landing and popularization in various scenes and various types of practical applications such as supply chain inventory management, dynamic vehicle scheduling, satellite data transmission and the like are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of networked control and optimization, specifically to a networked control method and system under strong constraints, and more particularly to a networked control method under strong constraints that is adaptable to large-scale, efficient training and applicable to multiple scenarios. Background Technology

[0002] Many real-world systems can be abstracted as networked structures composed of large-scale nodes and edges. Nodes can generate, randomly emerge, store, or consume entities, while edges handle the transmission of entities; consumption brings revenue, while production, storage, and transmission incur costs; nodes and edges are often constrained by capacity, rate, inventory, bandwidth, etc. The core problem is to jointly decide on "production-storage-transmission-consumption" to maximize long-term profits while satisfying system constraints. This problem can be uniformly formulated as a strongly constrained networked control problem, essentially a capacity-constrained dynamic optimization problem on a graph, characterized by multi-scale spatiotemporal coupling, strong constraints, randomness, and a high-dimensional continuous action space. Existing work can be divided into the following two categories of techniques.

[0003] Pure neural network methods learn the distribution of possible actions through diffusion models or deep generative models, relying on feasible samples in offline data. During the testing phase, reward signals guide the generation of actions to optimize returns. These methods suffer from difficulties in guaranteeing constraint feasibility under high-dimensional, cross-node coupling constraints. Because multi-node decisions influence each other, coupling constraints such as capacity and flow conservation require global consistency; relying solely on neural network approximations can easily lead to actions that violate these constraints during the inference phase.

[0004] Another approach uses a neural network as a feature extraction or parameterization module, solving a constrained optimization problem on its output to obtain actionable actions, and then training the neural network module using a reinforcement learning algorithm. This approach is effective in small-scale systems, but suffers from poor scalability. When the action dimension increases to hundreds or even higher, the efficiency of reinforcement learning exploration decreases significantly, policy convergence is slow or it gets stuck in suboptimal conditions, and its actual performance is inferior to simple heuristic methods.

[0005] In summary, existing technologies have bottlenecks in the following aspects: strong constraint feasibility is difficult to guarantee stably under high-dimensional and cross-node coupling conditions; training costs are high and exploration efficiency is low in large-scale networks, making it difficult to achieve efficient training and online deployment; cross-scenario adaptability is insufficient, and performance fluctuates greatly when facing different graph sizes, constraint forms and demand statistical characteristics.

[0006] Therefore, this invention aims to design a strongly constrained networked control method that is adaptable to large-scale, efficient training and applicable to multiple scenarios, and to provide a strongly constrained control solution for large-scale networked systems, supporting its implementation and promotion in various practical applications such as supply chain inventory management, dynamic vehicle scheduling, and satellite data transmission.

[0007] Patent application CN113255876A discloses an optimization method, apparatus, and application method / appliance for deep learning neural networks. The optimization method includes: constructing a recognition and classification neural network framework; constructing an optimization decision neural network framework; inputting image / video data into the recognition and classification neural network for training to obtain the network parameters of the trained recognition and classification neural network; and inputting image / video data of known categories into the optimization decision neural network for training to obtain the optimization parameters for each category of the trained optimization decision neural network. However, this patent cannot completely solve the existing technical problems, nor can it meet the needs of this invention. Summary of the Invention

[0008] In view of the shortcomings of the prior art, the purpose of this invention is to provide a networked control method and system under strong constraints.

[0009] The strongly constrained networked control method provided by the present invention is applied to resource scheduling and control of a physical network system containing multiple nodes and edges. The physical network system includes at least one of a supply chain logistics network, a transportation network, or a satellite communication network, wherein nodes represent resource production, storage, or consumption sites, edges represent resource transmission paths, and system operation is constrained by the capacity constraints of nodes and edges. The method includes an online control phase and an offline training phase; In the online control phase, a control strategy model consisting of a neural network model and an optimization solution layer connected sequentially is executed. The neural network model receives the current state of each node in the physical network system, which includes resource inventory, in-transit resources, and historical demand information, and outputs intermediate layer weights. The optimization solution layer receives the intermediate layer weights and the node capacity constraints, edge capacity constraints, and traffic allocation constraints of the physical network system. By solving the first constraint optimization problem, it outputs resource scheduling and control decisions that satisfy all constraints. These decisions are used to directly control the resource production, storage, and transmission allocation on the edges of each node in the physical network system. In the offline training phase, the parameters of the neural network model are updated through a training algorithm. This algorithm includes, sequentially, an expert decision acquisition submodule, a weight truth value acquisition submodule, and a loss calculation and backpropagation submodule. The expert decision acquisition submodule solves a second-constraint optimization problem based on the completed trajectory data of the physical network system to obtain expert decisions. The weight truth value acquisition submodule takes the expert decisions and the intermediate layer weights output by the neural network model as input, iteratively solves a third-constraint optimization problem, updates internal variables, and outputs the truth values ​​of the intermediate layer weights. The loss calculation and backpropagation submodule calculates a differentiable loss function based on the expert decisions, the truth values ​​of the intermediate layer weights, and the intermediate layer weights output by the neural network model, and backpropagates the gradient of the loss function with respect to the neural network model parameters to update the model. The current state includes the number of entities stored by the node itself, the number of entities in production, and the historical cumulative demand; the control decision is a vectorized representation of the proportion of entity transmission allocated to each node on each available path, used to determine the number of entities produced by the node and the number of entities transmitted on the path.

[0010] Preferably, the mathematical form of the first constrained optimization problem is:

[0011]

[0012]

[0013]

[0014]

[0015] in, Variables to be decided The vector representation of, Represents the node Arrival Node path The proportion of entity transmission allocated above, It is the index of a certain edge in the network. It is a node At any moment The size of the entity carried at that time These are the intermediate layer weights, and T is the transpose sign. For a set of nodes, for Time node To the node The number of transmissions required, For nodes Arrival Node exist The set of available paths at that time They are the edges and nodes Capacity limitations Time node The set of edges available within the system at that time.

[0016] Preferably, the mathematical form of the second constrained optimization problem is:

[0017]

[0018]

[0019]

[0020]

[0021]

[0022] in, For time slices The instantaneous benefits generated within; From to The index of a specific time slice within the time range is used to describe the state of the system at that specific time slice; H represents the number of future time steps. For time slices Inside, from the node Arrival Node The set of paths; For time slices Inside, at the node Arrival Node path The transmission ratio allocated above; For time slices Inside, node To the node The transmission requirements; For time slices Inside, path Maximum capacity limit; For time slices The set of usable edges inside; For time slices internal nodes The amount of physical storage; For time slices internal nodes The amount of physical storage; For time slices The total amount of products produced; The time required for product manufacturing; For time slices internal nodes The quantity of demand satisfied; For time slices internal nodes The amount of physical storage.

[0023] Preferably, the process of the weight truth value acquisition submodule executing the Lagrange iteration algorithm includes: performing... Round iteration, in the 1st round In each iteration, based on the current iteration value and Lagrange multipliers Solving the third constraint optimization problem yields the variables. Update the variables according to the following formula: After completing K iterations, output As the true value of the intermediate layer weights; The mathematical form of the third constraint optimization problem is:

[0024]

[0025]

[0026]

[0027]

[0028] in, The learning rate parameter, For expert decision-making, and For the updated iterative values ​​and Lagrange multipliers.

[0029] Preferably, the mathematical form of the loss function is:

[0030] in, Representative moment The constraint space is limited by the constraints. For loss function The first term of the optimization problem is the decision variable. For iteration The subsequent output The value, To use expert decision-making The instant benefits it brings.

[0031] The strongly constrained networked control system provided by the present invention is applied to resource scheduling and control of a physical network system containing multiple nodes and edges. The physical network system includes at least one of a supply chain logistics network, a transportation network, or a satellite communication network, wherein nodes represent resource production, storage, or consumption sites, edges represent resource transmission paths, and system operation is constrained by the capacity constraints of nodes and edges. The system includes an online control phase and an offline training phase; In the online control phase, a control strategy model consisting of a neural network model and an optimization solution layer connected sequentially is executed. The neural network model receives the current state of each node in the physical network system, which includes resource inventory, in-transit resources, and historical demand information, and outputs intermediate layer weights. The optimization solution layer receives the intermediate layer weights and the node capacity constraints, edge capacity constraints, and traffic allocation constraints of the physical network system. By solving the first constraint optimization problem, it outputs resource scheduling and control decisions that satisfy all constraints. These decisions are used to directly control the resource production, storage, and transmission allocation on the edges of each node in the physical network system. In the offline training phase, the parameters of the neural network model are updated through a training algorithm. This algorithm includes, sequentially, an expert decision acquisition submodule, a weight truth value acquisition submodule, and a loss calculation and backpropagation submodule. The expert decision acquisition submodule solves a second-constraint optimization problem based on the completed trajectory data of the physical network system to obtain expert decisions. The weight truth value acquisition submodule takes the expert decisions and the intermediate layer weights output by the neural network model as input, iteratively solves a third-constraint optimization problem, updates internal variables, and outputs the truth values ​​of the intermediate layer weights. The loss calculation and backpropagation submodule calculates a differentiable loss function based on the expert decisions, the truth values ​​of the intermediate layer weights, and the intermediate layer weights output by the neural network model, and backpropagates the gradient of the loss function with respect to the neural network model parameters to update the model. The current state includes the number of entities stored by the node itself, the number of entities in production, and the historical cumulative demand; the control decision is a vectorized representation of the proportion of entity transmission allocated to each node on each available path, used to determine the number of entities produced by the node and the number of entities transmitted on the path.

[0032] Preferably, the mathematical form of the first constrained optimization problem is:

[0033]

[0034]

[0035]

[0036]

[0037] in, Variables to be decided The vector representation of, Represents the node Arrival Node path The proportion of entity transmission allocated above, It is the index of a certain edge in the network. It is a node At any moment The size of the entity carried at that time These are the intermediate layer weights, and T is the transpose sign. For a set of nodes, for Time node To the node The number of transmissions required, For nodes Arrival Node exist The set of available paths at that time They are the edges and nodes Capacity limitations Time node The set of edges available within the system at that time.

[0038] Preferably, the mathematical form of the second constrained optimization problem is:

[0039]

[0040]

[0041]

[0042]

[0043]

[0044] in, For time slices The instantaneous benefits generated within; From to The index of a specific time slice within the time range is used to describe the state of the system at that specific time slice; H represents the number of future time steps. For time slices Inside, from the node Arrival Node The set of paths; For time slices Inside, at the node Arrival Node path The transmission ratio allocated above; For time slices Inside, node To the node The transmission requirements; For time slices Inside, path Maximum capacity limit; For time slices The set of usable edges inside; For time slices internal nodes The amount of physical storage; For time slices internal nodes The amount of physical storage; For time slices The total amount of products produced; The time required for product manufacturing; For time slices internal nodes The quantity of demand satisfied; For time slices internal nodes The amount of physical storage.

[0045] Preferably, the process of the weight truth value acquisition submodule executing the Lagrange iteration algorithm includes: performing... Round iteration, in the 1st round In each iteration, based on the current iteration value and Lagrange multipliers Solving the third constraint optimization problem yields the variables. Update the variables according to the following formula: After completing K iterations, output As the true value of the intermediate layer weights; The mathematical form of the third constraint optimization problem is:

[0046]

[0047]

[0048]

[0049]

[0050] in, The learning rate parameter, For expert decision-making, and For the updated iterative values ​​and Lagrange multipliers.

[0051] Preferably, the mathematical form of the loss function is:

[0052] in, Representative moment The constraint space is limited by the constraints. For loss function The first term of the optimization problem is the decision variable. For iteration The subsequent output The value, To use expert decision-making The instant benefits it brings.

[0053] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention solves the technical problem that existing solutions cannot meet the constraints by using a two-stage strategy design and explicitly applying constraints, thus achieving a strong guarantee of constraint satisfaction. (2) This invention provides a stable and unbiased update signal by designing a loss function that focuses on the merits of decisions, and focuses on the merits of decisions rather than the merits of intermediate output weight results. This solves the technical problems of unstable training and frequent noise interference in existing schemes, and realizes efficient and stable strategy training that is suitable for large-scale applications. (3) By unifying mathematical modeling, this invention deliberately introduces fewer assumptions and restrictions, solves the technical problem of limited applicability of existing solutions, and realizes networked control applicable to multiple scenarios. Attached Figure Description

[0054] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart of a networked control method under strong constraints in an embodiment of the present invention. Detailed Implementation

[0055] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.

[0056] Example 1 This invention provides a strongly constrained networked control method that is adaptable to large-scale, efficient training and applicable to multiple scenarios, such as... Figure 1 As shown, it includes: Step S1: Online Control Strategy Model , by neural network model With optimization solution layer Together they form; Step S2: The offline policy model training algorithm consists of an expert decision acquisition submodule, a Lagrange iteration submodule, and a loss function calculation and backpropagation submodule.

[0057] Step S1 includes: Step S1.1: During the online networked control process, at each decision-making time node... Each node, based on its own state This includes the number of entities stored within the system, the number of entities in production, and the historical cumulative demand. These are input into the neural network model to obtain the intermediate layer weights. The neural network model consists of one RNN layer that processes historical cumulative demand data and three MLP layers that process other state data. The activation function is the ReLU function.

[0058] Step S1.2: Based on the obtained intermediate layer weights The system considers constraints such as node storage, link capacity, and the effectiveness of allocation ratios. It then solves for the decision on the number of nodes produced and the number of entities transmitted along each path by minimizing linear weighted averages. Specifically, let... Time node The set of edges available within the system at that time. For a set of nodes, For nodes Arrival Node exist The set of available paths at that time for Time node To the node The number of transmissions required, They are the edges and nodes Capacity limitations These are the decision variables, representing the variables at the nodes. Arrival Node path The entity transmission ratio allocated above, and marked with symbols A unified vectorization substitution is performed. The specific constrained optimization problem solved in this step is shown below:

[0059]

[0060]

[0061]

[0062]

[0063] in, Variables to be decided The vector representation of represents, in practice, the proportion of entity transmission allocated to each node along each path. For example, in supply chain inventory management applications, This represents the proportion of goods that are transferred from the factory to various warehouses along each path. In practical terms, it represents the index of a specific edge in the network. For example, in supply chain inventory management applications, This represents an index that indicates a connection path between a factory and a warehouse. The actual meaning is a node At any moment The size of the entity carried at that time. For example, in supply chain inventory management applications, This represents the physical inventory size of each warehouse and each store.

[0064] Step S2 includes: Step S2.1: The expert decision acquisition submodule aims to acquire expert decisions to provide accurate and unbiased update signals. In the offline phase, after the strategy and the specific application system have completed a full episode, all entity consumption demands within that episode have been met. Based on these realized state transitions and demand information, an offline constrained optimization problem with the goal of maximizing long-term profit can be constructed and solved using the realized trajectory data. This constrained optimization problem explicitly incorporates the costs and benefits of production, storage, and transportation, as well as coupling conditions such as node and edge capacity constraints and inventory conservation constraints, to ensure the feasibility and economy of the solution. The transportation allocation ratio and production quantity obtained by solving this problem are considered expert decisions. In other words, expert behavior is generated by offline constrained optimization based on information realized within the episode; its essence is a decision scheme obtained by consistently optimizing long-term profit under strong constraints. Specifically, let... For time period Instantaneous gains To balance optimality and solution complexity, the time domain considered includes four main parts: production, consumption, transmission to outgoing nodes, and transmission to incoming nodes. Based on the above, the specific constrained optimization problem solved in this step is denoted as (OHO), as shown below:

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] in, The actual meaning is time slice The instantaneous benefits generated within. For example, in supply chain inventory management applications, This represents a time slice. The total revenue generated from customers purchasing physical products. From to An index for a specific time slice within a time range, used to describe the state of the system at that specific time slice. The actual meaning is time slice Inside, from the node Arrival Node The set of all possible paths. For example, in supply chain inventory management applications, the feasible paths from the factory to the warehouse in different time slices may change due to events such as traffic accidents or traffic congestion. Therefore, we use... Provide a rigorous description. The actual meaning is time slice Inside, at the node Arrival Node path The allocated transfer ratio. For example, in supply chain inventory management applications, different physical transfer ratios are allocated from a factory to a warehouse along different paths. The actual meaning is time slice Inside, node To the node The transmission demand. For example, in supply chain inventory management applications, the total amount of physical products produced by a factory represents the transmission demand. The actual meaning is time slice Inside, path Maximum capacity limitations. For example, in satellite data transmission applications, inter-satellite transmission links and satellite-to-ground transmission links each have their own maximum bandwidth limitations, which restrict the maximum amount of data that can pass through the link at the same time. The actual meaning is time slice The set of available edges. For example, in satellite data transmission applications, due to satellite flybys, interference from third-party satellites, electromagnetic influences, etc., the set of available links in different time slices is time-varying. Therefore, we use a set... Provide a rigorous description. The actual meaning is time slice internal nodes The physical inventory level. For example, in supply chain inventory management applications, the physical inventory level of a factory changes after production and transportation. Therefore, we adopt... Provide a rigorous description. The actual meaning is time slice internal nodes The physical storage volume. Its sum Together, they describe the changes in the storage volume of physical products. The actual meaning is time slice The total amount of products produced. For example, in supply chain inventory management applications, factories in time slices... Select to proceed The production of a product, in time slices Production completed. This refers to the time required for product production. For example, in supply chain inventory management applications, factories allocate time slices... Select to proceed The production of a product, in time slices Once production is complete, it will be added to the total product quantity, i.e., added to... middle. The actual meaning is time slice internal nodes The quantity of demand satisfied. For example, in supply chain inventory management applications, user purchase demands arise randomly, and the portion that is satisfied represents the quantity of demand met. . The actual meaning is time slice internal nodes The physical storage volume. Its sum Together, they describe the changes in the storage volume of physical products.

[0071] Step S2.2: The Lagrange iterative algorithm submodule aims to obtain the necessary inputs for the loss function of focused decision-making, namely the intermediate layer weights. truth value Specifically, the Lagrange iteration algorithm submodule performs... The next Lagrange iteration. In each of these iteration rounds... Based on the current iteration value Additional variables to consider With expert decision-making The difference between them is used as part of the objective to solve the problem. Then, gradient descent is used to iteratively update. and Lagrange multipliers Specifically, the iterative update process, denoted as (GDA), is... ,in Here is the learning rate parameter. This step involves solving a problem, denoted as (GetA), as shown below:

[0072]

[0073]

[0074]

[0075]

[0076] Step S2.3: The loss function calculation and feedback submodule aims to calculate the differentiable loss function for focused decision-making based on the expert decisions and intermediate weight truth values ​​obtained above, and then feed it back to the neural network policy model. An update will be performed. Specifically, the loss function used is... ,in Representative moment The loss is limited by the constraint space imposed by the constraints. This loss is differentiable and fully considers the existence of the constraint optimization problem in the second stage of the strategy, unlike previous solutions that completely ignored the constraint optimization problem in the second stage. Therefore, this loss design focuses on the final... Does the output closely reflect expert decision-making? The loss is not the loss of intermediate weights. Given that this loss is relative to the neural network policy model... The output is differentiable, so it can be used directly. As the loss function, gradient update signals are backpropagated to update the policy model, relying on the support of popular PyTorch or Tensorflow environments. .

[0077] in, For loss function The first term in the optimization problem represents the decision variable. For (GDA) iteration The subsequent output The value of . To use expert decision-making The resulting instantaneous benefits. Due to the lack of... Generate gradient signal, Value in loss function The middle part can be directly ignored.

[0078] Example 2 The present invention also provides a networked control system under strong constraints, which can be implemented by executing the process steps of the networked control method under strong constraints. That is, those skilled in the art can understand the networked control method under strong constraints as a preferred embodiment of the networked control system under strong constraints.

[0079] The system includes an online control phase and an offline training phase; In the online control phase, a control strategy model consisting of a neural network model and an optimization solution layer connected in sequence is executed: the neural network model receives the current state of each node in the system and outputs the intermediate layer weights; the optimization solution layer receives the intermediate layer weights and preset system constraints, solves the first constraint optimization problem, and outputs a control decision that satisfies all constraints. During the offline training phase, the parameters of the neural network model are updated through a training algorithm. This algorithm includes, sequentially, an expert decision acquisition submodule, a weight truth value acquisition submodule, and a loss calculation and backpropagation submodule. The expert decision acquisition submodule solves a second-constraint optimization problem based on completed system interaction trajectory data to obtain expert decisions. The weight truth value acquisition submodule takes the expert decisions and the intermediate layer weights output by the neural network model as input, iteratively solves a third-constraint optimization problem, updates internal variables, and outputs the truth values ​​of the intermediate layer weights. The loss calculation and backpropagation submodule calculates a differentiable loss function based on the expert decisions, the truth values ​​of the intermediate layer weights, and the intermediate layer weights output by the neural network model, and backpropagates the gradient of the loss function with respect to the neural network model parameters to update the model. The current state includes the number of entities stored by the node itself, the number of entities in production, and the historical cumulative demand; the control decision is a vectorized representation of the proportion of entity transmission allocated to each node on each available path, used to determine the number of entities produced by the node and the number of entities transmitted on the path.

[0080] The mathematical form of the first constrained optimization problem is:

[0081]

[0082]

[0083]

[0084]

[0085] in, Variables to be decided The vector representation of, Represents the node Arrival Node path The proportion of entity transmission allocated above, It is the index of a certain edge in the network. It is a node At any moment The size of the entity carried at that time These are the intermediate layer weights, and T is the transpose sign. For a set of nodes, for Time node To the node The number of transmissions required, For nodes Arrival Node exist The set of available paths at that time They are the edges and nodes Capacity limitations Time node The set of edges available within the system at that time.

[0086] The mathematical form of the second constrained optimization problem is:

[0087]

[0088]

[0089]

[0090]

[0091]

[0092] in, For time slices The instantaneous benefits generated within; From to The index of a specific time slice within the time range is used to describe the state of the system at that specific time slice; H represents the number of future time steps. For time slices Inside, from the node Arrival Node The set of paths; For time slices Inside, at the node Arrival Node path The transmission ratio allocated above; For time slices Inside, node To the node The transmission requirements; For time slices Inside, path Maximum capacity limit; For time slices The set of usable edges inside; For time slices internal nodes The amount of physical storage; For time slices internal nodes The amount of physical storage; For time slices The total amount of products produced; The time required for product manufacturing; For time slices internal nodes The quantity of demand satisfied; For time slices internal nodes The amount of physical storage.

[0093] The process of the weight truth value acquisition submodule executing the Lagrange iteration algorithm includes: performing... Round iteration, in the 1st round In each iteration, based on the current iteration value and Lagrange multipliers Solving the third constraint optimization problem yields the variables. Update the variables according to the following formula: After completing K iterations, output As the true value of the intermediate layer weights; The mathematical form of the third constraint optimization problem is:

[0094]

[0095]

[0096]

[0097]

[0098] in, The learning rate parameter, For expert decision-making, and For the updated iterative values ​​and Lagrange multipliers.

[0099] The mathematical form of the loss function is:

[0100] in, Representative moment The constraint space is limited by the constraints. For loss function The first term of the optimization problem is the decision variable. For iteration The subsequent output The value, To use expert decision-making The instant benefits it brings.

[0101] Those skilled in the art will understand that, in addition to implementing the system, apparatus, and their modules provided by this invention in purely computer-readable program code, the same program can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system, apparatus, and their modules provided by this invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; alternatively, modules for implementing various functions can be considered both software programs implementing the method and structures within the hardware component.

[0102] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A networked control method under strong constraints, characterized in that, Resource scheduling and control are applied to physical network systems containing multiple nodes and edges. The physical network system includes at least one of supply chain logistics networks, transportation networks, or satellite communication networks, wherein nodes represent resource production, storage, or consumption sites, edges represent resource transmission paths, and system operation is limited by the capacity constraints of nodes and edges. The method includes an online control phase and an offline training phase; In the online control phase, a control strategy model consisting of a neural network model and an optimization solution layer connected sequentially is executed. The neural network model receives the current state of each node in the physical network system, which includes resource inventory, in-transit resources, and historical demand information, and outputs intermediate layer weights. The optimization solution layer receives the intermediate layer weights and the node capacity constraints, edge capacity constraints, and traffic allocation constraints of the physical network system. By solving the first constraint optimization problem, it outputs resource scheduling and control decisions that satisfy all constraints. These decisions are used to directly control the resource production, storage, and transmission allocation on the edges of each node in the physical network system. In the offline training phase, the parameters of the neural network model are updated through a training algorithm. This algorithm includes, sequentially, an expert decision acquisition submodule, a weight truth value acquisition submodule, and a loss calculation and backpropagation submodule. The expert decision acquisition submodule solves a second-constraint optimization problem based on the completed trajectory data of the physical network system to obtain expert decisions. The weight truth value acquisition submodule takes the expert decisions and the intermediate layer weights output by the neural network model as input, iteratively solves a third-constraint optimization problem, updates internal variables, and outputs the truth values ​​of the intermediate layer weights. The loss calculation and backpropagation submodule calculates a differentiable loss function based on the expert decisions, the truth values ​​of the intermediate layer weights, and the intermediate layer weights output by the neural network model, and backpropagates the gradient of the loss function with respect to the neural network model parameters to update the model. The current state includes the number of entities stored by the node itself, the number of entities in production, and the historical cumulative demand; the control decision is a vectorized representation of the proportion of entity transmission allocated to each node on each available path, used to determine the number of entities produced by the node and the number of entities transmitted on the path.

2. The networked control method under strong constraints according to claim 1, characterized in that, The mathematical form of the first constrained optimization problem is: in, Variables to be decided The vector representation of, Represents the node Arrival Node path The proportion of entity transmission allocated above, It is the index of a certain edge in the network. It is a node At any moment The size of the entity carried at that time These are the intermediate layer weights, and T is the transpose sign. For a set of nodes, for Time node To the node The number of transmissions required, For nodes Arrival Node exist The set of available paths at that time They are the edges and nodes Capacity limitations Time node The set of edges available within the system at that time.

3. The networked control method under strong constraints according to claim 2, characterized in that, The mathematical form of the second constrained optimization problem is: in, For time slices The instantaneous benefits generated within; From to The index of a specific time slice within the time range is used to describe the state of the system at that specific time slice; H represents the number of future time steps. For time slices Inside, from the node Arrival Node The set of paths; For time slices Inside, at the node Arrival Node path The transmission ratio allocated above; For time slices Inside, node To the node The transmission requirements; For time slices Inside, path Maximum capacity limit; For time slices The set of usable edges inside; For time slices internal nodes The amount of physical storage; For time slices internal nodes The amount of physical storage; For time slices The total amount of products produced; The time required for product manufacturing; For time slices internal nodes The quantity of demand satisfied; For time slices internal nodes The amount of physical storage.

4. The networked control method under strong constraints according to claim 3, characterized in that, The process of the weight truth value acquisition submodule executing the Lagrange iteration algorithm includes: performing... Round iteration, in the 1st round In each iteration, based on the current iteration value and Lagrange multipliers Solving the third constraint optimization problem yields the variables. Update the variables according to the following formula: After completing K iterations, output As the true value of the intermediate layer weights; The mathematical form of the third constraint optimization problem is: in, The learning rate parameter, For expert decision-making, and For the updated iterative values ​​and Lagrange multipliers.

5. The networked control method under strong constraints according to claim 4, characterized in that, The mathematical form of the loss function is: in, Representative moment The constraint space is limited by the constraints. For loss function The first term of the optimization problem is the decision variable. For iteration The subsequent output The value, To use expert decision-making The instant benefits it brings.

6. A networked control system under strong constraints, characterized in that, Resource scheduling and control are applied to physical network systems containing multiple nodes and edges. The physical network system includes at least one of supply chain logistics networks, transportation networks, or satellite communication networks, wherein nodes represent resource production, storage, or consumption sites, edges represent resource transmission paths, and system operation is limited by the capacity constraints of nodes and edges. The system includes an online control phase and an offline training phase; In the online control phase, a control strategy model consisting of a neural network model and an optimization solution layer connected sequentially is executed. The neural network model receives the current state of each node in the physical network system, which includes resource inventory, in-transit resources, and historical demand information, and outputs intermediate layer weights. The optimization solution layer receives the intermediate layer weights and the node capacity constraints, edge capacity constraints, and traffic allocation constraints of the physical network system. By solving the first constraint optimization problem, it outputs resource scheduling and control decisions that satisfy all constraints. These decisions are used to directly control the resource production, storage, and transmission allocation on the edges of each node in the physical network system. In the offline training phase, the parameters of the neural network model are updated through a training algorithm. This algorithm includes, sequentially, an expert decision acquisition submodule, a weight truth value acquisition submodule, and a loss calculation and backpropagation submodule. The expert decision acquisition submodule solves a second-constraint optimization problem based on the completed trajectory data of the physical network system to obtain expert decisions. The weight truth value acquisition submodule takes the expert decisions and the intermediate layer weights output by the neural network model as input, iteratively solves a third-constraint optimization problem, updates internal variables, and outputs the truth values ​​of the intermediate layer weights. The loss calculation and backpropagation submodule calculates a differentiable loss function based on the expert decisions, the truth values ​​of the intermediate layer weights, and the intermediate layer weights output by the neural network model, and backpropagates the gradient of the loss function with respect to the neural network model parameters to update the model. The current state includes the number of entities stored by the node itself, the number of entities in production, and the historical cumulative demand; the control decision is a vectorized representation of the proportion of entity transmission allocated to each node on each available path, used to determine the number of entities produced by the node and the number of entities transmitted on the path.

7. The networked control system under strong constraints according to claim 6, characterized in that, The mathematical form of the first constrained optimization problem is: in, Variables to be decided The vector representation of, Represents the node Arrival Node path The proportion of entity transmission allocated above, It is the index of a certain edge in the network. It is a node At any moment The size of the entity carried at that time These are the intermediate layer weights, and T is the transpose sign. For a set of nodes, for Time node To the node The number of transmissions required, For nodes Arrival Node exist The set of available paths at that time They are the edges and nodes Capacity limitations Time node The set of edges available within the system at that time.

8. The networked control system under strong constraints according to claim 7, characterized in that, The mathematical form of the second constrained optimization problem is: in, For time slices The instantaneous benefits generated within; From to The index of a specific time slice within the time range is used to describe the state of the system at that specific time slice; H represents the number of future time steps. For time slices Inside, from the node Arrival Node The set of paths; For time slices Inside, at the node Arrival Node path The transmission ratio allocated above; For time slices Inside, node To the node The transmission requirements; For time slices Inside, path Maximum capacity limit; For time slices The set of usable edges inside; For time slices internal nodes The amount of physical storage; For time slices internal nodes The amount of physical storage; For time slices The total amount of products produced; The time required for product manufacturing; For time slices internal nodes The quantity of demand satisfied; For time slices internal nodes The amount of physical storage.

9. The networked control system under strong constraints according to claim 8, characterized in that, The process of the weight truth value acquisition submodule executing the Lagrange iteration algorithm includes: performing... Round iteration, in the 1st round In each iteration, based on the current iteration value and Lagrange multipliers Solving the third constraint optimization problem yields the variables. Update the variables according to the following formula: After completing K iterations, output As the true value of the intermediate layer weights; The mathematical form of the third constraint optimization problem is: in, The learning rate parameter, For expert decision-making, and For the updated iterative values ​​and Lagrange multipliers.

10. The networked control system under strong constraints according to claim 9, characterized in that, The mathematical form of the loss function is: in, Representative moment The constraint space is limited by the constraints. For loss function The first term of the optimization problem is the decision variable. For iteration The subsequent output The value, To use expert decision-making The instant benefits it brings.

Citation Information

Patent Citations

  • Deep learning neural network optimization method and device and application method and device

    CN113255876A