Industrial automation control method and device, equipment and storage medium
By performing dynamic topological modeling and distributed reinforcement learning training on multi-agent systems in industrial systems, combined with multi-objective game optimization and robust optimization, the problem of difficult multi-agent systems in the existing technology to take into account production efficiency, energy consumption and system reliability, and more stable and efficient automated collaborative control is achieved.
Patent Information
- Application Number
- CN202510191228.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When facing complex industrial scenarios and multi-dimensional uncertainties, existing multi-agent systems are difficult to take into account multi-target needs such as production efficiency, energy consumption and system reliability, and lack flexible topological modeling and collaborative game strategies.
By dynamic topological modeling of multi-agent systems in industrial systems, the adaptive clustering algorithm is used to decompose the interactive relationship of agents to obtain a local dynamic model. Then, based on the relationship diagram and local dynamic model between agents, a deep neural network is built for each agent, and the value network and state prediction function are obtained through distributed reinforcement learning algorithm training. Receive control instructions and current system status, use the state prediction function to generate future state predictions, build multi-objective game optimization problems and solve them, and obtain the Nash equilibrium solution and control priority order. Finally, a robust optimization problem is constructed for each agent and a robust control action is solved.
It achieves a balance between multi-target requirements, has high topological adaptation capabilities and robust control performance, and is suitable for complex multi-agent scenarios in industrial production, achieving more stable and efficient automated collaborative control.
Smart Images

Figure CN120029214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automation, and in particular to a control method, device, equipment and storage medium for industrial automation. Background Art
[0002] In the field of modern industrial automation, multi-agent systems (MAS) have gradually become an important technical solution for complex industrial scenarios due to their flexibility and scalability. However, existing multi-agent systems are often unable to cope with multi-dimensional uncertainties in real time when facing dynamic production environments, including equipment failures, production line changes, and multi-objective optimization requirements. Traditional control methods are usually based on centralized architectures or pre-defined topological structures, and rely on fixed model assumptions for the dynamic behavior of each agent, which makes it difficult to adapt to complex interactions and highly time-varying industrial production processes. In addition, as industrial systems continue to evolve towards intelligence and distribution, the need for collaborative control between agents has become increasingly prominent. It is difficult to achieve a global balance between various objectives (such as production efficiency, energy consumption, reliability, etc.) by relying solely on a single optimization algorithm or a simple game strategy. At the same time, the uncertainty in the environment also greatly increases the difficulty of robust control of the system. If there is a lack of consideration of uncertainty quantification and elastic control mechanisms, system bottlenecks or safety hazards are likely to occur in key links. Summary of the invention
[0003] The main purpose of this invention is to solve the technical problem that the existing multi-agent system lacks flexible topological modeling and collaborative game strategy when facing complex industrial scenarios and multi-dimensional uncertain factors, and it is difficult to take into account the multi-objective requirements such as production efficiency, energy consumption and system reliability; A first aspect of the present invention provides an industrial automation control method, the industrial automation control method comprising: Perform dynamic topological modeling on the multi-agent system in the industrial system to obtain a relationship diagram between agents, and use an adaptive clustering algorithm to decompose the multi-agent system based on the relationship diagram to obtain a local dynamic model of each agent; According to the relationship diagram between the agents and the local dynamic model, a deep neural network is constructed for each agent, and the deep neural network is trained by a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each agent; Receive control instructions and current system states, generate future state predictions using the state prediction function, construct and solve a multi-objective game optimization problem based on the prediction results, the control instructions, and the value network, and obtain Nash equilibrium solutions and control priority orders for each intelligent agent; According to the Nash equilibrium solution, control priority order and control instructions, a robust optimization problem considering uncertainty is constructed and solved for each intelligent agent, and each intelligent agent is controlled to perform robust control actions.
[0004] Optionally, in a first implementation of the first aspect of the present invention, the dynamic topological modeling of the multi-agent system in the industrial system is performed to obtain a relationship graph between the agents, and the multi-agent system is decomposed based on the relationship graph using an adaptive clustering algorithm to obtain a local dynamic model of each agent, including: Performing time-varying connection relationship analysis on the multi-agent system to obtain a time-varying connection matrix describing the connection relationship between agents, and constructing a relationship graph based on the time-varying matrix, wherein the nodes of the relationship graph represent agents and the edge weights represent the connection strengths; Applying an adaptive spectral clustering algorithm to calculate the Laplace matrix of the relationship graph, and performing clustering based on the eigenvectors corresponding to the Laplace matrix to obtain a clustering result of the intelligent agent; According to the clustering results, the multi-agent system is decomposed and allocated to different subsystems to obtain the subsystem division results. The subsystem division results are used to establish a local linear dynamic model for each agent based on the least squares method.
[0005] Optionally, in a second implementation of the first aspect of the present invention, decomposing the multi-agent system according to the clustering result, allocating it to different subsystems, obtaining subsystem division results, and using the subsystem division results to establish a local linear dynamic model for each agent based on the least squares method includes: Perform subsystem partitioning on the clustering results to obtain a preliminary subsystem partitioning scheme, in which the agents in the same cluster are assigned to the same subsystem; Identify boundary agents for the preliminary subsystem partitioning scheme to obtain a set of boundary agents, where boundary agents refer to agents that have direct connections with agents in other subsystems; The preliminary subsystem partitioning scheme is optimized according to the boundary agent set to obtain the subsystem partitioning result, and the local state space modeling of the agents in each subsystem is performed to obtain the state equation and output equation of each agent; The state equation and output equation of each intelligent agent are estimated by the least square method to obtain the local linear dynamic model of each intelligent agent.
[0006] Optionally, in a third implementation of the first aspect of the present invention, a deep neural network is constructed for each agent according to the relationship graph between the agents and the local dynamic model, and the deep neural network is trained by a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each agent, including: Based on the relationship graph between agents, determine the neighbor set of each agent, and build a simulation environment by combining the local dynamic model and the neighbor information, which can simulate the impact of the agent's actions on the state of itself and its neighbors; Constructing a dual deep Q network structure for each agent, and executing preset steps for each agent in a simulation environment to obtain a quadruple and store it in an experience replay buffer, wherein the dual deep Q network structure includes an online Q network and a target Q network; Randomly sampling batch data from the experience replay buffer, using the online Q network and the target Q network to calculate the current Q value and the target Q value based on the target Q network respectively; By minimizing the time difference error between the current Q value and the target Q value, the parameters of the online Q network are updated, and the parameters of the online Q network are soft-updated to the target Q network at every fixed number of steps to balance the stability and efficiency of learning; Return to the step of executing preset steps for each agent in the simulation environment until the Q network converges or reaches a preset number of training times, and use the trained online Q network as the value network; Based on the neighbor set, a graph convolutional neural network is constructed as the backbone network of the state prediction function, and combined with a long short-term memory network, the state prediction function is trained to predict future states.
[0007] Optionally, in a fourth implementation of the first aspect of the present invention, constructing a dual deep Q network structure for each agent, and executing preset steps for each agent in a simulation environment to obtain a quadruple and store it in an experience replay buffer comprises: Discretize the state space and action space of each agent to obtain a discrete state set and a discrete action set; A dual deep Q network structure is constructed based on a discrete state set and a discrete action set to obtain an online Q network and a target Q network. In the simulation environment, the epsilon-greedy strategy is used to select actions for each agent to obtain the current action, and based on the selected current action, the state transition is performed in the simulation environment to obtain the next state and immediate reward; The current state, current action, immediate reward, and next state are combined into a four-tuple and stored in the experience replay buffer.
[0008] Optionally, in a fifth implementation of the first aspect of the present invention, the receiving of control instructions and the current system state, generating a future state prediction using the state prediction function, constructing and solving a multi-objective game optimization problem based on the prediction result, the control instructions and the value network, and obtaining the Nash equilibrium solution and control priority order of each intelligent agent include: Receiving a control instruction and a current system state, and using a state prediction function to generate a state sequence in a prediction time domain based on the current system state as a prediction result; Based on the prediction results, control instructions and value network, a multi-objective optimization problem is constructed for each intelligent agent, and an improved non-dominated sorting genetic algorithm is used to solve the multi-objective optimization problem and obtain the Pareto optimal solution set; A multi-agent game problem is constructed based on the Pareto optimal solution set, and the Nash equilibrium is solved dynamically using the best response. The utility value of each agent is calculated according to the Nash equilibrium solution, and the control priority order is determined in descending order of utility value.
[0009] Optionally, in a sixth implementation of the first aspect of the present invention, constructing and solving a robust optimization problem considering uncertainty for each intelligent agent according to the Nash equilibrium solution, the control priority order, and the control instruction, and obtaining and controlling each intelligent agent to perform a robust control action includes: The Nash equilibrium solution and the control instruction are fused to obtain an initial control strategy and constraint conditions, and a robust optimization problem is constructed based on the initial control strategy and constraint conditions; The improved scenario tree method is used to model the system uncertainty of multi-agent systems and generate a limited number of typical scenario sets; The column generation algorithm is used to solve the robust optimization problem according to a set of typical scenarios to obtain a robust control strategy. Based on the obtained robust control strategy and control priority order, the control action sequence of the intelligent agent is generated and executed.
[0010] A second aspect of the present invention provides an industrial automation control device, the industrial automation control device comprising: A topological decomposition module is used to perform dynamic topological modeling on a multi-agent system in an industrial system, obtain a relationship diagram between agents, and decompose the multi-agent system based on the relationship diagram using an adaptive clustering algorithm to obtain a local dynamic model of each agent; A reinforcement training module is used to construct a deep neural network for each agent based on the relationship diagram between the agents and the local dynamic model, and train the deep neural network through a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each agent; A game optimization module is used to receive control instructions and current system states, generate future state predictions using the state prediction function, construct and solve a multi-objective game optimization problem based on the prediction results, the control instructions, and the value network, and obtain a Nash equilibrium solution and a control priority order for each intelligent agent; The robust control module is used to construct and solve a robust optimization problem taking uncertainty into consideration for each intelligent agent according to the Nash equilibrium solution, control priority order and control instructions, and obtain and control each intelligent agent to perform robust control actions.
[0011] A third aspect of the present invention provides an industrial automation control device, comprising: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected via lines; the at least one processor calls the instructions in the memory so that the industrial automation control device executes the steps of the above-mentioned industrial automation control method.
[0012] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the above-mentioned industrial automation control method.
[0013] The above-mentioned industrial automation control method, device, equipment and storage medium perform dynamic topological modeling on the multi-agent system and decompose the interaction relationship of each agent based on an adaptive clustering algorithm to obtain a local dynamic model. Distributed reinforcement learning is used to train the deep neural network of each agent, and the value network and state prediction function are output. After receiving the control instructions and the current system state, the future evolution is estimated by the state prediction function, and a multi-objective game optimization problem is constructed in combination with the value network to solve the Nash equilibrium and control priority of each agent. A robust optimization method is introduced for uncertain environments to generate control actions for each agent that can cope with environmental disturbances. The present invention achieves a balance between multi-objective requirements, has a high topological adaptability and robust control performance, is suitable for complex multi-agent scenarios in industrial production, and achieves more stable and efficient automated collaborative control.
[0014] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0015] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of a first embodiment of an industrial automation control method according to an embodiment of the present invention; Figure 2 A schematic diagram of an embodiment of an industrial automation control device in an embodiment of the present invention; Figure 3 It is a schematic diagram of an embodiment of an industrial automation control device in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0019] To facilitate understanding of this embodiment, a control method for industrial automation disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the method comprises the following steps: 101. Perform dynamic topological modeling on a multi-agent system in an industrial system to obtain a relationship diagram between agents, and use an adaptive clustering algorithm to decompose the multi-agent system based on the relationship diagram to obtain a local dynamic model of each agent; In one embodiment of the present invention, the dynamic topological modeling of the multi-agent system in the industrial system is performed to obtain a relationship graph between agents, and the multi-agent system is decomposed based on the relationship graph using an adaptive clustering algorithm to obtain a local dynamic model of each agent, including: performing a time-varying connection relationship analysis on the multi-agent system to obtain a time-varying connection matrix that describes the connection relationship between agents, and constructing a relationship graph based on the time-varying matrix, wherein the nodes of the relationship graph represent agents and the edge weights represent the connection strength; applying an adaptive spectral clustering algorithm to calculate the Laplace matrix of the relationship graph, and clustering based on the eigenvectors corresponding to the Laplace matrix to obtain clustering results of the agents; decomposing the multi-agent system according to the clustering results and assigning it to different subsystems to obtain subsystem partitioning results, and using the subsystem partitioning results to establish a local linear dynamic model for each agent based on the least squares method.
[0020] Specifically, the time-varying connection relationship of the multi-agent system is analyzed to obtain a time-varying connection matrix describing the connection relationship between the agents, and a relationship graph is constructed based on the time-varying matrix, wherein the nodes of the relationship graph represent the agents and the edge weights represent the connection strengths; in the implementation process, it is first necessary to record and organize the intersection data of each agent at the same sampling time, and then assign a value to each entry in the matrix based on the neighborhood intersection frequency or coupling strength and other weighted indicators. This value is recorded as , indicating the The agent and An agent at time In order to ensure the dynamic nature of the data, it is necessary to integrate the multi-time information and finally form a matrix ,In this matrix, if the coupling between a pair of agents is strong, then Get a relatively large value. If there is almost no interaction, the corresponding entry is close to zero, and then As a weighted adjacency form of the relationship graph, with vertices Represents the agent, using the edge weight Indicates the connection strength and maintains the continuity of the weighted relationship during visualization, thereby realizing the monitoring and analysis of time-varying connection relationships; Applying the adaptive spectral clustering algorithm to calculate the Laplace matrix of the relationship graph, and clustering based on the eigenvectors corresponding to the Laplace matrix to obtain the clustering results of the intelligent agent; When executing this step, it is necessary to first Constructing the graph Laplacian matrix ,in is the degree matrix, whose diagonal elements For the The sum of the edge weights of the vertices needs to be calculated separately at different times in order to reflect the adaptive characteristics , and compare the spectral changes at adjacent moments. By performing eigendecomposition on the Laplace matrix, the eigenvectors corresponding to the first few smallest non-zero eigenvalues can be extracted, and their coordinates in high-dimensional space can be used to measure the implicit relationship between agents. Adaptive spectral clustering will divide these coordinates using K-means or similar methods, so that agents with similar interaction characteristics can be automatically divided into the same category. In order to improve the accuracy of segmentation, the number of clusters or thresholds can be dynamically adjusted during the clustering process and tested in combination with system performance indicators. According to the clustering results, the multi-agent system is decomposed and assigned to different subsystems to obtain the subsystem division results. The subsystem division results are used to establish a local linear dynamic model for each agent based on the least squares method. After clustering is completed, the agents in the corresponding category are regarded as a relatively independent subsystem, and local modeling is performed based on the data of the agents within the subsystem. By applying least squares regression to the input and output observation sequence of each agent, it is assumed that the state vector of the agent is recorded as , the control input is denoted as , the local model can satisfy ,in and Depend on The solution is that in order to reduce the recognition error, the inertia term or regularization factor can be set in the regression process so that the matrix and More robust, all agents within each subsystem obtain their own local linear models through similar methods and form a hierarchical coupling relationship in the overall structure of the system.
[0021] Furthermore, the multi-agent system is decomposed according to the clustering results, allocated to different subsystems, and subsystem partitioning results are obtained, and the subsystem partitioning results are used to establish a local linear dynamic model for each agent based on the least squares method, including: subsystem partitioning processing is performed on the clustering results to obtain a preliminary subsystem partitioning scheme, in which the agents in the same cluster are allocated to the same subsystem; boundary agent identification is performed on the preliminary subsystem partitioning scheme to obtain a boundary agent set, wherein the boundary agent refers to an agent that has a direct connection with the agents in other subsystems; the preliminary subsystem partitioning scheme is optimized according to the boundary agent set to obtain subsystem partitioning results, and local state space modeling is performed on the agents in each subsystem to obtain the state equation and output equation of each agent; the state equation and output equation of each agent are estimated by the least squares method to obtain the local linear dynamic model of each agent.
[0022] Specifically, the clustering results are processed for subsystem division to obtain a preliminary subsystem division plan, where agents in the same cluster are assigned to the same subsystem. In this process, it is necessary to first map the agents according to the label information output by spectral clustering, and classify the agents with the same label into the same subsystem. In order to enable each subsystem to have a relatively tight internal interaction relationship, it is necessary to retrieve and count the total edge weight within the same cluster during this process. If there is a high edge connection strength between agents belonging to the same class, it is considered that they have similar functions or dynamic behaviors, and thus they are merged into the same subsystem. To save the data structure, a hash map can be established, with the cluster label as the key value, and a corresponding set of agent indexes stored in the table for quick retrieval and access in subsequent steps. For some samples at the junction, if the feature vector distances between adjacent clusters are close, they can be default assigned to the cluster with a smaller label number to form a preliminary subsystem division plan, and the belonging of these agents will be further processed in subsequent steps.
[0023] Specifically, boundary agents are identified for the preliminary subsystem division plan to obtain a boundary agent set, where boundary agents refer to agents that have a direct connection relationship with agents in other subsystems. In the specific implementation, it is necessary to traverse all agents within each subsystem in the preliminary division plan and retrieve the connections between them and agents in external subsystems. If in the adjacency matrix there exists , and agent and agent belong to different subsystems, then is marked as "the boundary agent of subsystem p". At this time, a boundary agent list needs to be created in the data structure to record the information of for subsequent integration. If an agent has multiple cross-subsystem connections, only one record of the agent is kept in the boundary list, and the information of the target subsystems of the external connections is merged and recorded. To ensure the queryability of this list, a dictionary structure can be used, with the agent number as the key and the set of its external connected subsystems as the value. The preliminary subsystem division plan is optimized according to the boundary agent set to obtain the subsystem division result, and a local state space model is established for the agents in each subsystem to obtain the state equation and output equation of each agent. In this step, it is necessary to determine its final belonging according to the interaction strength between the boundary agent and different subsystems. If the total edge weight of a certain boundary agent with subsystem A is much greater than its total edge weight with subsystem B, it can be assigned to subsystem A to reduce the correlation between subsystems. After completing this optimization, the agents within each subsystem have a more compact connection structure. Then, a state space model is established within each subsystem using the existing data. Assume that the state of the th agent is denoted as , and the input is denoted as , the output is recorded as , can be written in and Represent process noise and measurement noise respectively. In order to maintain the scalability of the subsystem, after the state and input dimensions are determined, the data at each moment can be integrated into a training set in time sequence, and then all the included intelligent agent models are managed within each subsystem. The state equation and output equation of each intelligent agent are estimated by the least squares method to obtain the local linear dynamic model of each intelligent agent; when implementing, the state quantity needs to be obtained in time sequence. , Input And the output , and combine them into the regression equation with noise. The optimal estimate is obtained by The fitting priority for balancing the state equation and the output equation can be used to constrain the size of the parameters and improve robustness. Regularization terms can be added to the loss function. Lagrange multipliers are given to constrain their size, and then numerical methods such as QR decomposition or SVD are used to obtain the global optimal solution, thereby outputting the local linear dynamic model of each agent.
[0024] 102. According to the relationship diagram between the intelligent agents and the local dynamic model, a deep neural network is constructed for each intelligent agent, and the deep neural network is trained by a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each intelligent agent; In one embodiment of the present invention, the method of constructing a deep neural network for each agent according to the relationship graph between the agents and the local dynamic model, and training the deep neural network through a distributed reinforcement learning algorithm to obtain the value network and state prediction function of each agent includes: determining the neighbor set of each agent based on the relationship graph between the agents, and constructing a simulation environment in combination with the local dynamic model and the neighbor information, wherein the simulation environment can simulate the impact of the action of the agent on the state of itself and its neighbors; constructing a dual deep Q network structure for each agent, and executing preset steps for each agent in the simulation environment to obtain a quadruple and store it in the experience playback buffer, wherein the dual deep Q network structure includes an online Q network and a target Q network; Batch data is randomly sampled from the experience replay buffer, and the current Q value and the target Q value are respectively calculated based on the target Q network using the online Q network and the target Q network; the parameters of the online Q network are updated by minimizing the temporal difference error between the current Q value and the target Q value, and the parameters of the online Q network are soft-updated to the target Q network at every fixed number of steps to balance the stability and efficiency of learning; the step of executing preset steps for each agent in the simulation environment is returned to until the Q network converges or a preset number of training times is reached, and the trained online Q network is used as the value network; based on the neighbor set, a graph convolutional neural network is constructed as the backbone network of the state prediction function, and combined with a long short-term memory network, the state prediction function is trained to predict future states.
[0025] Specifically, based on the relationship graph between agents, the neighbor set of each agent is determined, and a simulation environment is constructed in combination with the local dynamic model and the neighbor information. The simulation environment can simulate the impact of the agent's actions on its own and neighbor states. When implementing this step, it is necessary to first retrieve the direct edge information of each agent in the relationship graph, and mark the non-zero weight of the edge as a neighbor relationship mark. After obtaining the neighbor list, it is necessary to combine the state equation and output equation in the local dynamic model to construct a set of interactive environment mechanisms to characterize the state evolution caused by the agent after taking action. The environment should simultaneously consider the input of the current agent, the linkage response of the neighboring agents, and the coupling effect applied by the neighboring agents to the target agent. In order to make the simulation results closer to the real process, the system noise and measurement noise identified in the local dynamic model can be processed into random disturbance signals, and the state update equation can be disturbed and superimposed when the environment is running. It is necessary to maintain the state vectors of all agents within the environment. , and synchronously update the partial state or output of the neighboring agent when executing the action, so that the new state combination can be read from the environment at the next moment. In order to improve the flexibility and reproducibility of the system, a unified data interface is needed to map the discrete actions of the agent to the input quantity in the environment. , while recording the return value or reward signal generated by the environment for subsequent training of the reinforcement learning process. The environment can be designed as a distributed structure so that multiple agents can interact in parallel. If some neighbor sets in the relationship graph change, the topology can be updated by calling the reset or reconstruction function of the environment to keep the neighbor connections of each agent consistent, and use the local dynamic model to calculate the new state evolution trajectory when the environment steps. A dual-depth Q network structure is constructed for each agent, and preset steps are performed on each agent in the simulation environment to obtain a quadruple and store it in the experience replay buffer, wherein the dual-depth Q network structure includes an online Q network and a target Q network; when constructing the network, two neural networks need to be designed separately. The online Q network is used to make policy decisions and value evaluations based on current parameters. The parameters of the target Q network are consistent with the online Q network but are updated at a lower frequency, which can provide a more stable reference when calculating the target Q value. For each agent, its state can be The network is then concatenated with neighbor information to form the network input vector, or the self-state and neighbor state are processed in parallel in the network structure, and features are extracted through multiple convolutional or fully connected layers, and finally the Q value under different discrete actions is predicted in the output layer. When executing preset steps in a simulated environment, the agent needs to select actions, update the environment state, and obtain reward values in turn, and then assemble the current state, selected action, reward value, and next moment state into The experience quadruple in the form of a time-varying number of steps is stored in the experience replay buffer. In order to enable the network to better learn cross-time features, several jump steps or multi-step reward mechanisms can be introduced to record state transitions in different time spans while accumulating rewards, so as to include rich behavioral trajectories in the training process. Once the memory weight of the buffer exceeds the set threshold, the oldest experience can be deleted according to a certain strategy, and the more representative trajectories in the recent period can be retained to ensure that the training samples always contain enough new state and new action combinations. Randomly sample batches of data from the experience replay buffer, and use the online Q network and the target Q network to calculate the current Q value and the target Q value based on the target Q network respectively; update the parameters of the online Q network by minimizing the temporal difference error between the current Q value and the target Q value, and soft-update the parameters of the online Q network to the target Q network at every fixed number of steps to balance the stability and efficiency of learning; when implementing this link, it is necessary to first randomly extract batches of quadruple. As training samples, for each sample, the online Q network calculates the current state With action Predicted , the target Q network has the next state Calculate the next best move And get the target Q value Then, according to the idea of dual DQN, the timing difference target value is set to ,in is the discount factor, which is used to balance the current return and the future return. The difference between the predicted current Q value and the target Q value is constructed into the mean squared error or the root mean squared error, and the sum is obtained to get the loss function , where are the parameters of the online Q network. Subsequently, the backpropagation algorithm is applied to for gradient update, so that the online Q network can better fit the optimal Q value. A soft update operation is performed every fixed number of steps, which can make the parameters of the target Q network gradually approach the parameters of the online Q network , for example, by using the method of to avoid excessive drift of the target Q network in the initial stage of training. Through this method, the learning efficiency of the online Q network and the stability of the target Q network are balanced. Repeating this sampling and update step can make the network continuously converge to a better value function representation.
[0026] Return to the step of executing the preset steps for each agent in the simulation environment until the Q network converges or reaches the preset number of training times, and use the trained online Q network as the value network; in this process, after each training, it is necessary to re-enter the environment, let the agent continue to interact and execute the preset actions from the current policy to accumulate new experience trajectories, and sample and update again after collecting enough quadruples. The environment can adopt a multi-process or distributed framework to concurrently run multiple agents at this stage to shorten the training time and improve data diversity. For each agent, it can locally record its reward, state change, and action distribution information to track the convergence trend at the global level. If the online Q networks of all agents reach the minimum value in the temporal difference error or the mean squared loss, and if the return is stable, it is regarded as the basic convergence of the Q network, otherwise, it can continue to iterate and periodically evaluate the performance of the policy. After reaching the preset number of training times, the parameters of the final online Q network can be solidified and used as the value network to quickly output the action value estimate during the execution stage. For the target Q network, its significance is to provide a relatively stable training benchmark. If online re-learning is required, the dual structure of the online Q network - target Q network can be retained to maintain a high adaptability when the environment changes. This step is repeated until the value network has good generalization performance, and the goal of distributed reinforcement learning can be achieved.
[0027] Based on the neighbor set, a graph convolutional neural network is constructed as the backbone network of the state prediction function, and combined with the long short-term memory network, the state prediction function is trained to predict future states; to achieve this part, it is necessary to regard the neighbor set of each agent in the relationship graph and its edge weights as graph structure data, and the graph convolutional neural network (GCN) can update the node representation of the agent according to the representation information of the neighbor nodes in each layer of convolution operation. During training, the agent state vector can be regarded as a time series input, and the GCN is asked to perform a convolution aggregation on the adjacency matrix and state distribution at each time slice, and then the obtained high-dimensional embedding vector is input into the long short-term memory network (LSTM) to capture dynamic dependencies across time. Suppose an agent is at time The state is recorded as , the state set of the neighboring agent is denoted as ,The update process of graph convolution can be written as ; in Indicates Neighbors after layer convolution The node feature vector of represents the adjacency coefficient normalized based on weights, is the activation function, is a trainable parameter matrix; then the GCN output sequences at several moments are stacked and passed into the LSTM unit, allowing it to use the gating mechanism to extract short-term and long-term dependencies, and finally predict the state of the next one or more time steps at the output. During the training process, the mean square error or other regression loss function can be used, in conjunction with the stochastic gradient descent or Adam optimizer, and the parameters are updated after each batch of forward propagation is completed. When the mean square error decreases steadily on the validation set, it can be regarded as the model has mastered the propagation mode and time law between neighbors, and can provide more accurate future state predictions for the multi-agent system and assist the decision-making process of each agent.
[0028] Furthermore, the method of constructing a dual deep Q network structure for each agent and executing preset steps on each agent in a simulation environment to obtain a four-tuple stored in an experience replay buffer includes: discretizing the state space and action space of each agent to obtain a discrete state set and a discrete action set; constructing a dual deep Q network structure based on the discrete state set and the discrete action set to obtain an online Q network and a target Q network; in a simulation environment, executing an epsilon-greedy strategy for action selection on each agent to obtain a current action, and executing state transfer in the simulation environment based on the selected current action to obtain a next state and an immediate reward; and forming a four-tuple of the current state, current action, immediate reward and next state into an experience replay buffer.
[0029] Specifically, the state space and action space of each agent are discretized to obtain a discrete state set and a discrete action set. When implementing this process, it is necessary to first analyze the value range of the continuous state variable, and select an appropriate resolution based on actual needs to divide the interval into several finite segments, so as to obtain several discrete points in each dimension, and regard their Cartesian product as a discrete state set. , and subdivide it into The sampling points on this dimension can be recorded as , thus obtaining the corresponding discrete state axis. The action space usually contains several executable instructions or settable discrete input values. It is necessary to map the continuous action range to a set of feasible discrete action sets. If the action is multidimensional, it can be segmented and combined in each dimension in the same way. After discretization, the internal state of each agent will be classified into a discrete label, and the action will also become a finite number of optional discrete combinations, so that the reinforcement learning network can estimate the Q value on a finite state-action pair, and avoid the complex gradient problem caused by local-dimensional continuous operations in subsequent training, thereby enhancing computational controllability and convergence performance. According to the discrete state set and discrete action set, a dual deep Q network structure is constructed to obtain an online Q network and a target Q network; in this step, it is necessary to first determine the size of the network input to match it with the discrete state label or state vector. If the state is a multidimensional discrete number, it can be converted into a one-hot vector or an embedding vector, and then input into the convolution or fully connected layer of the neural network. The output dimension of the network is consistent with the number of discrete actions, and each output node corresponds to the Q value estimation of an action. To implement dual DQN, two networks need to be built: the online Q network is used for real-time prediction and to guide action selection based on its output, and the target Q network calculates the target Q value during the training phase to improve stability. The parameters of the target Q network can be the same as those of the online Q network at the time of initialization, and then gradually synchronized through soft updates or periodic hard updates. The online network receives the current discrete state input during forward propagation, and extracts features through the hidden layer in sequence, and finally outputs a set of Q value vectors that can measure the expected return of each action, while the target network uses the same state and action information when calculating the time difference target to avoid the occurrence of training instability. In the simulation environment, the epsilon-greedy strategy is executed for each agent to select actions to obtain the current action, and based on the selected current action, the state transfer is performed in the simulation environment to obtain the next state and immediate reward; in the specific implementation, it is necessary to randomly generate a numerical value that obeys a uniform distribution at each time step. ,like , then randomly select an action from the discrete action set and regard it as the exploration action at that time. , then the Q value vector output by the online Q network is greedily selected, and the action with the largest Q value is selected. This strategy uses more exploratory actions in the early stage of training to prevent falling into the situation of poor local decision-making, and gradually reduces it in the later stage of training. Enhance the utilization of the best action. After the action is selected, the simulation environment updates the discrete state of the agent according to the action applied, and calculates the effect of the action and the impact of the interaction with the neighbors on the system, thereby generating a new next state and immediate reward. If the reward function is recorded as , then the value is calculated according to the pre-defined logic or physical equation inside the environment, and it is used as the basis for subsequent Q value calculation. After recording the new state, you can continue to enter the next time step or execute the termination condition judgment to decide whether to reset the environment. The current state, current action, immediate reward and next state are combined into a four-tuple and stored in the experience playback buffer; in this process, after the action is executed and the state transfer is completed, Encapsulated into a data entry, where represents the old discrete state label or state vector, represents the discrete action applied, It’s an instant reward. Represents the next state label or vector obtained by the evolution of the environment. The quadruple is then inserted into a queue or ring buffer structure. If the buffer capacity reaches the upper limit, the earliest entry is popped out to free up the storage location, thereby retaining fresher interaction data for subsequent small batch random sampling. This operation enables the network's training samples to be covered at different times and under different exploration strategies, avoiding the problem of excessive correlation caused by continuous moment sampling, and forming a more uniform and stable training data stream. When enough quadruple groups have accumulated in the buffer, they can be randomly batch sampled, and the parameters of the online Q network can be updated by calculating the time difference target, and coordinated with the target Q network, so that the Q value estimation gradually approaches a better value function.
[0030] 103. Receive a control instruction and a current system state, generate a future state prediction using the state prediction function, construct and solve a multi-objective game optimization problem based on the prediction result, the control instruction and the value network, and obtain a Nash equilibrium solution and a control priority order for each intelligent agent; In one embodiment of the present invention, the receiving of control instructions and the current system state, generating future state predictions using the state prediction function, constructing and solving a multi-objective game optimization problem based on the prediction results, the control instructions and the value network, and obtaining the Nash equilibrium solution and control priority order of each intelligent agent includes: receiving control instructions and the current system state, and generating a state sequence in the prediction time domain based on the current system state using the state prediction function as a prediction result; constructing a multi-objective optimization problem for each intelligent agent based on the prediction results, the control instructions and the value network, and solving the multi-objective optimization problem using an improved non-dominated sorting genetic algorithm to obtain a Pareto optimal solution set; constructing a multi-agent game problem based on the Pareto optimal solution set, and using the best response to dynamically solve the Nash equilibrium, and calculating the utility value of each intelligent agent based on the Nash equilibrium solution, and determining the control priority order in descending order of the utility value.
[0031] Specifically, the control instructions and the current system state are received, and the state prediction function is used to generate a state sequence in the prediction time domain based on the current system state as a prediction result; in implementation, it is necessary to parse the external input control instructions, extract the actions or adjustment targets that each agent needs to perform, and then obtain the global or local state information at the current moment from the system monitoring layer, organize these state vectors into a suitable input form and feed them to the trained state prediction function. The state prediction function usually contains a graph convolution unit and a time series network structure to characterize the coupling relationship and time evolution law between the agent and its neighbors. In order to generate a state sequence in the prediction time domain, it is necessary to iterate forward propagation in time in sequence. Each iteration uses the prediction result of the previous moment as the input of the current moment, and is combined with a known control instruction or scheduling scheme. The state estimate of the next moment is obtained through the aggregation operation of the gating unit and the convolution kernel, which is saved in the prediction path and continues to roll forward until the entire prediction time domain is covered. For example, if the length of the prediction time domain is , which can be done in discrete time steps Call the prediction function in sequence and collect the output , and splice them into a complete state sequence as the prediction result. In order to ensure robustness and accuracy, the dynamic perception of the adjacency matrix can be maintained in the prediction function. When it is detected that the weights of some edges have changed significantly, the prediction function will automatically recalculate the convolution aggregation weights to reflect the latest pattern of interaction between agents. In this rolling prediction mode, each agent can generate a state trajectory estimate for the future moment, and couple it with the control instructions to provide a clear dynamic evolution background for subsequent multi-objective optimization and game reasoning. If there is external disturbance or fault information in the system, it can also be injected into the input layer of the prediction function through an additional channel so that the changing trend of the external environment can be reflected in the prediction output. Based on the prediction results, control instructions and value network, a multi-objective optimization problem is constructed for each agent, and the improved non-dominated sorting genetic algorithm is used to solve the multi-objective optimization problem to obtain the Pareto optimal solution set; when executing this process, it is necessary to determine the key indicators of each agent in the future time domain according to the predicted state sequence, such as energy consumption level, output quality, safety margin or economic cost, and combine the action space or the upper and lower limits of the adjustable parameters required in the control instructions to construct a multi-objective optimization model. Assume that for the first Agents need to optimize multiple objective functions simultaneously , there are also a series of inequalities and equality constraints that constitute the constraint set , which can be expressed as ; in Represents a control action sequence or scheduling plan. The value network provides a qualitative or quantitative value measure for each agent, which is complementary to these objective functions and is used to evaluate the contribution and benefits of the action plan to the overall system outside the objective function. When using the improved non-dominated sorting genetic algorithm (NSGA-Il or its variants) for solving, the initial population needs to be randomly generated or heuristically initialized based on historical experience. Each chromosome represents a set of feasible control action sequences. . During the iteration process, it is necessary to calculate the performance of all chromosomes on different objective functions, compare non-dominated relationships and perform fast non-dominated sorting, and then select the next generation of elite individuals based on indicators such as crowding distance, and generate new individuals through crossover and mutation operations, and continue to iterate until the population converges in the target space or reaches a given number of cycles. The converged elite cluster can be regarded as a Pareto optimal solution set, in which each solution cannot be completely dominated by other solutions for multiple objectives and can be further evaluated in subsequent game links. In order to ensure that the algorithm is adaptable to multi-agent parallel solving, the improved version can additionally encode neighbor information or value network feedback components in the chromosome structure, so that individuals can simultaneously consider the current agent's selfish goals and neighbor coordination needs during the evolution process. The final Pareto frontier will strike a balance between system performance and individual efficiency. Based on the Pareto optimal solution set, a multi-agent game problem is constructed, and the Nash equilibrium is solved dynamically using the best response. The utility value of each agent is calculated based on the Nash equilibrium solution, and the control priority order is determined in descending order of utility value. When carrying out this step, it is necessary to combine the optional strategies of each agent in the Pareto solution set, and regard them as optional action configurations of multiple players in the same game scenario, and then construct a utility matrix based on the value network or other external utility functions. Specifically, each player's game strategy set consists of its corresponding alternative solutions on the Pareto frontier. If a solution has a higher priority at the objective function level, it will usually get a better score in the utility function. The goal of the game problem is to find a Nash equilibrium, that is, no individual can improve its own utility value by unilateral deviation when the strategies of other players are fixed. When using the best response dynamic solution, it is necessary to start from an initial strategy configuration and let each agent update the strategy in a sequential or parallel manner. Each time the update is made, the action that maximizes its own utility value when the strategies of other players are fixed is selected. If the system reaches a state where all players have no intention of changing their strategies, it is considered a Nash equilibrium. The calculated Nash equilibrium solution can make an explicit evaluation of the utility value of each player, and use it as a basis to sort all agents in descending order, and regard those with higher utility values as having greater influence or higher priority in the current scheduling round, so as to gain dominance when multiple agents are coordinated and controlled. In this process, if multiple agents have similar utility values, they can be regarded as having a comparable level of control, and can also be further subdivided according to detailed rules or rotation mechanisms.
[0032] 104. According to the Nash equilibrium solution, control priority order and control instruction, a robust optimization problem considering uncertainty is constructed for each intelligent agent and solved, and each intelligent agent is controlled to perform a robust control action.
[0033] In one embodiment of the present invention, according to the Nash equilibrium solution, control priority order and control instructions, a robust optimization problem considering uncertainty is constructed and solved for each intelligent agent, and each intelligent agent is controlled to perform robust control actions, including: fusing the Nash equilibrium solution and the control instructions to obtain an initial control strategy and constraints, and constructing a robust optimization problem based on the initial control strategy and constraints; using an improved scenario tree method to model the system uncertainty of the multi-agent system and generate a limited number of typical scenario sets; using a column generation algorithm to solve the robust optimization problem according to the typical scenario set to obtain a robust control strategy, and based on the obtained robust control strategy and control priority order, generate and execute the control action sequence of the intelligent agent.
[0034] Specifically, the Nash equilibrium solution and the control instructions are fused to obtain the initial control strategy and constraints, and a robust optimization problem is constructed based on the initial control strategy and constraints. In the specific implementation, it is necessary to first extract the action configuration and utility performance of each agent from the Nash equilibrium solution, and compare and fuse them with the external input control instructions. If the control instructions set limits or priority start and stop requirements for certain variables, it is necessary to add these constraints to the feasible domain of the action configuration, thereby forming a set of initial control strategies and corresponding hard constraints or soft ends. On this basis, the state variables, control variables and uncertainty components can be combined in the optimization model to form constraints with interval fluctuations or probability distribution characteristics. If an agent has a higher control priority in the Nash equilibrium, a larger allowable range is given to the agent on the same objective function weight or constraint relaxation to ensure that it strikes a balance between the overall system performance and individual goals. For the entire multi-agent system, the control action u can be combined with the uncertain disturbance. is also included in the objective function, such as ,in represents the uncertainty set, It can include multiple indicators such as energy consumption, economic losses or default penalties. This form of two-layer structure can be transformed into a single-layer robust optimization model through further algorithm design, and solved in subsequent steps, and finally a set of feasible action strategies that can guarantee performance under multiple scenarios are obtained. The improved scenario tree method is used to model the system uncertainty of the multi-agent system and generate a limited number of typical scenario sets; when performing this step, it is necessary to first identify the random factors or external disturbances that may have a significant impact on the system, such as load fluctuations, energy price fluctuations, changes in production demand, etc., and sample these random variables or classify them based on the distribution of historical big data. The scenario tree method expresses the random evolution path at different times in a tree structure that splits layer by layer. However, if the standard scenario tree generation method is directly used, the number of scenarios will increase exponentially with the number of layers. In order to reduce the scale of calculation, a similarity threshold or K-means merging strategy can be introduced in the scenario generation stage to aggregate and prune similar paths, thereby obtaining a limited number of typical scenario nodes and retaining their occurrence probability at each node. This improved process can control the scale of the scenario tree while maintaining the diversity of uncertainty, thereby improving the feasibility of subsequent robust optimization. Each scenario node contains the specific implementation of state disturbances or external parameters, which is the environment faced by multiple agents in executing control actions in this scenario. If a detailed description of dynamic coupling is required, time division can be added to the high-level nodes of the tree so that the disturbance sequence of each scenario at different times can inherit the random information of the parent node and superimpose new fluctuations. The generated finite scenario set not only reflects the distribution characteristics of key uncertainties in the system, but also retains the necessary transfer relations and probability measurements between nodes. It provides a unified framework for the subsequent use of column generation algorithms to solve robust problems in blocks. The column generation algorithm is used to solve the robust optimization problem based on the typical scenario set to obtain a robust control strategy, and the control action sequence of the agent is generated and executed based on the obtained robust control strategy and control priority order; in this link, the robust optimization model needs to be placed in a multi-scenario mixed integer or continuous decision framework, and the candidate action or constraint column is gradually expanded by alternating iterations between the main problem and the sub-problem. The main problem can be solved in the initial scenario subset to obtain a set of tentative control strategies u∗. For each uncovered scenario, the feasibility and performance of u∗ need to be verified in the subproblem. Once a scenario is found that violates the constraint or leads to an excessively high target value, its corresponding invalidity proof or constraint cutting plane will be "generated" as a new column, which will be fed back to the main problem and re-optimized, and this process will be repeated until no new columns are triggered. This column generation process can greatly reduce the difficulty of solving large-scale scenario optimization, and output a set of robust control strategies on the basis of ensuring global consistency and controllability of the worst case scenario. If an agent ranks high in the priority order, a relatively loose tolerance interval can be added to its control variables in the main problem, or a higher safety redundancy can be reserved for the agent in the worst case scenario.Finally, when the algorithm converges, the control actions of each agent in different time periods or decision cycles can be read from the solution set, and arranged in time dimension to form action sequences, which are then sent to the system execution layer to enable all agents to operate in coordination. In actual deployment, if a scenario deviation or a significant change in the uncertainty range is detected, the scenario tree update and column generation steps can be triggered again to complete the online correction and re-optimization of the new robust strategy in a timely manner.
[0035] In this embodiment, a local dynamic model is obtained by performing dynamic topological modeling on the multi-agent system and decomposing the interaction relationship of each agent based on an adaptive clustering algorithm. The deep neural network of each agent is trained using distributed reinforcement learning, and a value network and a state prediction function are output. After receiving the control instructions and the current system state, the future evolution is estimated by the state prediction function, and a multi-objective game optimization problem is constructed in combination with the value network to solve the Nash equilibrium and control priority of each agent. A robust optimization method is introduced for uncertain environments to generate control actions for each agent that can cope with environmental disturbances. The present invention achieves a balance between multi-objective requirements, has a high degree of topological adaptability and robust control performance, is suitable for complex multi-agent scenarios in industrial production, and achieves more stable and efficient automated collaborative control.
[0036] The above describes the control method of industrial automation in the embodiment of the present invention. The following describes the control device of industrial automation in the embodiment of the present invention. Figure 2 , an embodiment of the industrial automation control device in the embodiment of the present invention includes: A topology decomposition module 201 is used to perform dynamic topology modeling on a multi-agent system in an industrial system, obtain a relationship diagram between agents, and decompose the multi-agent system based on the relationship diagram using an adaptive clustering algorithm to obtain a local dynamic model of each agent; A reinforcement training module 202 is used to construct a deep neural network for each agent based on the relationship diagram between the agents and the local dynamic model, and train the deep neural network through a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each agent; The game optimization module 203 is used to receive the control instruction and the current system state, generate the future state prediction using the state prediction function, construct and solve the multi-objective game optimization problem based on the prediction result, the control instruction and the value network, and obtain the Nash equilibrium solution and control priority order of each intelligent agent; The robust control module 204 is used to construct and solve a robust optimization problem taking uncertainty into consideration for each intelligent agent according to the Nash equilibrium solution, control priority order and control instructions, and obtain and control each intelligent agent to perform a robust control action.
[0037] In an embodiment of the present invention, the control device of industrial automation runs the above-mentioned control method of industrial automation. The control device of industrial automation performs dynamic topological modeling on the multi-agent system and decomposes the interaction relationship of each agent based on an adaptive clustering algorithm to obtain a local dynamic model. The deep neural network of each agent is trained by distributed reinforcement learning, and a value network and a state prediction function are output. After receiving the control instruction and the current system state, the future evolution is estimated by the state prediction function, and a multi-objective game optimization problem is constructed in combination with the value network to solve the Nash equilibrium and control priority of each agent. A robust optimization method is introduced for uncertain environments to generate control actions for each agent that can cope with environmental disturbances. The present invention achieves a balance between multi-objective requirements, has a high topological adaptability and robust control performance, is suitable for complex multi-agent scenarios in industrial production, and realizes more stable and efficient automated collaborative control.
[0038] above Figure 2 The control device for industrial automation in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The control device for industrial automation in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0039] Figure 3 3 is a structural diagram of an industrial automation control device provided by an embodiment of the present invention. The industrial automation control device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the industrial automation control device 300. Furthermore, the processor 310 can be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the industrial automation control device 300 to implement the steps of the above-mentioned industrial automation control method.
[0040] The industrial automation control device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that Figure 3 The illustrated industrial automation control device structure does not constitute a limitation on the industrial automation control device provided by the present invention, and may include more or fewer components than illustrated, or a combination of certain components, or a different arrangement of components.
[0041] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the industrial automation control method.
[0042] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0043] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0044] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A control method for industrial automation, characterized in that: The industrial automation control method comprises: Perform dynamic topological modeling on the multi-agent system in the industrial system to obtain a relationship diagram between the agents, and use an adaptive clustering algorithm to decompose the multi-agent system based on the relationship diagram to obtain a local dynamic model of each agent; According to the relationship diagram between the agents and the local dynamic model, a deep neural network is constructed for each agent, and the deep neural network is trained by a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each agent; Receive control instructions and current system states, generate future state predictions using the state prediction function, construct and solve a multi-objective game optimization problem based on the prediction results, the control instructions, and the value network, and obtain Nash equilibrium solutions and control priority orders for each intelligent agent; According to the Nash equilibrium solution, control priority order and control instructions, a robust optimization problem considering uncertainty is constructed and solved for each intelligent agent, and each intelligent agent is controlled to perform robust control actions.
2. The industrial automation control method according to claim 1, characterized in that: The multi-agent system in the industrial system is dynamically topologically modeled to obtain a relationship diagram between agents, and the multi-agent system is decomposed based on the relationship diagram using an adaptive clustering algorithm to obtain a local dynamic model of each agent, including: Performing time-varying connection relationship analysis on the multi-agent system to obtain a time-varying connection matrix describing the connection relationship between agents, and constructing a relationship graph based on the time-varying matrix, wherein the nodes of the relationship graph represent agents and the edge weights represent the connection strengths; Applying an adaptive spectral clustering algorithm to calculate the Laplace matrix of the relationship graph, and performing clustering based on the eigenvectors corresponding to the Laplace matrix to obtain a clustering result of the intelligent agent; According to the clustering results, the multi-agent system is decomposed and allocated to different subsystems to obtain the subsystem division results. The subsystem division results are used to establish a local linear dynamic model for each agent based on the least squares method.
3. The industrial automation control method according to claim 2, characterized in that: The multi-agent system is decomposed according to the clustering results, allocated to different subsystems, and the subsystem division results are obtained. The subsystem division results are used to establish a local linear dynamic model for each agent based on the least squares method, including: Perform subsystem partitioning on the clustering results to obtain a preliminary subsystem partitioning scheme, in which the agents in the same cluster are assigned to the same subsystem; Identify boundary agents for the preliminary subsystem partitioning scheme to obtain a set of boundary agents, where boundary agents refer to agents that have direct connections with agents in other subsystems; The preliminary subsystem partitioning scheme is optimized according to the boundary agent set to obtain the subsystem partitioning result, and the local state space modeling of the agents in each subsystem is performed to obtain the state equation and output equation of each agent; The state equation and output equation of each intelligent agent are estimated by the least square method to obtain the local linear dynamic model of each intelligent agent.
4. The industrial automation control method according to claim 1, characterized in that: According to the relationship diagram between the agents and the local dynamic model, a deep neural network is constructed for each agent, and the deep neural network is trained by a distributed reinforcement learning algorithm to obtain the value network and state prediction function of each agent, including: Based on the relationship graph between agents, determine the neighbor set of each agent, and build a simulation environment by combining the local dynamic model and the neighbor information, which can simulate the impact of the agent's actions on the state of itself and its neighbors; Constructing a dual deep Q network structure for each agent, and executing preset steps for each agent in a simulation environment to obtain a quadruple and store it in an experience replay buffer, wherein the dual deep Q network structure includes an online Q network and a target Q network; Randomly sampling batch data from the experience replay buffer, using the online Q network and the target Q network to calculate the current Q value and the target Q value based on the target Q network respectively; By minimizing the time difference error between the current Q value and the target Q value, the parameters of the online Q network are updated, and the parameters of the online Q network are soft-updated to the target Q network at every fixed number of steps to balance the stability and efficiency of learning; Return to the step of executing preset steps for each agent in the simulation environment until the Q network converges or reaches a preset number of training times, and use the trained online Q network as the value network; Based on the neighbor set, a graph convolutional neural network is constructed as the backbone network of the state prediction function, and combined with a long short-term memory network, the state prediction function is trained to predict future states.
5. The industrial automation control method according to claim 4, characterized in that: The dual-depth Q network structure is constructed for each agent, and preset steps are executed for each agent in the simulation environment to obtain a quadruple and store it in the experience replay buffer, including: Discretize the state space and action space of each agent to obtain a discrete state set and a discrete action set; A dual deep Q network structure is constructed based on a discrete state set and a discrete action set to obtain an online Q network and a target Q network. In the simulation environment, the epsilon-greedy strategy is used to select actions for each agent to obtain the current action, and based on the selected current action, the state transition is performed in the simulation environment to obtain the next state and immediate reward; The current state, current action, immediate reward, and next state are combined into a four-tuple and stored in the experience replay buffer.
6. The industrial automation control method according to claim 1, characterized in that: The receiving control instruction and the current system state, using the state prediction function to generate a future state prediction, constructing and solving a multi-objective game optimization problem based on the prediction result, the control instruction and the value network, and obtaining the Nash equilibrium solution and control priority order of each intelligent agent include: Receiving a control instruction and a current system state, and using a state prediction function to generate a state sequence in a prediction time domain based on the current system state as a prediction result; Based on the prediction results, control instructions and value network, a multi-objective optimization problem is constructed for each intelligent agent, and an improved non-dominated sorting genetic algorithm is used to solve the multi-objective optimization problem and obtain the Pareto optimal solution set; A multi-agent game problem is constructed based on the Pareto optimal solution set, and the Nash equilibrium is solved dynamically using the best response. The utility value of each agent is calculated according to the Nash equilibrium solution, and the control priority order is determined in descending order of utility value.
7. The industrial automation control method according to claim 1, characterized in that: According to the Nash equilibrium solution, the control priority order and the control instruction, a robust optimization problem considering uncertainty is constructed for each intelligent agent and solved, and each intelligent agent is controlled to perform a robust control action, which includes: The Nash equilibrium solution and the control instruction are fused to obtain an initial control strategy and constraint conditions, and a robust optimization problem is constructed based on the initial control strategy and constraint conditions; The improved scenario tree method is used to model the system uncertainty of multi-agent systems and generate a limited number of typical scenario sets; The column generation algorithm is used to solve the robust optimization problem according to a set of typical scenarios to obtain a robust control strategy. Based on the obtained robust control strategy and control priority order, the control action sequence of the intelligent agent is generated and executed.
8. A control device for industrial automation, characterized in that: The industrial automation control device comprises: A topological decomposition module is used to perform dynamic topological modeling on a multi-agent system in an industrial system, obtain a relationship diagram between agents, and decompose the multi-agent system based on the relationship diagram using an adaptive clustering algorithm to obtain a local dynamic model of each agent; A reinforcement training module is used to construct a deep neural network for each agent based on the relationship diagram between the agents and the local dynamic model, and train the deep neural network through a distributed reinforcement learning algorithm to obtain a value network and a state prediction function of each agent; A game optimization module is used to receive control instructions and current system states, generate future state predictions using the state prediction function, construct and solve a multi-objective game optimization problem based on the prediction results, the control instructions, and the value network, and obtain a Nash equilibrium solution and a control priority order for each intelligent agent; The robust control module is used to construct and solve a robust optimization problem taking uncertainty into consideration for each intelligent agent according to the Nash equilibrium solution, control priority order and control instructions, and obtain and control each intelligent agent to perform robust control actions.
9. An industrial automation control device, characterized in that: The industrial automation control device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the industrial automation control device executes the steps of the industrial automation control method as described in any one of claims 1-7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the industrial automation control method as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Data preset learning-based dynamic material allocation system for multiple stations
CN120317643A
Layered optimization regulation and control method for electricity-hydrogen coupling in multi-energy complementary system
CN120745439A
Hierarchical optimization and control method for electro-hydro coupling in polygeneration complementary systems
CN120745439B
Casting production line energy consumption optimization control method
CN120802759A
Thermal power plant operation control optimization method and system based on artificial intelligence
CN121300048A