Deep reinforcement learning industrial decision intelligent construction method based on action correlation
By building a digital twin model and a hierarchical decision-making model in Unreal 5 engine, combining deep learning and PPO algorithms, the learning difficulties of industrial decision-making intelligence in the existing technology in complex environments is solved, and efficient and accurate industrial production control is achieved.
Patent Information
- Application Number
- CN202510402012.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing intelligent construction methods for industrial decision-making are difficult to achieve accurate and efficient macro and micro control in the complex and changing real industrial process, and there are problems such as long learning time, slow learning rate and poor transferability.
Build a deep reinforcement learning industrial decision-making intelligent method based on action correlation. By simulating the industrial environment in Unreal 5 engine, using sensor data to build a digital twin model, combining deep learning algorithms to describe parameter relationships, and using hierarchical models and PPO algorithms to train upper and lower-level decision models to achieve action correlation optimization.
It improves the stability and accuracy of industrial decision-making, can dynamically reflect changes in the physical environment, reduces the complexity of strategy search, and enhances the adaptability and maintainability of the model.
Smart Images

Figure CN120258081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep reinforcement learning action correlation, and particularly to an intelligent construction method for industrial decision-making of deep reinforcement learning based on action correlation. Background Art
[0002] Artificial intelligence is applied to various aspects of industrial production, including material control, time planning, cost planning, safety monitoring, production process control, etc.
[0003] Industrial decision-making intelligence is an intelligent agent driven by information of a fully interconnected industrial environment based on a simulation environment, that is, the intelligence of industrial entities (including machine equipment, management control, logistics control, environmental regulation, resource allocation) is realized through data collection, data fusion, integrated processing, modeling analysis, optimization decision-making and feedback control, adaptive optimization, etc. Specifically, the actual performance of industrial decision-making intelligence is to use the calculations and algorithms of artificial intelligence to transform the traditional human-based decision-making process into an autonomous decision-making mode based on machines or systems. At present, the construction of industrial decision-making intelligence is relatively immature, and the biggest problem with the mainstream methods is that it is difficult to be implemented. The difficulties include both the difficult simulation of the environment for constructing the intelligent agent and the difficulties in model training and real-world migration.
[0004] The existing common algorithm for constructing industrial decision-making intelligence is reinforcement learning, which initializes the policy space and constrains the policy change through a random agent and a new region strategy optimization method. Most of this algorithm framework runs in an offline process model, and on average, it can barely reach the task completion level in terms of effect. It can simply combine the strategies at the scheduling level and the machine level, and perform a certain degree of hierarchical optimization, with a certain degree of transferability. However, its learning time is long, the learning rate is poor, and due to the huge gap between the abstract scenario and the physical environment, the transferability is poor. Of course, there are also reinforcement learning methods that directly learn in the physical environment, but this method has almost no possibility of implementation because the learning process in the real physical environment will bring more risks, such as safety problems, equipment damage problems, slow and uncontrollable learning rate, and high costs.
[0005] Another method that does not rely on the construction of a complete process model is to use deep learning methods to dynamically construct industrial decision-making tasks. In particular, when facing industrial decision-making tasks in knowledge-intensive production, there is a large amount of data to be analyzed. Industrial process mining is used to directly extract data from the empirical model for deep model processing and then industrial decision-making assistance. This method aims to mine the causal relationship of potential problems in the production process based on historical logs and professional knowledge, and apply it to task improvement. This data-driven method relies on the stability of the industrial production process itself. If there is more uncertainty or data volume, it will significantly affect the accuracy and efficiency of the analysis. The reason is that the key points that hinder the effect of deep learning models on the empirical model of industrial processes are the dependence on professional knowledge, the constraints of traditional log data, and the lack of information knowledge in the professional field of the deep model itself.
[0006] In general, the current algorithm models for industrial intelligent decision-making may only have good results in relatively simple industrial production environments and specific industrial processes. Faced with the complex and ever-changing real industrial processes, considering the needs of instant and specialized decision-making, the training effect of the model is unlikely to achieve good results. At the same time, in order to simplify the uncertainty problem in the decision-making process, the structure of the process model will also be relatively simple, which leads to poor portability and maintainability of the industrial decision-making model. Faced with a relatively real industrial environment, it cannot meet the practical requirements and is difficult to exist as a successful engineering achievement. Summary of the invention
[0007] In view of the shortcomings of the prior art, the present invention provides a method for constructing industrial decision-making intelligence based on deep reinforcement learning based on action correlation, which realizes macro-control and micro-control of decision-making problems in industrial production processes in a real industrial environment, accurately and efficiently.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0009] A method for building industrial decision-making intelligence based on deep reinforcement learning based on action correlation includes the following steps:
[0010] Step 1: Build an industrial production process model; specifically, build a digital twin industrial environment Virtual Factory based on Unreal Engine 5.
[0011] Step 1.1: Build a solid model:
[0012] The Virtual Factory environment is used to simulate industrial entities in the production process. Specifically, several sensors are installed on each industrial entity to collect data in the industrial production process. The data includes environmental data, equipment status, and resource parameters. The data collected by the sensors are stored in a unified database.
[0013] Step 1.2: Construct a parametric mathematical model:
[0014] Simulate the relationships between entities through the Virtual Factory environment. Specifically, use deep learning algorithms to describe the mathematical relationships of parameters between various industrial entities, and train the parameter relationships between different industrial entities;
[0015] Step 1.2.1: Structure sensor data;
[0016] Structurally process the sensor data corresponding to the parameters between industrial entities, store the sensor data in tabular form, and retain the title and table key information of different sensor data; The specific form of the table is that the title represents the type of sensor data, and the table key information represents the specific attributes or metrics under the current data type;
[0017] Step 1.2.2: Table key semantic recognition;
[0018] Use the method of character embedding and title features to extract the table key semantic features of characters, and use the neural network LSTM as the network entity of the semantic recognition model, and train the semantic recognition model to fit the mathematical relationship between parameter data; The trained semantic recognition model is the parametric mathematical model.
[0019] Step 1.2.2.1: Initialize the semantic recognition model;
[0020] Initialize the initial weights of the semantic recognition model with one-hot codes, and use a fully connected layer to extract the semantics of characters;
[0021] Step 1.2.2.2: Train the semantic recognition model;
[0022] Normalize the table key semantic features extracted from the title and table key information, process them into the form of a one-dimensional array to obtain a feature array, and use cross-entropy loss to calculate the distribution difference of semantics. When calculating the distribution difference, randomly mask the table key semantic features;
[0023] The specific form of the cross-entropy loss is: where N represents the length of the feature array, a i and b i represent the feature values of different feature arrays at the same position i;
[0024] Step 1.2.3: Title semantic recognition, introduce the title of the table as a semantic classification feature to jointly classify the table;
[0025] Specifically:
[0026] Construct and calculate the word frequency distribution of table titles, screen keywords with different frequencies, define a threshold, and use inverse word frequency to measure the universality of keywords; where the title set is defined as T, the keyword set is K, the total number of titles is N, and the number of titles containing keyword K i is U i ; for the keyword K of any table title T i , there is an inverse word frequency TF i The calculation formula is as follows:
[0027]
[0028] where U i +1 is to avoid a division-by-zero error when U i =0; for any title T j , construct a feature vector v j , each element in the vector represents the weight of a certain keyword in the title; assuming there are n keywords and m feature categories, then the feature vector v j has a dimension of n*m; for two titles T j and T k , their semantic classification feature vectors are v j and v k respectively; the cosine similarity calculation formula is where v j ·v k represents the dot product of the vectors, ‖v j ‖ and ‖v k ‖ represent the norms of the vectors respectively; furthermore, the loss function of the cosine similarity is Furthermore, obtain the semantic classification features of different titles, represent them using an n*m feature vector, use the constructed semantic classification features of different titles as the input of the semantic recognition model, and calculate the matching degree of the semantic classification features of different titles using the loss function based on cosine similarity.
[0029] Step 1.2.4, Feature fusion; fuse the table key semantic features in Step 1.2.2 and the semantic classification features in Step 1.2.3, and jointly train the semantic recognition model to obtain a parametric mathematical model;
[0030] Step 1.3, Construct a digital twin model;
[0031] Complete the visualization function through the Virtual Factory environment, and construct a three-dimensional digital environment to simulate the physical scene of industrial production; use the three-dimensional models of each entity in the factory environment as the twin mapping of the entity. Specifically, arrange the three-dimensional models of the entities according to the physical environment in the Unreal Engine 5, define their three-dimensional positions, sizes, and the parameters corresponding to each entity model.
[0032] Step 1.4: Construct the Virtual Factory environment;
[0033] Mount the parameter mathematical model in Step 1.2 and the digital twin model in Step 1.3 on the parameters corresponding to the twin entity, so as to simulate the physical environment and parameter change relationship of the real factory;
[0034] Step 2: Construct the hierarchical model of the industrial decision-making agent strategy:
[0035] The hierarchical model of the industrial decision-making agent strategy includes two hierarchical models: the upper-layer scheduling model and the lower-layer decision-making model. Each layer has three parts: an action model, a value model, and a decision-making model;
[0036] Step 2.1: Construct the hierarchical model:
[0037] Use the strategy model in the Virtual Factory environment, which is specifically divided into two upper and lower hierarchical models. The upper-layer scheduling model is used to schedule the task goals among industrial entities, and the lower-layer decision-making model is used for the optimization decision of the corresponding goals of each industrial entity;
[0038] Step 2.2: Construct the action model:
[0039] Step 2.2.1: Construct the action model of the upper-layer scheduling model:
[0040] The upper-layer scheduling model is used to schedule the task goals of each industrial entity. Let the total number of industrial entities be NF, and each industrial entity is represented by F i For the industrial entity F i There is an output parameter target set The scheduling of the upper-layer scheduling model itself is to schedule the output parameters of each industrial entity. Therefore, its action space is OU = {OU i , i ∈ (1, NF)}; In the Virtual Factory environment, before training the upper-layer scheduling model, the industrial entities will be numbered, and the output parameters of the industrial entities will be converted into one-dimensional One-hot vectors and input into the upper-layer scheduling model as key feature information;
[0041] Step 2.2.2: Construct the action model of the lower-layer decision-making model:
[0042] The lower-layer decision-making model is used to optimize the decision-making process within the industrial entity to obtain the optimal decision action for the current goal. Define the input parameter IN i of each industrial entity F i as the action space, and then perform distributed optimization and training under the condition that the output parameter OU i is fixed;
[0043] Step 2.3, Construct the value model:
[0044] Construct the value model based on the critic in the actor-critic structure. Specifically, it is a policy gradient algorithm that calculates the policy value using the value model. Use Generalized Advantage Estimation (GAE) to measure the variance and bias of the return value estimation as follows:
[0045]
[0046] where γ is the discount factor, λ is the smoothing parameter, and δ k+t is the temporal difference error. K is the index of the time step, used to represent the future time steps starting from the current time step t. Set the objective function of the value model as follows:
[0047]
[0048] where, is the value function target, is the empirical expectation, representing the expectation of the empirical data at time step t. V θ (s t ) is the value function, used to estimate the value of state s t , is the value function target, representing the target value of the value function. In both the upper-layer scheduling model and the lower-layer decision-making model, the value model is used to evaluate the value of decisions online.
[0049] Step 2.4, Construct the decision model:
[0050] Construct the decision model based on the actor in the actor-critic structure, and output different features and different actions for the different action spaces of the upper-layer scheduling model and the lower-layer decision-making model respectively;
[0051] Step 2.4.1, Importance sampling:
[0052] First, evaluate the difference between the new and old policies by calculating the correlation between them. Use the importance sampling formula expression as follows, where ρ t (θ) is the importance sampling ratio, representing the probability ratio of the new and old policies choosing action a t in state s t . π θ (a t |s t ) is the probability of the new policy choosing action a t in state s t , is the probability of the old policy choosing action a t in state st The probability, θ is the parameter of the new policy, and θ old is the parameter of the old policy:
[0053]
[0054] Step 2.4.2, Gradient clipping:
[0055] Use the method of gradient clipping to control the update step size of the policy model. Let where θ is the policy parameter, is the estimator of the advantage function at time step t, clip is the clipping operation, restricting the value range of ρ t (θ). After clipping, the parameters obtained by importance sampling vary between 1 - ∈ and 1 + ∈, where ∈ is a very small positive number.
[0056] Step 2.4.3, Value function calculation:
[0057] Use L CLIP +β S S[π] to optimize the value function of the policy network, where β S is the weight constant of entropy addition, and S is the entropy addition of the policy.
[0058] Step 2.4.4, Construct the decision model of the upper - layer scheduling model:
[0059] Use the deep reinforcement learning algorithm to perform online training on the decision model constructed in Steps 2.4.1 - 2.4.3;
[0060] Step 2.4.5, Construct the decision model of the lower - layer decision model:
[0061] The construction of the decision model of the lower - layer decision model is optimized by adding action - correlation information on the basis of the above - mentioned decision model, and is divided into an out - of - group policy stage and an inter - group optimization stage;
[0062] Step 2.4.5.1, Action grouping;
[0063] Use the noise interference degree clustering method to group the action space of the lower - layer decision model, that is, add a noise σ to an action in the pre - experiment, and set a threshold by observing the perturbation of the remaining actions. Group the actions with perturbation interference greater than the set threshold for clustering.
[0064] Step 2.4.5.2, Out - of - group optimization stage:
[0065] Process action groups separately using a parallel path network, and fit the relationships between groups via a combination network; the decision model additionally outputs k group unit value heads to evaluate the value of each group's strategy in the action group, and trains the objective based on the clip method;
[0066] Step 2.4.5.3, Inter-group Optimization Phase:
[0067] In the inter-group optimization phase, use the joint objective of the group unit loss to optimize the policy network, where α i , i = 1, 2, …, k is the weight function to measure the weight of each group unit loss in the total loss; is the group unit loss;
[0068] where the group unit loss is calculated as:
[0069] The group unit loss where is the group unit value head of the decision model, and the superscript i indicates which group of action groups the loss corresponds to,
[0070] Step 3, Select a training algorithm to train the upper-layer scheduling model and the lower-layer decision model;
[0071] The industrial decision-making agent policy hierarchical model is an upper-layer scheduling model and several lower-layer decision models. At the beginning of training, the action space of the upper-layer scheduling model is determined by the output of the lower-layer decision models; both the upper-layer scheduling model and the lower-layer decision models are trained using the PPO algorithm;
[0072] Step 4, Design the model architecture;
[0073] The output of the upper-layer scheduling model is directly used as the target of the lower-layer decision model, and specifically, it is an overall directly hierarchical association relationship in the model architecture; the training algorithms for both the upper-layer scheduling model and the lower-layer decision models are the PPO algorithm, and the network architectures of both the upper-layer scheduling model and the lower-layer decision models are designed based on the standard Actor-Critic network architecture; in this architecture, the lower-layer decision model is used to parameterize the representation of the policy, and the policy is represented using a linear function after parameterization or in combination with a neural network of deep learning. At the same time, the value model estimates the state value using the state value network, calculates the policy value in parallel, and updates the decision model;
[0074] Step 5, Model Training:
[0075] Based on the PPO algorithm and the model architecture constructed in Step 4, train the corresponding agent models in the upper-layer scheduling model and the lower-layer decision models; first train the lower-layer decision models, and then train the upper-layer scheduling model, and iterate in a loop;
[0076] Step 5.1. Agent Training of the Lower-Level Decision Model:
[0077] The task objective of the lower-level decision model is to control the current industrial entity to reach the task objective parameters with the optimal strategy, and the PPO algorithm is used for training.
[0078] Step 5.1.1. Training Process of the Lower-Level Decision Model:
[0079] When the lower-level decision model is trained for the first time, the upper-level scheduling model has not been trained yet. At this time, an initial task objective parameter is set for the industrial entity, and the setting standard of the parameter is determined by the average value or mode of the usual task objective parameters. Specifically, a target parameter sequence is constructed; the training target parameters are randomly set in this sequence and iterated multiple times.
[0080] Step 5.1.2. Training Optimization of the Lower-Level Decision Model:
[0081] During the initial stage of training, some entity adjustable parameters are fixed, that is, the fixed parameters during training will not participate in the update. As the effect of the lower-level decision model is optimized, the fixed parameters will be set to the updatable state, which is called thawing. The selection of parameter thawing is random, so as to improve the effect and robustness of the lower-level decision model.
[0082] Step 5.2. Agent Training of the Upper-Level Scheduling Model:
[0083] The task objective of the upper-level policy is to schedule the target parameters of several lower-level decision models, allocate appropriate task objectives for the industrial entity, and overall schedule the industrial production process. It is also trained using the PPO algorithm.
[0084] When training the upper-level scheduling model, the lower-level decision model has completed at least one round of training. At this time, the upper-level scheduling model is trained using unified constraints; the constraints of the upper-level scheduling model are directly specified by people; when training the upper-level policy, the unified constraints are used to dynamically calculate the value function to measure the upper-level scheduling policy reward; the training process is that the upper-level policy allocates task objectives to the lower-level policy model in stages, and the lower-level decision model makes specific production decisions. In the actual target allocation stage, each stage is five minutes.
[0085] Step 5.3. Agent Training of the Overall Model;
[0086] Steps 5.1 and 5.2 are continuously repeated until the training effect is evaluated manually.
[0087] Step 6. Decision Making Construction;
[0088] Through Steps 1 to 5, the trained industrial decision-making agent model is obtained, and industrial decision-making is constructed based on this.
[0089] Step 6.1. Real-Time Data Collection and Preprocessing:
[0090] Collect environmental data, equipment status, and resource parameters in real time through sensors installed on industrial entities; the data format needs to be consistent with the structured table in the training phase;
[0091] Step 6.2, Data Input and Decision Inference:
[0092] Input the data obtained in Step 6.1 into the trained industrial decision intelligent agent model, and let it output the corresponding decision actions;
[0093] Step 6.3, Decision Execution and Dynamic Adjustment: Send the decision actions obtained in Step 6.2 to the industrial entity in the form of control instructions.
[0094] The beneficial effects produced by adopting the above technical solutions are as follows:
[0095] The present invention provides a method for constructing an industrial decision intelligence based on action correlation and deep reinforcement learning. Compared with the prior art, the present invention decouples the high-dimensional action space into low-dimensional action subspaces through the action grouping method, reduces the complexity of policy search, and optimizes by introducing action correlation to ensure the coordination of actions within the group and the independence of actions between groups, thereby improving the stability and accuracy of decision-making. At the same time, in the training environment, the present invention constructs a high-fidelity virtual factory model through digital twin environment and semantic feature fusion, which can dynamically reflect the changes in the physical environment, and to a certain extent makes up for the phenomenon that traditional methods rely on accurate process models and are difficult to adapt to the dynamically changing industrial environment. Brief Description of the Drawings
[0096] Figure 1 It is a schematic diagram of the entity of the digital twin environment provided by an embodiment of the present invention;
[0097] Figure 2 It is a schematic diagram of the digital twin process model construction architecture provided by an embodiment of the present invention;
[0098] Figure 3 It is a schematic diagram of the actual influence model of different action groupings provided by an embodiment of the present invention;
[0099] Figure 4 It is a schematic diagram of the architecture of the overall policy model provided by an embodiment of the present invention;
[0100] Figure 5 It is a network architecture diagram of the upper-layer policy model and the lower-layer policy model provided by an embodiment of the present invention. Detailed Embodiments
[0101] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0102] The purpose of this example is to use the twin model to abstract the entities in the industrial production process in the physical environment and to summarize the relationship between the entities. In this example, the basic entities of industrial production are divided into three parts: resources, logistics, and machine tools. Logistics and machine tools are controlled and scheduled through the Internet of Things system. In addition, the relationship between the three parts is connected by the cyclic iterative production process of the workpiece. The overall entity summary categories and relationships are as follows Figure 1 shown.
[0103] The purpose of this example is to build a digital twin environment that serves as a process model and also as a training environment for a policy model. First, the establishment of a process model relies on data from a large number of sensors on industrial entities. Secondly, in order to use the processing data of the model batch, the above data needs to be structured. Since the type of this data is tabular data, its tabular features need to be extracted. This extraction is completed by the semantic recognition model. Thirdly, after the feature data is processed by the semantic recognition model, a parameter mathematical model will be formed. The purpose of this model is to describe the mathematical relationship between each parameter and to fit the changes in each parameter in the environment. Finally, the parameter mathematical model and the three-dimensional digital model obtained above are combined to jointly construct a digital twin process model. The overall construction content is as follows: Figure 2 shown.
[0104] In the construction of the lower-level decision model, the action space of the decision model is clustered using the action grouping method. This is to reduce the negative impact of the dynamic uncertain correlation between the actions of the model on the model output results. In addition, the method of selecting action grouping is the noise interference degree clustering method, that is, by adding a noise σ to the action in the experiment, by observing the disturbance of the remaining actions, setting a threshold, and grouping the actions with greater disturbance interference into one group to cluster the actions. Substituting the clustered groups into PPO for experiments, it is found that the effect is the best, such as Figure 3 As shown in the figure, a comparative experiment was conducted between cluster grouping (grouping 1) and two other different groupings (grouping 2 and grouping 3). Note that when grouping in this way, the actions can be divided into 2-4 groups according to the size of the threshold. The more groups there are, the more training memory is required during training. Therefore, in actual practice, only selecting two groups does not necessarily select the grouping with the best effect.
[0105] The purpose of this example is to use the deep reinforcement learning method to optimize different task objectives respectively. First, it is necessary to analyze the organic unity of the overall scheduling task objective and the entity optimization task objective, and abstract the non-deterministic objectives of the upper-layer scheduling model. For example, in actual industrial production, in addition to tasks such as maximizing production and minimizing costs, there may also be tasks in categories such as fire protection and environmental protection. Such objectives obviously require a configurable overall scheduling. Therefore, the actual upper and lower layer model architectures are as follows Figure 4 shown. It is necessary to manually set constraints in the upper-layer scheduling model to optimize the task objectives of the lower-layer decision-making model. The task of the lower-layer decision-making model is to optimize the decision under the current objective.
[0106] The purpose of this example is to illustrate the specific implementation structure of the lower-layer decision-making model and compare it with the structure trained by the PPO algorithm to clarify the training conditions of the lower-layer decision-making model. As Figure 5 shown, disjoint networks are used to represent the policy unit and the value unit. The value unit consists of a value network and its corresponding loss function, etc., and the same is true for the policy unit.
[0107] Among them, the policy unit (lower) and the value unit (upper) are independent of each other to reduce interference between objectives. The policy unit uses a parallel path network to process action groups separately and fits the inter-group relationship through a combination network. The policy unit will additionally output n group unit value heads to evaluate the value of each group of policies of the action groups. Parameter sharing is carried out through the group unit value heads output by the policy unit. That is to say, the group unit value heads are used to connect the originally independent policy unit and value unit, extract information related to the action group characteristics of the policy, and share all parameters except the network structure information. The general policy unit is refined into two training stages, while the original objective of the value unit is retained. During the training process, E1 and E2 are set as the number of training rounds for the two stages respectively, and the Round parameter is set as the start period of the intra-group optimization stage. This structure can theoretically control the training ratio of the two learning units by adjusting these three hyperparameters, and then adjust the data reusability of the two learning units.
[0108] A method for intelligent construction of industrial decision-making based on action correlation includes the following steps:
[0109] Step 1, construct an industrial production process model;
[0110] The training environment for industrial decision-making intelligence is the industrial production process model (referred to as the process model in this paper), and the virtual environment is the test bed for the deep reinforcement learning algorithm. To verify the effectiveness of the deep reinforcement learning method for industrial decision-making agents based on action correlation, a digital twin industrial environment, Virtual Factory, was specifically constructed based on the Unreal Engine 5.
[0111] Step 1.1: Construct the entity model:
[0112] Simulate industrial entities in the production process through the Virtual Factory environment. Specifically: First, a number of sensors are installed on each industrial entity to collect data in the industrial production process. The data includes environmental data, equipment status, and resource parameters. The data collected by the sensors are all stored in a unified management database.
[0113] Step 1.2: Construct the parameter mathematical model:
[0114] Simulate the relationship between entities through the Virtual Factory environment. Specifically, use a deep learning algorithm to describe the mathematical relationship between the parameters of each industrial entity, and train the parameter relationship between different industrial entities.
[0115] Step 1.2.1: Structuring sensor data;
[0116] Structurally process the sensor data corresponding to the parameters between industrial entities, store the sensor data in tabular form, and retain the title and table key information of different sensor data; The tabular form is specifically that the title represents the type of sensor data, and the table key information represents the specific attributes or indicators under the current data type.
[0117] Step 1.2.2: Table key semantic recognition;
[0118] In order to more accurately extract the semantic information of the table keys, the present invention uses the method of character embedding and title features to extract the table key semantic features of characters, and uses the neural network LSTM as the network entity of the semantic recognition model, and trains the semantic recognition model to fit the mathematical relationship between parameter data; The trained semantic recognition model is the parameter mathematical model.
[0119] Step 1.2.2.1: Initializing the semantic recognition model;
[0120] Initialize the weights of the semantic recognition model with one-hot codes, and use a fully connected layer to extract the semantics of characters.
[0121] Step 1.2.2.2: Training the semantic recognition model;
[0122] Normalize the table key semantic features extracted from the title and table key information, process them into the form of a one-dimensional array to obtain a feature array, calculate the distribution difference of semantics using cross-entropy loss, and in order to improve robustness, perform random masking processing on the table key semantic features when calculating the distribution difference;
[0123] The specific form of the cross-entropy loss is as follows: where N represents the length of the feature array, a i and b i represent the feature values of different feature arrays at the same position i;
[0124] Step 1.2.3, Title semantic recognition;
[0125] Considering that there are cases where the table keys of different tables are the same, which may lead to table ambiguity problems, the title of the table is introduced as a semantic classification feature to jointly classify the tables;
[0126] Specifically:
[0127] Construct and calculate the word frequency distribution of the table title, screen keywords with different frequencies, and define a threshold and use inverse word frequency to measure the universality of keywords; among them, define the title set as T, the keyword set as K, the total number of titles as N, and the number of titles containing keyword K i is U i ; for any keyword K i in the table title T, there is an inverse word frequency TF i The calculation formula is as follows:
[0128]
[0129] where, U i +1 is to avoid division by zero error when U i =0; for any title T j , construct a feature vector v j , each element in the vector represents the weight of a certain keyword in the title; assuming there are n keywords and m feature categories, then the dimension of the feature vector v j is n*m; for two titles T j and T k , their semantic classification feature vectors are v j and v k respectively; the cosine similarity calculation formula is where v j ·v k represents the dot product of vectors, ‖v j ‖ and ‖v k ‖ represent the norms of the vectors respectively; furthermore, the loss function of the cosine similarity is Furthermore, semantic classification features of different titles are obtained and represented using an n*m feature vector. The constructed semantic classification features of different titles are used as the input of the semantic recognition model, and a loss function based on cosine similarity is used to calculate the matching degree of the semantic classification features of different titles.
[0130] Step 1.2.4, Feature fusion;
[0131] Fuse the table key semantic features in Step 1.2.2 and the semantic classification features in Step 1.2.3, and jointly train the semantic recognition model to obtain a parameter mathematical model;
[0132] Step 1.3, Construct a digital twin model;
[0133] Complete the visualization function through the Virtual Factory environment, and construct a three-dimensional digital environment to simulate the physical scene of industrial production; use the three-dimensional models of each entity in the factory environment as the twin mapping of the entity. Specifically, arrange the three-dimensional models of the entities according to the physical environment in the Unreal Engine 5, define their three-dimensional positions, sizes, and the parameters corresponding to each entity model;
[0134] Step 1.4, Construct the Virtual Factory environment;
[0135] Mount the mathematical model in Step 1.2's parameter mathematical model and Step 1.3's digital twin model on the parameters corresponding to the twin entities to simulate the physical environment and parameter change relationship of the real factory;
[0136] Step 2, Construct an industrial decision-making intelligent agent policy hierarchical model:
[0137] The industrial decision-making intelligent agent policy hierarchical model includes two hierarchical models: an upper-layer scheduling model and a lower-layer decision-making model. Each layer has three parts: an action model, a value model, and a decision-making model;
[0138] Step 2.1, Construct the hierarchical model:
[0139] Use the policy model in the Virtual Factory environment, which is specifically divided into two upper and lower hierarchical models. The upper-layer scheduling model is used to schedule the task objectives among industrial entities, and the lower-layer decision-making model is used for the optimal decision-making of the objectives corresponding to each industrial entity;
[0140] Step 2.2, Construct the action model:
[0141] Step 2.2.1, Construct the action model of the upper-layer scheduling model:
[0142] The upper-layer scheduling model is used to schedule the task objectives of each industrial entity. Let the total number of industrial entities be NF, and each industrial entity is represented by Fi It is shown that for industrial entity F i there is a target set of output parameters The scheduling of the upper-layer scheduling model itself schedules the output parameters of each industrial entity. Therefore, its action space is OU = {OU i , i ∈ (1, NF)}; In the Virtual Factory environment, the industrial entities will be numbered before training the upper-layer scheduling model, and the output parameters of the industrial entities will be converted into one-dimensional One-hot vectors and input into the upper-layer scheduling model as key feature information;
[0143] Step 2.2.2, Construct the action model of the lower-layer decision-making model:
[0144] The lower-layer decision-making model is used to optimize the decision-making process inside the industrial entity to obtain the optimal decision-making action for the current goal. Define the input parameters (also called adjustable parameters) IN i of each industrial entity F i as the action space, and then perform distributed optimization and training under the condition that the output parameter OU i is fixed;
[0145] Step 2.3, Construct the value model:
[0146] Construct the value model based on the critic in the actor-critic structure, which is a policy gradient algorithm. Calculate the policy value using the value model; Use the Generalized Advantage Estimation (GAE) to measure the variance and bias of the return value estimation, as follows:
[0147]
[0148] where γ is the discount factor, λ is the smoothing parameter, and δ k+t is the temporal difference error. K is the index of the time step, used to represent the future time steps starting from the current time step t; Set the objective function of the value model, as follows:
[0149]
[0150] where, is the value function target, is the empirical expectation, representing the expectation of the empirical data at time step t. V θ (s t ) is the value function, used to estimate the value of state s t ; is the value function target, representing the target value of the value function. In both the upper-layer scheduling model and the lower-layer decision-making model, the value model is used to evaluate the value of the decision online.
[0151] Step 2.4. Construct a decision-making model:
[0152] Construct a decision-making model based on the actor in the actor-critic structure, and output different features and different actions for the different action spaces of the upper-level scheduling model and the lower-level decision-making model respectively.
[0153] Step 2.4.1. Importance sampling:
[0154] First, evaluate the difference between the new and old policies by calculating the correlation between them; use to limit the deviation between the new and old policies to be too large, ensuring the stability of the new policy when the difference is large or small. The importance sampling formula expression is as follows, where ρ t (θ) is the importance sampling ratio, representing the probability ratio of the new and old policies choosing action a t in state s t . π θ (a t |s t ) is the probability of the new policy choosing action a t in state s t , is the probability of the old policy choosing action a t in state s t , θ is the parameter of the new policy, and θ old is the parameter of the old policy:
[0155]
[0156] Step 2.4.2. Gradient clipping:
[0157] Use the method of gradient clipping to control the update step size of the policy model, and let where θ is the policy parameter, is the estimator of the advantage function at time step t, clip is the clipping operation, restricting the value range of ρ t (θ), and after clipping, the parameters obtained by importance sampling vary between 1 - ∈ and 1 + ∈, where ∈ is a very small positive number.
[0158] Step 2.4.3. Value function calculation:
[0159] Use L CLIP +β S S[π] to optimize the value function of the policy network, where β S is the weight constant of entropy addition, and S is the entropy addition of the policy.
[0160] Step 2.4.4. Construct the decision-making model of the upper-level scheduling model:
[0161] Use the deep reinforcement learning algorithm to perform online training on the decision-making model constructed in steps 2.4.1 - 2.4.3;
[0162] Step 2.4.5, construct the decision-making model of the lower-level decision-making model:
[0163] The construction of the decision-making model of the lower-level decision-making model is optimized by adding action correlation information on the basis of the above decision-making model, and is divided into an out-group strategy stage and an inter-group optimization stage;
[0164] Step 2.4.5.1, action grouping;
[0165] Use the noise interference degree clustering method to group the action space of the lower-level decision-making model, that is, add a noise σ to an action in the pre-experiment, and set a threshold by observing the perturbation of the remaining actions, and group the actions with perturbation interference greater than the set threshold for action clustering.
[0166] Step 2.4.5.2, out-group optimization stage:
[0167] Use parallel path networks to process the action grouping respectively, and fit the inter-group relationship through a combination network; the decision-making model additionally outputs k group unit value heads to evaluate the value of each group's strategy in the action grouping, and trains the target based on the clip method;
[0168] Step 2.4.5.3, inter-group optimization stage:
[0169] In the inter-group optimization stage, use the joint target of the group unit loss to optimize the policy network, where α i , i = 1, 2,..., k is a weight function used to measure the weight of each group unit loss in the total loss; is the group unit loss.
[0170] Among them, the group unit loss calculation:
[0171] Group unit loss where is the group unit value head of the decision-making model, and the superscript i indicates which group of action groupings the corresponding loss is;
[0172] Step 3, select a training algorithm to train the upper-level scheduling model and the lower-level decision-making model;
[0173] The industrial decision-making agent policy hierarchical model consists of an upper-layer scheduling model and several lower-layer decision-making models. At the beginning of training, the action space of the upper-layer scheduling model is determined by the output of the lower-layer decision-making models. Due to the digital twin characteristics of the Virtual Factory environment, both the upper-layer model and the lower-layer models can be trained online. That is, in this experiment, the Proximal Policy Optimization (PPO) algorithm is used to train both the upper-layer scheduling model and the lower-layer decision-making models.
[0174] Step 4: Design the model architecture, as Figure 5 shown;
[0175] The output of the upper-layer scheduling model directly serves as the target for the lower-layer decision-making models. Therefore, the lower-layer decision-making models are directly influenced by the upper-layer scheduling model, which means that the two are an integrated whole with a direct sequential relationship in the architecture, specifically a directly hierarchical and associated overall relationship in the model architecture. The training algorithms for both the upper-layer scheduling model and the lower-layer decision-making models are the PPO algorithm. The network architectures of both the upper-layer scheduling model and the lower-layer decision-making models are designed based on the standard Actor-Critic network architecture. In this architecture, the lower-layer decision-making models are used to parameterize the representation of the policy, using a linear function to parameterize or combining a neural network of deep learning to represent the policy. At the same time, the value model uses a state value network to estimate the state value, calculate the policy value in parallel, and update the decision-making models.
[0176] Step 5: Model training:
[0177] Based on the training algorithm selected in Step 3 and the model architecture constructed in Step 4, train the corresponding agent models in the upper-layer scheduling model and the lower-layer decision-making models.
[0178] Although the upper and lower models are generally trained simultaneously, in specific implementation, it is still necessary to first train the lower-layer decision-making models, and then train the upper-layer scheduling model, and iterate cyclically. That is, the initial training process follows the bottom-up order, which can greatly improve the training speed of the model.
[0179] Step 5.1: Agent training for the lower-layer decision-making models:
[0180] The task objective of the lower-layer decision-making models is to control the current industrial entity to reach the task target parameters with the optimal policy, and the PPO algorithm is used for training.
[0181] Step 5.1.1: Training process of the lower-layer decision-making models:
[0182] When the upper-layer scheduling model is not trained during the first training of the lower-layer decision model, an initial task target parameter is set for the industrial entity, and the setting standard of the parameter is determined by the average or mode of the usual task target parameters. Specifically, a target parameter sequence is constructed; in order to increase the robustness and generalization of the lower-layer decision model, the target parameters for training are randomly set in this sequence and iterated multiple times;
[0183] Step 5.1.2, Training and optimization of the lower-layer decision model:
[0184] In order to further reduce the learning difficulty of the lower-layer decision model and control the resource consumption during the training process, on the one hand, the present invention limits the maximum number of action groups of the lower-layer decision model, and on the other hand, adopts the method of curriculum learning to gradually increase the feature complexity of the lower-layer decision model; during the initial stage of training, some entity adjustable parameters will be fixed, that is, the fixed parameters during training will not participate in the update. As the effect of the lower-layer decision model is optimized, the fixed parameters will be set to the updatable state, which is called thawing. The selection of parameter thawing is random, so as to improve the effect and robustness of the lower-layer decision model;
[0185] Step 5.2, Agent training of the upper-layer scheduling model:
[0186] The task target of the upper-layer policy is to schedule the target parameters of several lower-layer decision models, allocate appropriate task targets for industrial entities, and overall schedule the industrial production process, and also use the PPO algorithm for training;
[0187] Training process of the upper-layer scheduling model:
[0188] When training the upper-layer scheduling model, the lower-layer decision model has completed at least one round of training. At this time, the upper-layer scheduling model is trained using unified constraints; the constraints of the upper-layer scheduling model are directly specified by people; when training the upper-layer policy, the unified constraints are used to dynamically calculate the value function to measure the reward of the upper-layer scheduling policy; the training process is that the upper-layer policy allocates task targets to the lower-layer policy model in stages, and the lower-layer decision model makes specific production decisions. Each stage in the actual target allocation stage is five minutes;
[0189] Step 5.3, Agent training of the overall model;
[0190] Steps 5.1 and 5.2 are continuously repeated until the training effect is evaluated manually. Note that when optimizing the update of one layer of agents, the agents of the other layer are not updated and optimized.
[0191] Step 6, Decision construction;
[0192] The industrial decision-making agent model that has completed training is obtained through Steps 1 to 5, and industrial decision-making is constructed based on this;
[0193] Step 6.1, Real-time data collection and preprocessing:
[0194] Through sensors installed on industrial entities (equipment, logistics units, etc.), environmental data (such as temperature, humidity), equipment status (such as rotational speed, energy consumption), and resource parameters (such as material inventory, energy consumption) are collected in real time. The data format needs to be consistent with the structured table in the training phase (refer to the table title and table key semantics in Step 1.2.1). Step 6.2, Data input and decision-making inference:
[0195] Input the data obtained in Step 6.1 into the trained industrial decision-making intelligent agent model to make it output the corresponding decision actions.
[0196] Step 6.3, Decision execution and dynamic adjustment:
[0197] Send the decision actions obtained in 6.2 to the industrial entity in the form of control instructions.
[0198] The above description is only the preferred embodiment of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for intelligent construction of industrial decision-making based on action correlation in deep reinforcement learning, characterized in that, It includes the following steps: Step 1: Construct an industrial production process model; specifically, a digital twin industrial environment Virtual Factory is constructed based on the Unreal Engine 5. Step 2: Construct a hierarchical model of industrial decision-making agent strategies: The hierarchical model of industrial decision-making agent strategies includes an upper-layer scheduling model and a lower-layer decision-making model, which are two hierarchical models. Each layer has three parts: an action model, a value model, and a decision-making model. Step 3: Select a training algorithm to train the upper-layer scheduling model and the lower-layer decision-making model. Specifically: At the beginning of training, the action space of the upper-layer scheduling model is determined by the output of the lower-layer decision-making model; the PPO algorithm is used to train both the upper-layer scheduling model and the lower-layer decision-making model. Step 4: Design the model architecture. The output of the upper-layer scheduling model is directly used as the target of the lower-layer decision-making model. Specifically, in the model architecture, it is a directly hierarchical and associated overall relationship; the training algorithms of both the upper-layer scheduling model and the lower-layer decision-making model are the PPO algorithm, and the network architectures of both the upper-layer scheduling model and the lower-layer decision-making model are designed based on the standard Actor-Critic network architecture; in this architecture, the lower-layer decision-making model is used to parameterize the strategy, and the strategy is represented by a linear function after parameterization or in combination with a neural network of deep learning. At the same time, the value model uses a state value network to estimate the state value, calculate the strategy value in parallel, and update the decision-making model. Step 5: Model training: Based on the PPO algorithm and the model architecture constructed in Step 4, train the corresponding intelligent agent models in the upper-layer scheduling model and the lower-layer decision-making model; first train the lower-layer decision-making model, and then train the upper-layer scheduling model, and iterate cyclically. Step 6: Decision-making construction; Obtain the trained industrial decision-making agent model through Steps 1 to 5, and use this to construct industrial decisions.
2. The intelligent construction method for industrial decision-making based on action correlation in deep reinforcement learning according to claim 1, wherein, Step 1 includes the following steps: Step 1.1: Construct an entity model: Simulate industrial entities in the production process through the Virtual Factory environment. Specifically: First, a number of sensors are installed on each industrial entity to collect data in the industrial production process. The data includes environmental data, equipment status, and resource parameters. The data collected by the sensors are all stored in a unified management database. Step 1.2: Construct a parameter mathematical model: Simulate the relationships between entities through the Virtual Factory environment. Specifically, a deep learning algorithm is used to describe the mathematical relationships between the parameters of each industrial entity, and the parameter relationships between different industrial entities are trained. Step 1.2.1: Structuring of sensor data; Structurally process the sensor data corresponding to the parameters between industrial entities, store the sensor data in tabular form, and retain the title and table key information of different sensor data; the tabular form is specifically that the title represents the type of sensor data, and the table key information represents the specific attributes or indicators under the current data type. Step 1.2.2: Table key semantic recognition; Extract the semantic features of the table keys of characters using the method of character embedding and title features, use the neural network LSTM as the network entity of the semantic recognition model, and train the semantic recognition model to fit the mathematical relationship between parameter data; the trained semantic recognition model is the parameter mathematical model; Step 1.2.2.1, Initialize the semantic recognition model; Initialize the initial weights of the semantic recognition model with one-hot codes, and use a fully connected layer to extract the semantics of characters; Step 1.2.2.2, Train the semantic recognition model; Normalize the semantic features of the table keys extracted from the title and table key information, process them into the form of a one-dimensional array to obtain a feature array, and use cross-entropy loss to calculate the distribution difference of semantics. When calculating the distribution difference, perform random masking processing on the semantic features of the table keys; The specific form of the cross-entropy loss is as follows: where N represents the length of the feature array, a i and b i represent the feature values of different feature arrays at the same position i; Step 1.2.3, Title semantic recognition, introduce the title of the table as a semantic classification feature to jointly classify the table; Specifically: Construct and calculate the word frequency distribution of table titles, screen keywords with different frequencies, and define a threshold and use inverse word frequency to measure the universality of keywords; where the set of titles is defined as T, the set of keywords is defined as K, the total number of titles is N, and the number of titles containing keyword K i is U i ; for the keyword K of any table title T i , there is an inverse word frequency TF i The calculation formula is as follows: Among them, U i +1 is to avoid division by zero error when U i = 0; for any title T j , construct a feature vector v j , and each element in the vector represents the weight of a certain keyword in the title; assume there are n keywords and m feature categories, then the dimension of the feature vector v j is n*m; for two titles T j and T k , their semantic classification feature vectors are v j and v k respectively; the cosine similarity calculation formula is where v j ·v k represents the dot product of the vectors, ‖v j ‖ and ‖v k ‖ represent the norms of the vectors respectively; furthermore, the loss function of the cosine similarity is Furthermore, obtain the semantic classification features of different titles, represent them using an n*m feature vector, use the constructed semantic classification features of different titles as the input of the semantic recognition model, and calculate the matching degree of the semantic classification features of different titles using the loss function based on cosine similarity; Step 1.2.4, Feature fusion; fuse the semantic features of the table keys in Step 1.2.2 and the semantic classification features in Step 1.2.3, and jointly train the semantic recognition model to obtain the parameter mathematical model; Step 1.3, Build a digital twin model; Complete the visualization function through the Virtual Factory environment, and build a three-dimensional digital environment to simulate the physical scene of industrial production; use the three-dimensional models of each entity in the factory environment as the twin mapping of the entity. Specifically, arrange the three-dimensional models of the entities according to the physical environment in the Unreal Engine 5, define their three-dimensional positions and sizes, and the parameters corresponding to each entity model; Step 1.4, Construct the Virtual Factory environment; Mount the parameter mathematical model in Step 1.2 and the digital twin model in Step 1.3 on the parameters corresponding to the twin entities to simulate the physical environment and parameter change relationship of the real factory.
3. The intelligent construction method for industrial decision-making based on action correlation in deep reinforcement learning according to claim 1, wherein Step 2 includes the following steps: Step 2.1, Build a hierarchical model: Use the policy model in the Virtual Factory environment, which is specifically divided into an upper-layer scheduling model and a lower-layer decision-making model. The upper-layer scheduling model is used to schedule the task goals among industrial entities, and the lower-layer decision-making model is used for the optimization decision of the goals corresponding to each industrial entity; Step 2.2, Build an action model: Step 2.3, Build a value model: Build a value model based on the critic in the actor-critic structure, which is specifically a policy gradient algorithm, and calculate the policy value using the value model; use the Generalized Advantage Estimation (GAE) to measure the variance and bias of the return value estimation, as follows: where γ is the discount factor, λ is the smoothing parameter, and δ k+t is the temporal difference error; K is the index of the time step, used to represent the future time steps starting from the current time step t; the objective function of the value model is set as follows: Among them, is the value function target, is the empirical expectation, representing the expectation of the empirical data at time step t; V θ (s t ) is the value function, used to estimate the value of state s t , and V t targ is the value function target, representing the target value of the value function; in the upper-layer scheduling model and the lower-layer decision-making model, the value model is used to evaluate the value of decisions online; Step 2.4, Build a decision model: Build a decision model based on the actor in the actor-critic structure, and output different features and different actions for the different action spaces of the upper-layer scheduling model and the lower-layer decision-making model respectively.
4. A method for intelligent construction of industrial decision-making based on action correlation in deep reinforcement learning according to claim 3, characterized in that, Step 2.2 includes the following steps: Step 2.2.1, Build the action model of the upper-layer scheduling model: The upper-level scheduling model is used to schedule the task objectives of each industrial entity. Let the total number of industrial entities be NF, and each industrial entity is represented by F i For industrial entity F i there is a set of output parameter objectives The scheduling of the upper-level scheduling model itself is to schedule the output parameters of each industrial entity. Therefore, its action space is OU = {OU i , i ∈ (1, NF)}; In the Virtual Factory environment, before training the upper-level scheduling model, the industrial entities will be numbered, and the output parameters of the industrial entities will be converted into one-dimensional One-hot vectors and input into the upper-level scheduling model as key feature information; Step 2.2.2, Build the action model of the lower-layer decision-making model: The lower-level decision-making model is used to optimize the decision-making process within an industrial entity to obtain the optimal decision-making actions for the current goal; each industrial entity F i 's inputtable parameter IN i is defined as the action space, and then distributed optimization and training will be carried out under the condition that the output parameter OU i is fixed.
5. A method for intelligent construction of industrial decision-making based on action correlation in deep reinforcement learning according to claim 3, characterized in that, Step 2.4 includes the following steps: Step 2.4.1, Importance sampling: First, the difference between the new and old policies is evaluated by calculating the correlation between them; the importance sampling formula is expressed as follows, where ρ t (θ) is the importance sampling ratio, representing the probability ratio of the new and old policies choosing action a t in state s t ; π θ (a t |s t ) is the probability of the new policy choosing action a t in state s t , is the probability of the old policy choosing action a t in state s t , θ is the parameter of the new policy, and θ old is the parameter of the old policy: Step 2.4.2, Gradient clipping: Use the gradient clipping method to control the update step size of the policy model. Among them, θ is the strategy parameter, is the estimator of the advantage function at time step t, clip is the clipping operation, and the limit ρ t The value range of (θ) varies between 1-∈ and 1+∈ after clipping, where ∈ is a minimum positive number; Step 2.4.3, Value function calculation: Use L CLIP +β S to optimize the value function of the policy network with S[π], where β S is the weight constant of entropy addition, and S is the entropy addition of the policy; Step 2.4.4, Construct the decision model of the upper-level scheduling model: Use the deep reinforcement learning algorithm to perform online training on the decision model constructed in Steps 2.4.1 - 2.4.3; Step 2.4.5, Construct the decision model of the lower-level decision model: The construction of the decision model of the lower-level decision model is optimized by adding action correlation information on the basis of the above decision model, and is divided into an out-group policy stage and an inter-group optimization stage; Step 2.4.5.1, Action grouping; Use the noise interference degree clustering method to group the action space of the lower-level decision model, that is, add a noise σ to an action in the pre-experiment, set a threshold by observing the perturbation of the remaining actions, and group the actions with perturbation interference greater than the set threshold for action clustering; Step 2.4.5.2, Out-group optimization stage: Use parallel path networks to process the action grouping respectively, and fit the inter-group relationship through a combination network; the decision model additionally outputs k group unit value heads to evaluate the value of each group's strategy of the action grouping, and train the target based on the clip method; Step 2.4.5.3, Inter-group optimization stage: During the inter-group optimization phase, the joint objective of the group unit losses is used to optimize the policy network, where α i , i = 1, 2, …, k are the weight functions, which are used to measure the weight of each group unit loss in the total loss; i = 1, 2, …, n are the group unit losses; Among them, the calculation of the group unit loss: Group unit loss where is the group unit value header of the decision model, and the superscript i indicates the loss corresponding to which group of action groupings 6. The method for intelligently constructing an industrial decision based on action correlation in deep reinforcement learning according to claim 1, characterized in that Step 5 includes the following steps: Step 5.1, Agent training of the lower-level decision model: The task objective of the lower-level decision model is to control the current industrial entity to reach the task objective parameters with the optimal strategy, and use the PPO algorithm for training; Step 5.1.1, Training process of the lower-level decision model: When the lower-level decision model is trained for the first time, the upper-level scheduling model is not trained. At this time, an initial task objective parameter is set for the industrial entity, and the setting standard of the parameter is determined by the average or mode of the usual task objective parameters; specifically, construct a target parameter sequence; the training target parameters are randomly set in this sequence and iterated multiple times; Step 5.1.2, Training optimization of the lower-level decision model: Fix some entity adjustable parameters at the initial stage of training, that is, the fixed parameters during training will not participate in the update. As the effect of the lower-level decision model is optimized, the fixed parameters will be set to the updatable state, which is called thawing. The selection of parameter thawing is random, so as to improve the effect and robustness of the lower-level decision model; Step 5.2, Agent training of the upper-level scheduling model: The task objective of the upper-level policy is to schedule the target parameters of several lower-level decision models, allocate appropriate task objectives for the industrial entity, and overall schedule the industrial production process, and also use the PPO algorithm for training; When training the upper-level scheduling model, the lower-level decision model has completed at least one round of training. At this time, use unified constraints to train the upper-level scheduling model; the constraints of the upper-level scheduling model are directly specified by people; when training the upper-level policy, use unified constraints to dynamically calculate the value function to measure the upper-level scheduling policy reward; the training process is that the upper-level policy allocates task objectives to the lower-level policy model in stages, and the lower-level decision model makes specific production decisions. The actual target allocation stage is five minutes per stage; Step 5.3, Agent training of the overall model; Continuously repeat Steps 5.1 and 5.2 until the training effect passes the manual evaluation.
7. A method for intelligent construction of industrial decision-making based on action correlation in deep reinforcement learning according to claim 1, characterized in that Step 6 includes the following steps: Step 6.1, Real-time data collection and preprocessing: Through sensors installed on industrial entities, collect environmental data, equipment status, and resource parameters in real time; the data format needs to be consistent with the structured table in the training phase; Step 6.2, Data input and decision-making inference: Input the data obtained in Step 6.1 into the trained industrial decision-making intelligent agent model, and let it output the corresponding decision actions; Step 6.3, Decision execution and dynamic adjustment: Send the decision actions obtained in Step 6.2 to the industrial entity in the form of control instructions.