Electromechanical equipment adaptive control system based on depth deterministic strategy gradient algorithm
Through the deep deterministic strategy gradient algorithm, the adaptive control system of electromechanical equipment is solved, and the problem that electromechanical equipment cannot adaptively adjust in complex environments is realized, adaptive adjustment of future control parameters is achieved, and the robustness and intelligence level of the equipment are improved.
Patent Information
- Application Number
- CN202510819426.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The control system of electromechanical equipment cannot be adaptively adjusted in complex environments, resulting in the operating state deviating from the ideal goal and affecting the performance and efficiency of the equipment.
Adaptive control system of electromechanical equipment based on deep deterministic strategy gradient algorithm is adopted to realize adaptive adjustment of future control parameters through data acquisition, state transfer model training, operation target data prediction and control parameter adaptive model training.
It improves the robustness and adaptability of electromechanical equipment in dynamic environments, optimizes equipment performance, and improves the intelligence level of equipment.
Smart Images

Figure CN120335318A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent control of electromechanical equipment, and particularly relates to an adaptive control system for electromechanical equipment based on the deep deterministic policy gradient algorithm. Background Art
[0002] During the operation of electromechanical equipment, due to the complex coupling relationship between subsystems, the parameters of the control system often undergo slight mutations, which may lead to unstable operating states of the electromechanical equipment. For example, for testing equipment, the abnormal operating state of the equipment during material performance testing may directly affect the accuracy of test results; while in production equipment, since the current operating state of the equipment affects the execution of control parameters, it may lead to energy waste and a decline in production efficiency. The former belongs to the situation where instability is caused by sudden changes in the equipment state, while the latter is due to the fact that the control parameters in actual operation fail to adapt to the current state of the equipment, resulting in the inability to achieve the ideal output or production.
[0003] Currently, most control systems of electromechanical equipment adopt traditional control strategies such as PID control and fuzzy control. These control methods can meet the conventional operation requirements in most cases. However, with the increasing complexity of the equipment operating environment and working conditions, the control systems of traditional electromechanical equipment show certain limitations when dealing with complex, non-linear, and time-varying systems. Especially when real-time adjustment of control parameters and response to dynamic working conditions are required, the adaptability of traditional control systems is poor, and they cannot effectively optimize the operating performance of the equipment. Adaptive control technology has gradually become an important direction for solving the control problems of electromechanical equipment. Traditional adaptive control systems improve system performance by adjusting controller parameters. However, many adaptive control systems are prone to over-adjustment or instability problems in complex systems, and lack the ability to accurately predict and real-time optimize the operating state of the equipment. Summary of the Invention
[0004] In order to solve the problem that the control system of electromechanical equipment cannot perform adaptive adjustment during operation, resulting in the operating state of the electromechanical equipment deviating from the ideal operating target, the present invention provides an adaptive control system for electromechanical equipment based on the deep deterministic policy gradient algorithm. This system can, under the influence of the current operating state and control parameters of the electromechanical equipment, achieve adaptive adjustment of future control system parameters to approach the ideal operating target of the electromechanical equipment as much as possible, thereby improving the performance of the electromechanical equipment and realizing the intelligent upgrade of the electromechanical equipment.
[0005] The technical solutions adopted by the present invention to solve the above technical problems are as follows:
[0006] An adaptive control system for electromechanical equipment based on the deep deterministic policy gradient algorithm, the system includes:
[0007] A data acquisition module, which is used to acquire the operation status data, control parameter data, and operation target data of the electromechanical equipment, and process the acquired data;
[0008] A data storage module, which is used to store the operation status data, control parameter data, and operation target data;
[0009] A state transition model training module, which is used to splice the control parameter data and operation status data of time units to form attribute data, use the operation status data of the next 1 time unit as label data, divide the attribute data and the corresponding label data into a first training set and a first test set, and then use the first training set and the first test set to train a long short-term memory neural network. After training, a state transition model is obtained, and this state transition model is used to predict and output the operation status data of the electromechanical equipment in the next 1 time unit;
[0010] An operation target data prediction model training module, which is used to form attribute data with the control parameter data and operation status data of 1 time unit, use the operation target data of the same time unit as label data, divide the attribute data and the label data into a second training set and a second test set, and then use the second training set and the second test set to train a deep neural network. After training, an operation target data prediction model is obtained, and this operation target data prediction model is used to predict and output the operation target data of the electromechanical equipment in this time unit;
[0011] A control parameter adaptive model training module, which uses the control parameter adaptive model as an agent and the electromechanical equipment operation system as the environment, and trains the control parameter adaptive model based on the deep deterministic policy gradient algorithm. When training, the adjustable range of the control parameter data is set to times the control parameter data of the previous time unit, and the size range of the control parameter data of the th time unit conforms to formula (1):
[0012] (1)
[0013] Where, represents the control parameter data of the electromechanical equipment at the th time unit;
[0014] The learning rate functions of the set policy network and value network are shown in formula (2):
[0015] (2)
[0016] Where, represents the learning rate reference coefficient, represents the total number of iterations of the deep deterministic policy gradient algorithm, represents a positive integer greater than 0, and are hyperparameters set during the operation of the deep deterministic policy gradient algorithm;
[0017] State is the operating state data of the electromechanical device at the -th time unit, and Action is the adjusted control parameter data from the control parameter data of the electromechanical device at the -th time unit to the next time unit. The relationship between the control parameter data of the electromechanical device and the action is shown in Equation (3):
[0018] (3)
[0019] Reward is shown in Equation (4):
[0020] (4)
[0021] where is the actual operating target data of the electromechanical device at the -th time unit, is the operating target data output by the target data prediction model at the -th time unit, is the relationship function;
[0022] The interaction process between the agent and the environment during the training of the control parameter adaptive model is as follows:
[0023] Collect the operating state data and control parameter data of the electromechanical device for time units. After splicing and then standardization, they are respectively input into the state transition model and the control parameter adaptive model. The state transition model outputs the standardized operating state data for the next time unit, and the control parameter adaptive model outputs the adjusted control parameter data. The standardized operating state data forms the actual operating state data after inverse standardization. The adjusted control parameter data is added to the current control parameter data to obtain the control parameter data for the next time unit. The inverse standardized operating state data and the control parameter data for the next time unit are spliced and input into the operating target data prediction model. The operating target data prediction model outputs the predicted operating target data. This operating target data and the actual operating target data jointly determine the reward value obtained from the electromechanical device operating system for this interaction. The number of interactions between the agent and the environment in each round is
[0024] The data generation module is used to collect the most recent The operation status data and control parameter data of the electromechanical equipment for one time unit, after data standardization and data splicing, are input into the trained control parameter adaptive model. The control parameter adaptive model outputs the adjustable control parameter data for the next one time unit, and the adjustable control parameter data is added to the current control parameter data to obtain the control parameter data for the next one time unit.
[0025] The control parameter execution module is used to output the control parameter data for the next one time unit obtained by the data generation module to the electromechanical equipment, so that the electromechanical equipment operates according to the control parameter data.
[0026] The beneficial effects of the electromechanical equipment adaptive control system based on the deep deterministic policy gradient algorithm proposed by the present invention are as follows:
[0027] The present invention solves the limitations of the existing control system in complex environments by introducing the deep deterministic policy gradient algorithm. Compared with the traditional control system, the present invention realizes the intelligent optimization of the performance of electromechanical equipment by adaptively adjusting control parameters, combining the prediction of historical data and real-time feedback. By combining deep learning and reinforcement learning, the present invention not only improves the control accuracy, but also enhances the robustness and adaptive ability of the system in dynamic environments. Brief Description of the Drawings
[0028] Figure 1 is the architecture diagram of the electromechanical equipment adaptive control system based on the deep deterministic policy gradient algorithm described in the embodiment of the present invention;
[0029] Figure 2 is the training framework of the control parameter adaptive model;
[0030] Figure 3 is the reward convergence curve for training the control parameter adaptive model. Detailed Embodiments
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] Such as Figure 1As shown in the figure, the present invention proposes an adaptive control system for electromechanical equipment based on the Deep Deterministic Policy Gradient (DDPG) algorithm, aiming to optimize the operating state of electromechanical equipment and improve equipment performance through intelligent adjustment of control parameters. This control system is based on a deep reinforcement learning algorithm, automatically adjusting control parameters to make the equipment approach the preset ideal operating target as much as possible under different operating conditions. The core modules of this control system mainly include:
[0033] A data acquisition module, which is used to collect the operating state data, control parameter data, and operating target data of the electromechanical equipment in real time, and process the collected data. The data processing includes standardization processing, outlier removal, and null value filling;
[0034] A data storage module, which is used to store the operating state data, control parameter data, and operating target data;
[0035] A state transition model training module, which is used to train a long short-term memory neural network. After training, a state transition model is obtained to predict the operating state data of the electromechanical equipment in the next 1 time unit;
[0036] An operating target data prediction model training module, which is used to train a deep neural network. After training, an operating target data prediction model is obtained to predict the operating target data of the electromechanical equipment in 1 time unit;
[0037] A control parameter adaptive model training module, which is used to train the control parameter adaptive model based on the Deep Deterministic Policy Gradient algorithm;
[0038] A data generation module, which is used to optimize the control parameters based on the trained control parameter adaptive model to obtain the control parameter data in the next 1 time unit;
[0039] A control parameter execution module, which is used to output the control parameter data in the next 1 time unit to the electromechanical equipment, so that the electromechanical equipment operates according to the control parameter data to achieve target optimization.
[0040] Specifically, the data acquisition module collects and processes data, and the data processing includes data standardization, outlier removal, and null value filling, etc.
[0041] The data acquisition module respectively collects the data reflecting the operating state of the electromechanical equipment, i.e., the operating state data, the parameter data for controlling the operation of the electromechanical equipment, i.e., the control parameter data, and the operating target data of the electromechanical equipment. Among them, the operating state data is collected through sensors, the control parameter data is collected from the control system of the electromechanical equipment, and the operating target data is collected through sensors.
[0042] Data processing refers to performing standardization processing, outlier removal, and null value filling on the above three types of data.
[0043] When outliers are removed, the data acquisition module uses the box plot outlier removal method, and the data above the upper whisker and below the lower whisker of the box plot of each type of attribute value is replaced with null values.
[0044] When null values are filled, the data acquisition module uses the K-nearest neighbor algorithm, where the size of the neighborhood scale of the K-nearest neighbor algorithm is set to 5.
[0045] The processed operation status data, control parameter data, and operation target data are stored in the data storage module.
[0046] The state transition model can dynamically reflect the state relationship during the operation of the electromechanical device. The state transition model training module trains the state transition model of the electromechanical device operation state, so that based on the known electromechanical device operation status data for a time unit can accurately predict the operation status data for the next 1 time unit, and its training method is based on deep learning.
[0047] In the state transition model training module, The control parameter data and operation status data for a time unit are concatenated to form attribute data, and the operation status data for the next 1 time unit is used as label data. The attribute data and the corresponding label data are divided into a first training set and a first test set in a ratio of 7:3, and then the first training set and the first test set are used to train the Long Short Term Memory Network (LSTM). After training, a state transition model is obtained, which is used to predict and output the operation status data of the electromechanical device for the next 1 time unit. The LSTM network architecture is an LSTM network with an input layer of 16, a hidden layer with 40 neurons, and 4 stacked layers, an input layer with 240 neurons, four hidden layers with 512 neurons each, and an output layer with 8 neurons. The ReLU activation function is used for the input layer and the hidden layers. The running parameters of the LSTM include: the total number of iterations is 400, the learning rate is 0.0001, the optimizer is the Adam optimizer, the loss function is the mean squared error, and the training batch size is 256.
[0048] The operation target data prediction model training module trains the operation target data prediction model, so that the operation target data of the electromechanical device for a time unit can be obtained based on the operation status data and control parameter data of the electromechanical device for 1 time unit, and its training method is also based on deep learning.
[0049] In the operation target data prediction model training module, the attribute data is composed of control parameter data and operation status data with a time unit of 1, and the operation target data with the same time unit is used as the label data. The attribute data and the label data are divided into a second training set and a second test set at a ratio of 7:3. Then, the second training set and the second test set are used to train a deep neural network (DNN). After training, an operation target data prediction model is obtained, which is used to predict and output the operation target data of the electromechanical equipment for this time unit. The network architecture of the DNN is an input layer with 16 neurons, three hidden layers with 512 neurons each, and an output layer with 1 neuron. The ReLU activation function is used in the input layer and the hidden layers. The operation parameters of the DNN include: the total number of iterations is 200, the learning rate is 0.001, the optimizer is the Adam optimizer, the loss function is the mean squared error, and the training batch size is 256.
[0050] The role of the control parameter adaptive model training module is to use the control parameter adaptive model as an agent and the electromechanical equipment operation system as the environment, and train the control parameter adaptive model based on the deep deterministic policy gradient algorithm. Finally, a trained control parameter adaptive model is obtained, so as to use the trained control parameter adaptive model to adjust the control parameter data of the electromechanical equipment for the next 1 time unit.
[0051] In the control parameter adaptive model training module, after building the neural network of the control parameter adaptive model, using the control parameter adaptive model as an agent and the electromechanical equipment operation system as the environment, the interaction framework between the control parameter adaptive model as an agent and the environment is as Figure 2 shown, and the control parameter adaptive model is trained based on the deep deterministic policy gradient algorithm. When training, the adjustable range of the control parameter data is set to times the control parameter data of the previous time unit, and the size range of the control parameter data of the th time unit conforms to formula (1):
[0052] (1)
[0053] where, represents the control parameter data of the electromechanical equipment for the th time unit, represents the control parameter data of the electromechanical equipment for the th time unit, and the value of is 0.06.
[0054] The learning rate functions of the set policy network (representing the control parameter adaptive model) and the value network are shown in formula (2):
[0055] (2)
[0056] Among them, represents the learning rate, represents the learning rate reference coefficient, represents the total number of DDPG iterations, represents a positive integer greater than 0, and belong to the hyperparameters newly added in the DDPG runtime settings of the present invention. and These two hyperparameters can adjust the convergence speed of the learning rate so that when the training control parameter adaptive model converges, it can converge to a better value.
[0057] The elements required for DDPG training are designed as follows: State is the operating state data of the electromechanical device at the th time unit; Action is the control parameter data adjusted from the control parameter data of the electromechanical device at the th time unit to the next time unit. The relationship between the control parameter data of the electromechanical device and the action is shown in formula (3):
[0058] (3).
[0059] Reward is the value output by taking the real operation target data of the electromechanical device and the operation target data predicted by the operation target data prediction model at the th time unit in the historical dataset as independent variables through a functional relationship as shown in formula (4):
[0060] (4)
[0061] Among them, is the real operation target data of the electromechanical device at the th time unit, is the operation target data predicted by the operation target data prediction model at the th time unit, is a known relationship function in advance. The larger the value of the reward , the better.
[0062] The interaction process between the agent and the environment when training the DDPG training control parameter adaptive model is as follows. Here, the agent is the control parameter adaptive model, and the environment is the electromechanical equipment operation system:
[0063] Collect The operation status data and control parameter data of the electromechanical equipment for a certain number of time units are collected. After concatenating the operation status data and control parameter data and normalizing them, they are respectively input into the state transition model and the control parameter adaptive model. The state transition model and the control parameter adaptive model respectively output the normalized operation status data and the adjusted control parameter data for the next time unit. The normalized operation status data forms the real operation status data after inverse normalization. The adjusted control parameter data is added to the current control parameter data to obtain the control parameter data for the next time unit. The concatenated inverse-normalized operation status data and control parameter data for the next time unit are input into the operation target data prediction model to output the predicted operation target data. The predicted operation target data and the real target value, that is, the real operation target data, jointly determine the reward value obtained from the electromechanical equipment operation system in this interaction. The return obtained from each round of interaction between the agent and the environment As shown in formula (5):
[0064] (5)
[0065] Among them, represents the number of interactions between the agent and the environment in each round.
[0066] The network architecture of the policy network of the control parameter adaptive model is an LSTM network with an input layer of 16, a hidden layer with 40 neurons, and 4 stacked layers, an input layer with 240 neurons, four hidden layers with 256 neurons each, and an output layer with 8 neurons. The input layer and the hidden layers use the RELU activation function. The network architecture of the value network is an input layer with 24 neurons, two hidden layers with 256 neurons each, and an output layer with 1 neuron. The input layer and the hidden layers use the RELU activation function.
[0067] When training the control parameter adaptive model based on DDPG, the learning rate functions of the policy network and the value network are set as shown in formula (2). The parameters set during the operation of DDPG include: the total number of iterations is 260, the discount rate is 0.96, the soft update parameter is 0.05, the capacity of the experience replay pool is 1000, the minimum capacity of the experience replay pool during model training is 256, the training data batch size is 128, and the learning rate benchmark coefficients representing the value network and the policy network is 0.00001, a positive integer is 10, the number of times the agent interacts with the environment during iteration is 32.
[0068] The data generation module collects the operating state data and control parameter data of the electromechanical equipment in the recent time units. After the above data is normalized and spliced, it is input into the trained control parameter adaptive model. The control parameter adaptive model outputs the adjustable control parameter data for the next 1 time unit. After adding the adjustable control parameter data for the next 1 time unit to the current control parameter data, the control parameter data for the next 1 time unit is obtained.
[0069] The control parameter execution module outputs the control parameter data for the next 1 time unit obtained by the data generation module to the electromechanical equipment, so that the electromechanical equipment operates according to the control parameter data. The adjustment of the control parameter data can be compensated by a compensator, and only the current control parameter data set needs to be compensated by the compensator.
[0070] The present invention solves the limitations of the existing control system in a complex environment by introducing the deep deterministic policy gradient algorithm. Compared with the traditional control system, the present invention realizes the intelligent optimization of the performance of the electromechanical equipment by adaptively adjusting the control parameters, combining the prediction of historical data and real-time feedback. By combining deep learning and reinforcement learning, the present invention not only improves the control accuracy, but also enhances the robustness and adaptive ability of the system in a dynamic environment. The present invention can adjust the operating state of the electromechanical equipment in real time, dynamically and automatically, optimize the operating target of the electromechanical equipment, so that it always has the best performance, and improves the intelligent level of the electromechanical equipment.
[0071] Taking the electromechanical equipment as a rotary kiln equipment for uniformly producing cement as an example, the technical solution of the present invention will be described in detail below. During the operation of the rotary kiln equipment, the goal of adjusting its control parameter data is to minimize the coal energy consumption when uniformly producing cement.
[0072] The data acquisition module collects and processes the data reflecting the operating state of the rotary kiln equipment. There are 8 attributes in the state data; the data acquisition module collects and processes the parameter data for controlling the operation of the rotary kiln equipment. There are 8 attributes in the parameter data; the data acquisition module collects and processes the data of the coal consumption per unit time of the rotary kiln equipment. There is 1 feature in the coal consumption data.
[0073] The state transition model training module concatenates the control parameter data and the operating state data of 7 time units to form attribute data, and uses the operating state data of the next 1 time unit as label data. The above data is divided into a training set and a test set in a ratio of 7:3, and the state transition model is trained based on the long short-term memory neural network (LSTM). The constructed network architecture is an LSTM network with an input layer of 16, a hidden layer with 40 neurons, and 4 stacked layers, an input layer with 240 neurons, four hidden layers with 512 neurons each, and an output layer with 8 neurons. The RELU activation function is used in the input layer and the hidden layers. The operating parameters of the LSTM include: the total number of iterations is 400, the learning rate is 0.0001, the optimizer is the Adam optimizer, the loss function is the mean squared error, and the training batch size is 256.
[0074] The operating target data prediction model training module forms attribute data with the control parameter data and the operating state data of 1 time unit, and uses the coal consumption data of the same time unit as label data. The above data is divided into a training set and a test set in a ratio of 7:3, and the operating target data prediction model is trained based on the deep neural network (DNN). The network architecture of the DNN is an input layer with 16 neurons, three hidden layers with 512 neurons each, and an output layer with 1 neuron. The RELU activation function is used in the input layer and the hidden layers. The operating parameters of the DNN include: the total number of iterations is 200, the learning rate is 0.001, the optimizer is the Adam optimizer, the loss function is the mean squared error, and the training batch size is 256.
[0075] The control parameter adaptive model training module trains the control parameter adaptive model. The control parameter adaptive model is used as the interaction framework between the intelligent agent and the environment as Figure 2 shown. The adjustable range of the control parameter is -0.06 to 0.06 times the control parameter data of the previous time unit, that is . The specific implementation process is as follows, where the intelligent agent is the control parameter adaptive model and the environment is the rotary kiln equipment operation system.
[0076] Collect The operation status data and control parameter data of the rotary kiln equipment for one time unit are spliced, standardized, and then input into the trained state transition model and control parameter adaptive model respectively. The state transition model and control parameter adaptive model respectively output the standardized operation status data and the adjusted control parameter data for the next time unit. The standardized operation status data forms the real operation status data after inverse standardization. The adjusted control parameter data is added to the current control parameter data to obtain the control parameter data for the next time unit. The inverse-standardized operation status data for the next time unit and the control parameter data for the next time unit are spliced and input into the trained operation target data prediction model to output the predicted operation target data. The reward value is the real coal consumption minus the predicted coal consumption. The return obtained from each round of interaction between the agent and the environment As shown in formula (5).
[0077] The policy network architecture of the control parameter adaptive model is an LSTM network with an input layer of 16, a hidden layer with 40 neurons, and 4 stacked layers, an input layer with 240 neurons, four hidden layers with 256 neurons each, and an output layer with 8 neurons. The ReLU activation function is used in the input layer and hidden layers; the value network architecture is an input layer with 24 neurons, two hidden layers with 256 neurons each, and an output layer with 1 neuron. The ReLU activation function is used in the input layer and hidden layers.
[0078] When training the control parameter adaptive model based on DDPG, the learning rate functions of the policy network and value network are as shown in formula (2). The parameters set during DDPG operation include: the total number of iterations is 260, the discount rate is 0.96, the soft update parameter is 0.05, the capacity of the experience replay pool is 1000, the minimum capacity of the experience replay pool during training the model is 256, the training data batch size is 128, and the learning rate benchmark coefficients representing the value network and policy network is 0.00001, a positive integer is 10, the number of times the agent interacts with the environment during iteration is 32.
[0079] The return convergence curve of the DDPG training process is as Figure 3 shown. Since the curve converges to a return value greater than 0, it proves that the method proposed in the present invention can reduce the coal consumption during the operation of the rotary kiln equipment.
[0080] The data generation module collects the operation status data and control parameter data of the rotary kiln equipment in the most recent 7 time units. After data standardization and data splicing, the data is input into the trained control parameter adaptive model. The control parameter adaptive model outputs the adjusted control parameter data, which is added to the current control parameter data to form the control parameter data for the next 1 time unit.
[0081] The control parameter execution module outputs the control parameter data for the next 1 time unit obtained by the data generation module to the rotary kiln equipment, enabling the rotary kiln equipment to operate according to this control parameter data.
[0082] The present invention aims to optimize the operation of electromechanical equipment by training a control parameter adaptive model that can output the control parameter data for the next 1 time unit for electromechanical equipment based on the operation status data and control parameter data in historical time units. It can not only adaptively adjust the control parameter data to make the electromechanical equipment approach the ideal operation target as much as possible, thereby improving the performance of the electromechanical equipment, but also contribute to the intelligent upgrade of the electromechanical equipment.
[0083] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0084] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. An adaptive control system for electromechanical equipment based on the Deep Deterministic Policy Gradient algorithm, characterized in that, including: a data acquisition module, configured to acquire the operation status data, control parameter data, and operation target data of the electromechanical equipment, and process the acquired data; a data storage module, configured to store the operation status data, control parameter data, and operation target data; A state transition model training module, which is used to concatenate the control parameter data and the operating state data of time units to form attribute data, use the operating state data of the next 1 time unit as label data, divide the attribute data and the corresponding label data into a first training set and a first test set, and then use the first training set and the first test set to train a long short-term memory neural network. After training, a state transition model is obtained, and this state transition model is used to predict and output the operating state data of the electromechanical equipment in the next 1 time unit; an operation target data prediction model training module, configured to form attribute data with the control parameter data and operation status data in one time unit, use the operation target data in the same time unit as label data, divide the attribute data and the label data into a second training set and a second test set, and then use the second training set and the second test set to train a deep neural network. After training, an operation target data prediction model is obtained, and the operation target data prediction model is used to predict and output the operation target data of the electromechanical equipment in this time unit; The control parameter adaptive model training module is used to take the control parameter adaptive model as an agent and the electromechanical equipment operation system as the environment, and train the control parameter adaptive model based on the deep deterministic policy gradient algorithm. When training, the adjustable range of the control parameter data is set to be times that of the control parameter data of the previous time unit. The size range of the control parameter data of the th time unit conforms to formula (1): (1) Among them, represents the control parameter data of the electromechanical device at the th time unit; the learning rate functions of the set policy network and value network are as shown in formula (2): (2) Among them, represents the learning rate benchmark coefficient, represents the total number of iterations of the deep deterministic policy gradient algorithm, represents a positive integer greater than 0, and are hyperparameters set when the deep deterministic policy gradient algorithm runs; Status is the operating status data of the electromechanical device for the th time unit, and the action is the control parameter data adjusted from the control parameter data of the electromechanical device for the th time unit to the next time unit. The relationship between the control parameter data of the electromechanical device and the action is shown in Equation (3): (3) Reward As shown in formula (4): (4) Among them, is the true operating target data of the electromechanical equipment at the th time unit, is the operating target data output by the target data prediction model at the th time unit, is the relationship function; the interaction process between the agent and the environment during the training of the control parameter adaptive model is as follows: Collect The operation status data and control parameter data of the electromechanical equipment for time units are collected. After splicing and then standardization, they are respectively input into the state transition model and the control parameter adaptive model. The state transition model outputs the standardized operation status data for the next time unit, and the control parameter adaptive model outputs the adjusted control parameter data. The standardized operation status data forms the real operation status data after inverse standardization. The adjusted control parameter data is added to the current control parameter data to obtain the control parameter data for the next time unit. The inverse standardized operation status data and the control parameter data for the next time unit are spliced and then input into the operation target data prediction model. The operation target data prediction model outputs the predicted operation target data. This operation target data and the real operation target data jointly determine the reward value obtained from the electromechanical equipment operation system in this interaction. The number of times the agent interacts with the environment in each round is times, and finally the trained control parameter adaptive model is obtained; A data generation module, which is used to collect the operation status data and control parameter data of the electromechanical equipment in the most recent time units. After data standardization and data splicing, the data is input into the trained control parameter adaptive model. The control parameter adaptive model outputs the adjustable control parameter data for the next 1 time unit, and adds the adjustable control parameter data to the current control parameter data to obtain the control parameter data for the next 1 time unit; a control parameter execution module, configured to output the control parameter data of the next one time unit obtained by the data generation module to the electromechanical equipment, so that the electromechanical equipment operates according to the control parameter data.
2. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1, characterized in that, the network architecture of the long short-term memory neural network is an LSTM network with an input layer of 16, a hidden layer number of 40, and a stacked layer number of 4, an input layer with 240 neurons, four hidden layers with 512 neurons each, and an output layer with 8 neurons. The RELU activation function is used in the input layer and hidden layers; the operation parameters of the long short-term memory neural network include: the total number of iterations is 400, the learning rate is 0.0001, the optimizer is the Adam optimizer, the loss function is the mean squared error, and the training batch size is 256.
3. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, the network architecture of the deep neural network is an input layer with 16 neurons, three hidden layers with 512 neurons each, and an output layer with 1 neuron. The RELU activation function is used in the input layer and hidden layers; the operation parameters of the deep neural network include: the total number of iterations is 200, the learning rate is 0.001, the optimizer is the Adam optimizer, the loss function is the mean squared error, and the training batch size is 256.
4. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, the network architecture of the policy network is a long short-term memory neural network with an input layer of 16, a hidden layer number of 40, and a stacked layer number of 4, an input layer with 240 neurons, four hidden layers with 256 neurons each, and an output layer with 8 neurons. The RELU activation function is used in the input layer and hidden layers.
5. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, the network architecture of the value network is an input layer with 24 neurons, two hidden layers with 256 neurons each, and an output layer with 1 neuron. The RELU activation function is used in the input layer and hidden layers.
6. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that The parameters set during the operation of the Deep Deterministic Policy Gradient algorithm include: the total number of iterations is 260, the discount rate is 0.96, the soft update parameter is 0.05, the capacity of the experience replay pool is 1000, the minimum capacity of the experience replay pool during training the model is 256, the training data batch size is 128, the learning rate base coefficient is 0.00001, a positive integer is 10, the number of times the agent interacts with the environment during iteration is 32.
7. The adaptive control system for electromechanical equipment based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, the reward obtained by the agent in each round of interaction with the environment is as shown in formula (5): (5) Among them, represents the reward obtained by the agent in each round of interaction with the environment.
8. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, The value of is 6.
9. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, the data acquisition module uses the box plot outlier removal method to remove outliers in the operation status data, control parameter data, and operation target data.
10. The electromechanical device adaptive control system based on the deep deterministic policy gradient algorithm according to claim 1 or 2, characterized in that, The data acquisition module uses the K-nearest neighbor algorithm to fill in the null values in the operation status data, control parameter data, and operation target data, and the size of the neighborhood scale of the K-nearest neighbor algorithm is set to 5.
Citation Information
Patent Citations
Multi-maneuvering target tracking method based on depth deterministic strategy gradient DDPG
CN111027677A
Method and device for predicting service life of chemical filter
CN119227550A
Inventory model training method and device based on deep reinforcement learning
CN119398655A
Power grid load optimization regulation and control method and system
CN120049422A