Adaptive control optimization method based on deep learning and reinforcement learning
By combining spatiotemporal fusion deep learning with reinforcement learning, predicted values of future steps and optimal control strategies are generated, which solves the problems of data loss and insufficient control accuracy in nonlinear scenarios in traditional industrial intelligent control, and achieves high-precision and robust adaptive control optimization.
Patent Information
- Application Number
- CN202510945559.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-14
AI Technical Summary
Existing industrial intelligent control technologies lack control accuracy when faced with data loss and nonlinear industrial scenarios, and fail to effectively process the multi-scale temporal dynamics and spatial topological correlations of industrial data, resulting in control strategy failure and error accumulation.
A spatiotemporal fusion deep learning model combining multi-scale temporal convolutional layers and graph convolutional networks with mask matrices is used to generate predicted values for future steps. Reinforcement learning is used to generate the optimal control strategy, construct a fully closed-loop industrial control architecture, and achieve high-precision adaptive control in the absence of sensor data.
It achieves high-precision and robust adaptive control in complex industrial scenarios with partial loss of sensor data and multi-variable nonlinear coupling, solves the problem of insufficient control accuracy and control failure caused by data loss of traditional models in nonlinear industrial scenarios, and fully considers the spatiotemporal characteristics of industrial data.
Smart Images

Figure CN120779741A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial intelligent control, and particularly relates to a self-adaptive control optimization method based on deep learning and reinforcement learning. BACKGROUND
[0002] The existing industrial intelligent control is mainly a traditional control algorithm based on a mathematical model, a linear or simplified nonlinear mathematical model is established to realize control parameter adjustment, but the traditional control algorithm is designed based on linear or local linear assumption, and it is difficult to accurately describe the nonlinear dynamic characteristics in a complex industrial scene, and model error accumulation will lead to control instability.
[0003] The existing industrial intelligent control also uses a time sequence model such as a recurrent neural network and a long short-term memory network to predict future sensor data, and then generates an optimal control strategy through a reinforcement learning model, however, the deep learning model usually assumes that the input data is complete, and does not have a built-in robustness mechanism for data loss, when part of the sensor data is lost, the model directly uses the original input for prediction, and the reinforcement learning model generates a control strategy by using the predicted value with error, which causes error accumulation, and therefore, a production accident may occur in an industrial scene with high precision requirements.
[0004] The prior art such as the invention application patent with the publication number CN113325721A discloses an industrial system model-free adaptive control method and system, which comprises: obtaining historical monitoring data of various devices in an industrial process; generating a control instruction set using the controllable data; the control instruction set comprises a plurality of control instructions generated at the next moment; constructing a prediction simulation model according to the historical monitoring data; training a reinforcement learning-based control model based on the control instruction set according to the prediction simulation model, generating a trained reinforcement learning-based control model; obtaining current monitoring data; inputting the current monitoring data into the trained reinforcement learning-based control model to adaptively control the production process of the industrial system, and outputting the optimal set target of the industrial system.
[0005] For the above-mentioned scheme, the following technical problems exist: 1. The current technology mainly establishes a prediction simulation model according to historical data, and then trains a reinforcement learning-based control model to adaptively control the industrial production process, but the current technology does not consider the scenario of data loss, when data loss occurs, the effect of the control strategy generated by the model is poor, and the reliability is low, the deep learning model usually assumes that the input data is complete, and does not have a built-in robustness mechanism for data loss, when part of the sensor data is lost, the model directly uses the original input for prediction, and the reinforcement learning model generates a control strategy by using the predicted value with error, which causes error accumulation.
[0006] 2、Current technology ignores the multiscale time dynamics and spatial topology correlation of industrial data, leading to inaccurate prediction and delayed control. Current technology does not fully consider the multisource heterogeneity, spatiotemporal coupling and strong nonlinear characteristics of industrial data, resulting in failure to consider the real physical relationship between devices when designing the model, ignoring the correlation between devices when processing data, and not taking into account the changing rules of different times and spaces when training the model, ultimately leading to the inability of the control strategy to accurately match the running state of the actual device. SUMMARY
[0007] The purpose of the present application is to provide a self-adaptive control optimization method based on deep learning and reinforcement learning, which solves the problems in the background technology.
[0008] To solve the above technical problems, the technical scheme adopted by the present application is as follows: The present application provides a self-adaptive control optimization method based on deep learning and reinforcement learning, comprising: step one, processing data, and generating a corresponding mask matrix according to the processed data.
[0009] Step two, constructing a spatiotemporal fusion deep learning model, combining a multiscale time series convolution layer and a graph convolution network with the mask matrix of the data to capture the time and space features between the time series data with missing values, and finally generating a prediction value for the future number of steps, and optimizing the spatiotemporal fusion deep learning model.
[0010] Step three, constructing a state vector according to the prediction value output by the spatiotemporal fusion deep learning model, and then generating an optimal control strategy.
[0011] Step four, deploying the trained spatiotemporal fusion deep learning model and reinforcement learning strategy network model to an edge controller to realize multi-step prediction and real-time optimal control decision generation under missing sensor data, drive the actuator to realize closed-loop control, and complete self-adaptive control optimization.
[0012] The present application has the following advantages: 1. The self-adaptive control optimization method based on deep learning and reinforcement learning provided by the present application constructs a mask matrix according to the collected data, then combines a multiscale time series convolution layer and a graph convolution network with the mask matrix of the data to generate a prediction value for the future number of steps, constructs a spatiotemporal fusion deep learning model accordingly, and then constructs a state vector according to the prediction value of the spatiotemporal fusion deep learning model, so as to generate an optimal control strategy according to the state vector. The present application embeds spatiotemporal fusion deep learning and reinforcement learning in a collaborative manner to construct a full-closed-loop industrial control architecture of prediction-decision-execution, realizes inference-control integration in an edge controller for complex industrial scenes with partial missing sensor data and multivariate nonlinear coupling, and achieves high-precision, strong-robustness self-adaptive control optimization.
[0013] 1. The application first solves the problem of insufficient control accuracy of traditional mathematical models in nonlinear industrial scenarios. Through an end-to-end reinforcement learning strategy, the traditional PID parameter setting is bypassed, and a nonlinear mapping from the state space to the control strategy is directly constructed to achieve accurate action output under complex working conditions. Second, it solves the problem of control algorithm failure when sensor data is partially missing. By introducing a mask matrix in the spatio-temporal fusion deep learning model, the mask shields the missing positions to prevent invalid data from polluting the features. Finally, it solves the problem of existing deep learning models ignoring the spatio-temporal characteristics of industrial data. Through spatio-temporal fusion modeling, the serial structure of extracting temporal features first and then spatial features fully considers the spatio-temporal characteristics of industrial data. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0015] Figure 1 The method embodiment steps flowchart of the present application.
[0016] Figure 2 The spatio-temporal fusion deep learning model in the present application.
[0017] Figure 3 The reinforcement learning decision model in the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] Referring to Figure 1 The present application provides a self-adaptive control optimization method based on deep learning and reinforcement learning, comprising the following steps: step one, processing the data, and generating the corresponding mask matrix according to the processed data.
[0020] In one specific example, the data processing process is as follows: the collected real-time industrial data is denoted as data, and the data is treated as multi-sensor time series data, and the dimension is defined as X ∈ R N*T*D, wherein N represents the number of sensors, T represents the historical time step, D represents the feature dimension of each sensor, R represents the real set, X represents the multi-sensor time series data, the data is cleaned, and the format of the cleaned data is converted into a time window sliding format, wherein the window size is H, and finally, complete data segments in the data are randomly added with noise.
[0021] It should be noted that the multi-sensor time series data includes a data set collected by a plurality of sensors such as temperature sensors, image sensors, and vibration sensors in a continuous time sequence.
[0022] It should be noted that the data cleaning includes removing outliers, etc., for example, when a certain data is out of the 3σ range data, the data is removed, wherein σ is the standard deviation; the window size H is the historical step.
[0023] In one specific example, the corresponding mask matrix is generated according to the processed data, and the specific process is as follows: data traversal is performed on X according to the processed data, wherein M ∈ {0, 1} N*T is 1 when the data exists, and 0 when the data is missing, and then a mask matrix M is constructed, and finally data X ∈ R N*H*D , the mask matrix M ∈ R N*H .
[0024] Referring to Figure 2 , step two, a spatio-temporal fusion deep learning model is constructed, a multi-scale time series convolution layer and a graph convolution network are used in combination with the mask matrix of the data to capture the time and space features between the time series data with missing values, and finally a prediction value of a future step is generated, and the spatio-temporal fusion deep learning model is optimized.
[0025] In one specific example, the spatio-temporal fusion deep learning model is constructed, and the specific process is as follows: L cascaded spatio-temporal fusion blocks are used to construct the model architecture of the spatio-temporal fusion deep learning model, the multi-scale time series convolution layer and the graph convolution network are included in each spatio-temporal fusion block, the time-space features are gradually fused, and the output of each spatio-temporal fusion block is used as the input of the next spatio-temporal fusion block, and accordingly the spatio-temporal fusion deep learning model is constructed.
[0026] In one specific example, the time and space features between the time series data with missing values are as follows: first, the time features are captured by the multi-scale time series convolution layer: the feature X (L-1) ∈ R N*H*(L-1) and the mask M (L-1) ∈ R N*H, three different convolution kernels are used for processing, for each convolution kernel size k∈K={3, 5, 7}, the calculation of the time convolution layer is as follows: Wherein Conv_k(X (L-1) ) represents the convolution operation of the feature with the convolution kernel size k, according to which the time pattern and feature of the data are extracted, Mask_k(M (L-1) ) represents the convolution operation of the mask with the convolution kernel size k, according to which the effective feature is obtained; by distributing the channels to the three value convolution kernels for operation, and then merging the features of each scale, the fusion feature is finally obtained:
[0027] In one specific example, the final generated prediction value of the future step number is as follows: A1, based on the fusion feature The maximum value of the fusion feature corresponding to each value of the convolution kernel is obtained as the mask value of each position, and then the mask matrix is generated: Accordingly, the output of the time convolution layer And M′ L ∈R N*H , the output of the time convolution layer And M′ L ∈R N*H are taken as the input of the graph convolution network, first, the preliminary adjacency matrix B is generated based on the mask M′ L ∈R N*H Wherein b is learned through the model training process, represents the learnable bias vector, and the dimension is N; 1 is a full 1 vector, and the dimension is N; T is the transpose matrix symbol, M t-H:t is an N*H submatrix, is an H*N transpose matrix, An N*N matrix is obtained, which is the time window association matrix between each sensor, wherein t is the dynamic time index.
[0028] It should be noted that the graph convolution network mainly performs spatial feature capturing operation.
[0029] It should be noted that b1 T -1b T adds asymmetry to the adjacency matrix, so that the graph structure can express the directed relationship.
[0030] A2, based on the preliminary adjacency matrix B, the final adjacency matrix A of the node embedding is generated: Wherein E1 and E2 are generated using a learnable embedding layer, and represent the node embedding matrix; β is a balance coefficient, which controls the influence degree of the mask information.
[0031] A3, calculate the row sum of the final adjacency matrix A according to the out-degree matrix D0, and calculate the row sum of the final adjacency matrix A according to the in-degree matrixi The column sum of the final adjacency matrix A is obtained by bidirectional convolution operation where X L is the output of the fusion spatial feature, I unit matrix, θ L and B L It is obtained through the model training process and is the convolution parameter. i represents the node number and is a positive integer.
[0032] It should be noted that the node's own information is retained through the identity matrix.
[0033] It should be noted that the out-degree matrix D0 is a diagonal matrix, D0i represents the out-degree of node i, that is, the number of edges from node i to other nodes, and the in-degree matrix D i It is a diagonal matrix, D0j represents the in-degree of node j, that is, the number of edges from other nodes to node j.
[0034] A4. Obtain the top K connections of each node according to the adjacency matrix A, and update the mask matrix to obtain: M ( i ) =MAX(M j ),j∈{i}∪z i , where Z i The connection strength of each node is sorted into the top K connection sets, and the corresponding mask values are selected through the Topk connection, and the maximum value of each mask value is taken as the new mask value. Based on this, the mask values of N nodes are updated to obtain the mask output: M L ∈R N*H .
[0035] It should be noted that TopK connection is an existing technology and will not be described in detail.
[0036] A5, then the input X of the Lth cascaded spatiotemporal fusion block (L-1) Compared with X after deep learning of spatiotemporal fusion L Add together to generate the input X of the next cascaded spatiotemporal fusion block L After L cascaded spatiotemporal fusion blocks are processed, the final feature is mapped to the next h-step prediction value through the linear layer: Y h ∈R N*h*D .
[0037] In a specific example, the spatiotemporal fusion deep learning model is optimized, specifically through the spatiotemporal fusion model loss function: The model parameters are continuously adjusted to optimize the model, where n represents the number of sensors, n is a positive integer, and η represents the number of dynamic time indexes, η is a positive integer.
[0038] Reference Figure 3As shown, step three, according to the predicted value output by the spatio-temporal fusion deep learning model, a state vector is constructed, and then an optimal control strategy is generated.
[0039] It should be noted that the state vector includes sensor data, predicted future state, control action history, and the like.
[0040] In one specific example, the state vector is constructed according to the predicted value output by the test block fusion deep learning model, and the specific process is as follows: Y h ∈R N*h*D is spliced with the current state of the programmable logic controller, and then a state vector S(t) is constructed: d where d represents the number of parameters of the state vector, i.e., the state vector is: where e(t) is the current step error, is the inverse error, and U (t-h):(t-1) is the control action sequence of the last h steps, and e(t+h) is the error of the future hth step.
[0041] It should be noted that the number of parameters of the state vector includes current sensor data, error terms, control action history, device process constraints, target information, and the like.
[0042] In one specific example, the optimal control strategy is generated, and the specific process is as follows: C1, first, based on the Actor network, an action is generated, the state vector S(t) is input into the Actor network, and the optimal control strategy is obtained, i.e., the output action a(t) = clip(π(s(t)|θ π )+λ t ,a min ,a max ), where π(s(t)|θ π ) represents the deterministic action base value output by the Actor network after inputting the state s(t), and clip(π(s(t)|θ π )+λ t ,a min ,a max ) represents that the generated action is within the control range of the actuator.
[0043] C2, Critic evaluation and TD-error calculation are performed, the Critic evaluation calculates the current Q value according to the input state vector S(t) and action a(t): Q = Q(S(t), a(t)), the Q value is the action value function, the target network calculates the Q value according to the next state and action: Q' = Q'(S'(t+1), μ'(S'(t+1))), where Q' is the action value function of the Critic target network, and μ' is the policy function of the Actor target network.
[0044] It should be noted that Q is the action value function, which is used to quantify the long-term value of an action in a certain state.
[0045] C3, TD-error calculation is performed again, and TD-error is calculated by formula: TD-error = R t + γ * Q' - Q, where γ is learned through the model training process, indicating a discount factor, R t is an immediate reward parameter, R(t) = -(o * tracking error + p * control energy consumption + q * parameter fluctuation penalty), where o, p and q are weight factors of tracking error, weight factors of control energy consumption and weight factors of parameter fluctuation penalty, respectively.
[0046] It should be noted that the specific values of o, p and q need to be determined according to the actual scene, and are not specifically limited here; wherein the tracking error measures the degree of deviation of the control output from the target value, the control energy consumption evaluates the energy consumed by executing the control action, and the parameter fluctuation penalty measures the degree of control quantity.
[0047] C4, finally, Critic network update, Actor network update and target network soft update are performed, wherein the Critic network is updated, indicating that the mean square loss of TD-error is minimized by expression: , the Actor network is updated, indicating that the Q value is maximized by expression: , and finally the target network is soft updated, and the best control strategy is obtained accordingly, wherein τ is a model parameter.
[0048] Step four, the trained spatio-temporal fusion deep learning model and reinforcement learning strategy network model are lightweight deployed to the edge controller to realize multi-step prediction and real-time optimal control decision generation under the condition of missing sensor data, drive the actuator to realize closed-loop control, and complete adaptive control optimization.
[0049] In one specific example, the trained spatio-temporal fusion deep learning model and reinforcement learning strategy network model are lightweight deployed to the edge controller to realize multi-step prediction and real-time optimal control decision generation under the condition of missing sensor data, drive the actuator to realize closed-loop control, and complete adaptive control optimization, the specific process is as follows: the trained spatio-temporal fusion model and reinforcement learning decision module are lightweight deployed to the edge controller, a data buffer is established, preprocessed real-time data is collected and stored, and the data buffer is updated in real time, while the longest historical data in the buffer is deleted, new collected data is added, and the data in the buffer is ensured to be the sensor data of the last H time points.
[0050] The sliding window buffer is maintained by the edge controller to store the sensor data of the last H time steps, and the real-time data is dynamically preprocessed; after receiving new data, the buffer is updated and input into the deep learning model, the future h-step prediction value is generated in real time through the spatio-temporal fusion model, the optimal control strategy is generated in real time through the reinforcement learning decision module, and the actuator is driven to realize adaptive control optimization.
[0051] It should be noted that the dynamic preprocessing includes operations such as cleaning, format conversion and mask generation.
[0052] The adaptive control optimization method based on deep learning and reinforcement learning provided in the application generates a prediction value of future steps by constructing a mask matrix according to the collected data, combining a multi-scale time series convolution layer and a graph convolution network with the mask matrix of the data, constructing a spatio-temporal fusion deep learning model according to the prediction value of the spatio-temporal fusion deep learning model, and constructing a state vector according to the prediction value of the spatio-temporal fusion deep learning model, so as to generate an optimal control strategy according to the state vector. The application constructs a full-closed-loop industrial control architecture of prediction-decision-execution by embedding spatio-temporal fusion deep learning and reinforcement learning, realizes inference-control integration in the edge controller for complex industrial scenes with partial missing of sensor data and multivariate nonlinear coupling, and realizes high-precision and strong-robustness adaptive control optimization.
[0053] The above is only an example and description of the concept of the application, and those skilled in the art can make various modifications, supplements or substitutions of the described specific embodiments or use similar ways to replace them, as long as they do not deviate from the concept of the application or exceed the scope defined by the application, and they should belong to the protection scope of the application.
Claims
1. An adaptive control optimization method based on deep learning and reinforcement learning, characterized in that: include: Step 1: Process the data and generate the corresponding mask matrix based on the processed data; Step 2: Build a spatiotemporal fusion deep learning model. This model uses multi-scale temporal convolutional layers and graph convolutional networks, combined with the data mask matrix, to capture the temporal and spatial features between time series data with missing values. This model ultimately generates a prediction of the number of future steps and optimizes the spatiotemporal fusion deep learning model. Step 3: Based on the predicted value output by the spatiotemporal fusion deep learning model, a state vector is constructed to generate the optimal control strategy. Step 4: Deploy the trained spatiotemporal fusion deep learning model and reinforcement learning strategy network model to the edge controller in a lightweight manner to achieve multi-step prediction and real-time optimal control decision generation in the absence of sensor data, drive the actuator to achieve closed-loop control, and complete adaptive control optimization.
2. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 1, characterized in that: The specific process of processing the data is as follows: The collected real-time industrial data is recorded as data, and the data is used as multi-sensor time series data, and the dimension is defined as X∈R N*T*D , where N represents the number of sensors, T represents the historical time step, D represents the feature dimension of each sensor, R represents the real number set, and X represents the multi-sensor time series data. At the same time, the data is cleaned and the format of the cleaned data is converted into a time window sliding format with a window size of H. Finally, noise is randomly added to the complete data segment in the data.
3. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 2, characterized in that: The corresponding mask matrix is generated according to the processed data. The specific process is as follows: According to the processed data, data traversal is performed on X, where M∈{0,1} is set N*T , when the data exists, it is 1, and when it is missing, it is 0, and then the mask matrix M is constructed, and finally the data X∈R is obtained N*H*D , the mask matrix M∈R N*H .
4. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 3, characterized in that: The specific process of constructing the spatiotemporal fusion deep learning model is as follows: The model architecture of the spatiotemporal fusion deep learning model is composed of L cascaded spatiotemporal fusion blocks. Each spatiotemporal fusion block contains a multi-scale temporal convolution layer and a graph convolution network to achieve step-by-step fusion of time-space features. The output of each spatiotemporal fusion block is used as the input of the next spatiotemporal fusion block, thereby constructing a spatiotemporal fusion deep learning model.
5. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 4, characterized in that: The temporal and spatial characteristics between the time series data with missing values are specifically processed as follows: First, the temporal features are captured through a multi-scale temporal convolution layer: The feature X processed by the L-1 cascaded spatiotemporal fusion block is input into the L cascaded spatiotemporal fusion block. (L-1) ∈R N*H*(L-1) and mask M (L-1) ∈R N*H , using three different sets of convolution kernels for processing. For each convolution kernel size k∈K={3,5,7}, the temporal convolution layer is calculated as follows: Where Conv_k(X (L-1) ) represents the convolution operation on the feature when the convolution kernel size is k, based on which the time pattern and features of the data are extracted, Mask_k(M (L-1) ) represents the convolution operation on the mask when the convolution kernel size is k, and the effective features are obtained accordingly; by evenly distributing the channels to the convolution kernels with 3 values, the features of each scale are merged and the fusion features are finally obtained:
6. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 5, characterized in that: The final predicted value of the number of future steps is generated, and the specific process is as follows: A1. Based on fusion features The maximum value of the fusion feature corresponding to the convolution kernel of each value is obtained as the mask value of each position, and then the mask matrix is generated: Based on this, the output of the temporal convolution layer can be obtained and M′ L ∈R N*H , the output of the temporal convolutional layer and M′ L ∈R N*H As the input of the graph convolutional network, we first L ∈R N*H Generate a preliminary adjacency matrix B: Where b is learned through the model training process and represents a learnable bias vector with a dimension of N; 1 is a full 1 vector with a dimension of N; T is the transposed matrix symbol, M t-H:t is a sub-matrix of N*H, is the transposed matrix of H*N, The N*N matrix is obtained, which is the time window correlation matrix between each sensor, where t is the dynamic time index; A2. Generate the final adjacency matrix A of node embedding based on the preliminary adjacency matrix B: Where E1 and E2 are generated using a learnable embedding layer and represented as a node embedding matrix; β is the balance coefficient, which controls the influence of mask information; A3. Calculate the row sum of the final adjacency matrix A based on the out-degree matrix D0, and calculate D based on the in-degree matrix i The column sum of the final adjacency matrix A is obtained by bidirectional convolution operation where X L is the output of the fusion spatial feature, I unit matrix, θ L and B L Obtained through the model training process, it is the convolution parameter, i represents the node number, and i is a positive integer; A4. Obtain the top K connections of each node according to the adjacency matrix A, and update the mask matrix to obtain: M (i) =MAX(M j ),j∈{i}∪z i , where Z i The connection strength of each node is sorted into the top K connection sets, and the corresponding mask values are selected through the Topk connection, and the maximum value of each mask value is taken as the new mask value. Based on this, the mask values of N nodes are updated to obtain the mask output: M L ∈R N*H ; A5, then the input X of the Lth cascaded spatiotemporal fusion block (L-1) Compared with X after deep learning of spatiotemporal fusion L Add together to generate the input X of the next cascaded spatiotemporal fusion block L After L cascaded spatiotemporal fusion blocks are processed, the final feature is mapped to the next h-step prediction value through the linear layer: Y h ∈R N*h*D .
7. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 6, characterized in that: The above-mentioned optimization of the spatiotemporal fusion deep learning model is specifically achieved through the spatiotemporal fusion model loss function: The model parameters are continuously adjusted to optimize the model, where n represents the number of sensors, n is a positive integer, and η represents the number of dynamic time indexes, η is a positive integer.
8. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 7, characterized in that: The predicted value output by the deep learning model based on the test block is used to construct the state vector. The specific process is as follows: The Y output of the spatiotemporal fusion deep learning model h ∈R N*h*D Combined with the current state of the programmable logic controller, the state vector is constructed: S(t)∈R d , where d represents the number of parameters of the state vector, that is, the state vector is: Where e(t) is the current step error, is the inverse of the error, U (t-h):(t-1) is the control sequence of the historical h steps, and e(t+h) is the error of the hth step in the future.
9. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 8, characterized in that: The specific process of generating the optimal control strategy is as follows: C1. First, generate actions based on the Actor network and input the state vector S(t) into the Actor network to obtain the optimal control strategy, that is, the output action: in Represents the deterministic action base value output after inputting the state s(t) into the Actor network, Indicates that the generated action is within the control volume of the actuator; C2. Critic evaluation and TD-error calculation are performed again. Critic evaluation calculates the current Q value based on the input state vector S(t) and action a(t): Q = Q(S(t), a(t)). The Q value is the action value function. The target network calculates the Q value based on the next state and action: Q′ = Q′(S′(t+1), μ′(S′(t+1))). Q′ is the action value function of the Critic target network, and μ′ is the policy function of the Actor target network. C3, then calculate TD-error, using the formula: TD-error = R t +γ*Q′-Q is used to calculate TD-error, where γ is learned through the model training process and represents the discount factor, R t is the immediate reward parameter, R(t) = -(o*tracking error + p*control energy consumption + q*parameter fluctuation penalty), where o, p, and q are the weight factors of tracking error, control energy consumption, and parameter fluctuation penalty, respectively; C4. Finally, perform the Critic network update, Actor network update, and target network soft update. The Critic network is updated, which is expressed by the expression: Minimize the mean square loss of TD-error; Update the Actor network, expressed as follows: Maximize the Q value and finally pass the expression: The target network is soft-updated to obtain the optimal control strategy, where τ is the model parameter.
10. The adaptive control optimization method based on deep learning and reinforcement learning according to claim 9, characterized in that: The trained spatiotemporal fusion deep learning model and reinforcement learning strategy network model are lightweight deployed to the edge controller to achieve multi-step prediction and real-time optimal control decision generation under the lack of sensor data, drive the actuator to achieve closed-loop control, and complete adaptive control optimization. The specific process is as follows: The trained spatiotemporal fusion model and reinforcement learning decision module are lightweight and deployed to the edge controller. A data buffer is established to collect and store pre-processed real-time data. The data buffer is then updated in real time. The oldest data in the buffer is deleted and newly collected data is added to ensure that the data in the buffer contains sensor data from the most recent H time points. The edge controller maintains a sliding window buffer, stores sensor data from the last H time steps, and dynamically preprocesses real-time data. Every time new data is received, the buffer is updated and input into the deep learning model. The spatiotemporal fusion model generates real-time predictions for the next h steps. The reinforcement learning decision module then generates the optimal control strategy in real time to drive the actuator, thereby achieving adaptive control optimization.
Citation Information
Patent Citations
Model-free adaptive control method and system for industrial system
CN113325721A