Digital twinborn calibration framework construction method based on Lyapunov strategy

By constructing a digital twin calibration framework based on the Liyapunov strategy, the problems of traditional methods dependence on labeled data and noise deviation are solved, and efficient training and strategy stability are achieved under labeled data, which is suitable for joint control of industrial robots.

CN120524985APending Publication Date: 2025-08-22CHINA YANGTZE POWER

Patent Information

Application Number
CN202510434941.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Traditional model calibration methods rely on a large amount of labeled data and are difficult to cope with sensor noise and model deviations. Existing reinforcement learning algorithms such as SAC lack stability guarantees, resulting in action oscillation and parameter divergence.

Method used

A digital twin calibration framework based on the Liyapunov strategy is built, and the dynamics of the physical system are simulated through the digital twin model, combined with specific network structures and optimization methods, efficient training of label-free data is achieved, and Lyapunov function constraints and dynamic smoothness optimization are introduced to ensure the stability of the strategy.

Benefits of technology

Efficient training of labelless data is achieved in high noise and deviation environments, ensuring the stability of the strategy and multiple constraint requirements, and is suitable for scenarios such as industrial robot joint control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524985A_ABST
    Figure CN120524985A_ABST
Patent Text Reader

Abstract

The invention discloses a digital twin calibration framework construction method based on a Lyapunov strategy, and the method comprises the steps: constructing a digital twin model containing a proxy network to simulate the dynamic state of a system, and defining a state, an action space and a reward function through a Markov decision process. A constrained Lyapunov action-commentator (CLAC) algorithm is introduced, a strategy network and a Lyapunov network are optimized, and the algorithm can stably act under high noise and deviation. Real-time parameter optimization is realized by means of single-time neural network forward propagation, an experience playback pool and the like. Each network structure is clear, and a specific initialization and optimization method is adopted. According to the method, a calibration problem can be converted into a parameter tracking task, efficient training is performed under unmarked data, constraint requirements such as stability can be met, and the method is suitable for scenes such as industrial robot joint control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and in particular relates to a method for constructing a digital twin calibration framework based on the Lyapunov strategy. Background Art

[0002] Traditional model calibration methods rely on large amounts of labeled data and struggle to cope with sensor noise and model bias. Existing reinforcement learning algorithms (such as SAC), while capable of unsupervised learning, lack stability guarantees, leading to motion oscillation and parameter divergence. This paper addresses these issues by introducing Lyapunov function constraints and dynamic smoothness optimization. This paper proposes a real-time calibration framework specifically for industrial scenarios, which is urgently needed. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method for constructing a digital twin calibration framework based on the Lyapunov strategy. The present invention constructs a digital twin calibration framework and uses an algorithm based on the Lyapunov strategy to convert calibration into a tracking task. Combined with a specific network structure and optimization method, efficient training of unlabeled data is achieved in a high-noise and deviation environment, ensuring the stability of the strategy and meeting various constraint requirements. It can be applied to scenarios such as industrial robot joint control. In order to solve the above technical problems, the technical solution adopted by the present invention is: A method for constructing a digital twin calibration framework based on the Lyapunov strategy, the steps are as follows: S1. Construct a digital twin model and a dynamic parameter real-time calibration module: The dynamic behavior of the physical system is simulated by the digital twin model. The model includes a proxy network of the dynamic model and uses a multilayer perceptron to approximate the system dynamics. The total number of layers is 4, including 3 hidden layers and 1 output layer. Each hidden layer contains 100 neurons. The ReLU activation function is uniformly used in the hidden layer. The dimension of the output layer is consistent with the sensor reading vector, which is recorded as , the dimension is ,Right now , the output layer does not use other compression or scaling operations, uses the identity mapping as the "activation function" of the output layer, and directly outputs the predicted value of the sensor reading; S2. Define the Markov decision process: the state space is output by the current model , actual sensor measurement value and environmental conditions Composition, that is ; The action space is the dynamic parameter to be calibrated ,Right now ; Reward function Based on model output With measured value Dynamic calculation of matching error; S3. Design an ActorCritic algorithm based on Lyapunov constraints: Parameters and network initialization: Input for Lyapunov network The learning rate , policy network The learning rate and Lagrange multipliers 、 、 Parameters; Randomly initialize the parameters of the Lyapunov network and the policy network, respectively and ; Set the target network parameters corresponding to the Lyapunov network to be the same as the current network, that is ; Main loop: distributed from the environment state Initial state of sampling , execute subsequent time step cycles and update cycles for this initial state; Time step loop: At each time step , using the current policy network From the status Medium sampling action , apply it to the environment and obtain the state of the next moment and rewards , and the quadruple Recorded to experience replay pool middle; Update loop: From the experience replay pool Randomly sample small batches of transfer samples in the Lyapunov network , policy network , Lagrange multipliers 、 and Perform gradient updates; adopt soft update strategy for target network parameters ,in is the soft update coefficient; in CLAC, the objective function has an additional term that aims to make the policy network in a given similar or close state ( ) to obtain similar optimal actions, the objective function is defined as follows: , in, is a positive Lagrange multiplier; Repeat process: Repeat the above steps until the preset training termination condition is reached; S4. Real-time parameter optimization and training: Real-time parameter inference through a single neural network forward propagation ; An experience replay pool is used to store interaction data, and the network weights are updated by combining mini-batch stochastic gradient descent with the Adam optimizer. The learning rate and soft update factor are configured according to the preset hyperparameters. The Xavier initialization method is used to initialize the weights of the policy network and the Lyapunov network, and the bias term is initialized to zero or a minimum value.

[0004] Preferably, the policy network: The input layer receives the current environment state vector , the dimension of the state vector depends on the specific application; Hidden layer 1 uses a fully connected multilayer perceptron structure with 256 neurons. It uses learnable weights and bias terms to perform linear transformation and nonlinear activation on the input. The activation function can be Leaky ReLU or other compatible functions. Hidden layer 2 is a fully connected structure with 256 neurons, using Leaky ReLU or other activation functions, and the output is passed to the output layer; The output layer outputs the mean of the Gaussian distribution and standard deviation , and post-processed with reversible compression function, through the formula Constrain the output to range, where is Gaussian noise; the objective function of the policy network is ,in is the action smoothness constraint weight, is the dynamic Lagrange multiplier.

[0005] Preferably, the Lyapunov network: The input layer receives the state vector and motion vector , and concatenate the two into a comprehensive input vector; Hidden layer 1 is a fully connected structure with 256 neurons and the activation function is Leaky ReLU; Hidden layer 2 is a fully connected structure with 256 neurons and uses Leaky ReLU or similar activation function; The output layer outputs the Lyapunov function value , used to evaluate policy stability; based on the maximum entropy action critic framework, Lyapunov function As the critic in the policy gradient formulation, The objective function is defined as ,in yes The approaching goal, The structure and Same, through hyperparameters The exponential weighted average of the control is used to update its parameters, and its target parameters are updated through the soft update strategy renew, is the soft update coefficient.

[0006] Preferably, the dimension of the proxy network output layer of the dynamic model is consistent with the sensor reading vector, and the physical quantity prediction value is directly generated using identity mapping, and the training data is generated through real-time interaction with the physical system through online simulation, without the need for pre-labeling of the data set. Preferably, the hidden layer of the strategy network adopts a Leaky ReLU activation function with a negative interval slope of 0.01. Preferably, the experience replay pool sampling mechanism is priority experience replay, and the sample weight is dynamically adjusted according to the temporal difference error. Preferably, the Gaussian noise input The variance of is dynamically updated through an adaptive adjustment mechanism to match the ambient noise level. Preferably, the objective function: Dynamic Lagrange multipliers in Online adjustments are made through gradient descent to balance policy performance and constraint satisfaction. Preferably, when optimizing the policy network and Lyapunov network, the weights and bias terms are iteratively updated using a combination of mini-batch stochastic gradient descent and the Adam algorithm. The target entropy of the policy network is set according to the soft actor critic results, and the remaining hyperparameters are adjusted according to practical requirements. Preferably, the digital twin calibration framework is applied to industrial robot joint control, including joint angle calibration, torque compensation and real-time optimization of dynamic friction parameters.

[0007] The present invention can achieve the following beneficial effects: The system dynamics are simulated via an agent network, and the calibration problem is converted into a parameter tracking task.

[0008] Stability is evaluated via a Lyapunov network, and motion smoothness constraints are introduced into the objective function.

[0009] Single forward propagation, experience replay pooling and adaptive noise processing are used to achieve efficient training with unlabeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The present invention will be further described below with reference to the accompanying drawings and examples: Figure 1 This is a flow chart of the model calibration of the present invention.

[0011] In the figure: Model Parameters: model parameters; Action: action; Actor: executor; Reinforcement learning in simulation: reinforcement learning in simulation; State: status; Policy Network: Policy Network; Action: action; Real Process: Actual process; Physics-based System Model: Physics-based system model; Deep Neural Network: Deep neural network. DETAILED DESCRIPTION

[0012] The preferred solution is Figure 1 As shown in FIG, a method for constructing a digital twin calibration framework based on the Lyapunov strategy includes the following steps: S1. Build a digital twin model and dynamic parameter real-time calibration module: The dynamic behavior of the physical system is simulated by the digital twin model. The model contains a proxy network of the dynamic model and uses a multi-layer perceptron (MLP) to approximate the system dynamics. The total number of layers is 4 ( ), which contains 3 hidden layers and 1 output layer, and each hidden layer contains 100 neurons (i.e. ). The ReLU activation function (or a ReLU variant similar to the previous network) is uniformly used in the hidden layer, and the dimension of the output layer is consistent with the sensor reading vector, denoted as , the dimension is ,Right now , the output layer does not use other compression or scaling operations, uses identity mapping (Identity) as the "activation function" of the output layer, and directly outputs the predicted value of the sensor reading.

[0013] S2. Define the Markov decision process: the state space is output by the current model , actual sensor measurement value and environmental conditions Composition, that is ; The action space is the dynamic parameter to be calibrated ,Right now ; Reward function Based on model output With measured value Dynamic calculation of matching error.

[0014] S3. Design of the ActorCritic (CLAC) algorithm based on Lyapunov constraints: In situations with high sensor noise and simulator bias (i.e., irreducible reality gap), the policy network may exhibit large variance due to an incomplete representation of the system model. This situation is undesirable in many real-world applications where obtaining stable or smooth action variations over time is crucial.

[0015] In order to stabilize the action, the constrained Lyapunov action critic (CLAC) algorithm is introduced, which is an improvement to the Lyapunov-based action critic (LAC) and significantly improves the stability of the action under model uncertainty and sensor noise.

[0016] Parameters and network initialization: Input for Lyapunov network The learning rate , policy network The learning rate and Lagrange multipliers 、 、 etc.; randomly initialize the parameters of the Lyapunov network and the policy network, respectively and ; Set the target network parameters corresponding to the Lyapunov network to be the same as the current network, that is .

[0017] Main loop: distributed from the environment state Initial state of sampling , subsequent time step cycles and update cycles are performed on this initial state.

[0018] Time step loop: At each time step , using the current policy network From the status Medium sampling action , apply it to the environment and obtain the state of the next moment and rewards , and the quadruple Recorded to experience replay pool middle.

[0019] Update loop: From the experience replay pool Randomly sample small batches of transfer samples in the Lyapunov network , policy network , Lagrange multipliers 、 and Perform gradient updates; adopt soft update strategy for target network parameters ,in is the soft update coefficient. In CLAC, the objective function has an additional term that aims to make the policy network ) to obtain similar optimal actions, the objective function is defined as follows: , in, is a positive Lagrange multiplier.

[0020] Repeat process: Repeat the above steps until the preset training termination condition is reached.

[0021] S4. Real-time parameter optimization and training: Real-time parameter inference through a single neural network forward propagation ; An experience replay pool is used to store interaction data, and the network weights are updated by combining mini-batch stochastic gradient descent with the Adam optimizer. The learning rate and soft update factor are configured according to the preset hyperparameters. The Xavier initialization method is used to initialize the weights of the policy network and the Lyapunov network, and the bias term is initialized to zero or a minimum value.

[0022] Furthermore, the policy network: The input layer receives the current environment state vector , the dimension of the state vector depends on the specific application.

[0023] Hidden layer 1 uses a fully connected multilayer perceptron structure with 256 neurons. It uses learnable weights and bias terms to perform linear transformation and nonlinear activation on the input. The activation function can be Leaky ReLU (leaky rectified linear unit) or other compatible functions.

[0024] Hidden layer 2 is a fully connected structure with 256 neurons, using Leaky ReLU or other activation functions, and the output is passed to the output layer.

[0025] The output layer outputs the mean of the Gaussian distribution and standard deviation , and post-processed with reversible compression function, through the formula Constrain the output to range, where is Gaussian noise. The objective function of the policy network is ,in is the action smoothness constraint weight, is the dynamic Lagrange multiplier.

[0026] Furthermore, the Lyapunov network: The input layer receives the state vector and motion vector , and concatenate the two into a comprehensive input vector.

[0027] Hidden layer 1 is a fully connected structure with 256 neurons and the activation function is Leaky ReLU.

[0028] Hidden layer 2 is a fully connected structure with 256 neurons and uses Leaky ReLU or similar activation functions.

[0029] The output layer outputs the Lyapunov function value , used to evaluate policy stability. Based on the maximum entropy action critic framework, the Lyapunov function As the critic in the policy gradient formulation, The objective function is defined as ,in yes The approaching goal, The structure and Same, through hyperparameters The exponential weighted average of the control is used to update its parameters, and its target parameters are updated through the soft update strategy renew, is the soft update coefficient ( ).

[0030] Furthermore, the output layer dimension of the proxy network of the dynamic model is consistent with the sensor reading vector, and the physical quantity prediction value is directly generated using identity mapping. The training data is generated through real-time interaction with the physical system through online simulation, without the need for pre-labeled data sets. Furthermore, the hidden layer of the policy network adopts a Leaky ReLU activation function with a negative interval slope of 0.01. Furthermore, the experience replay pool sampling mechanism is priority experience replay, and the sample weight is dynamically adjusted according to the temporal difference error. Furthermore, the Gaussian noise input The variance of is dynamically updated through an adaptive adjustment mechanism to match the ambient noise level. Furthermore, the objective function: Dynamic Lagrange multipliers in Online adjustments are made through gradient descent to balance policy performance and constraint satisfaction. Furthermore, when optimizing the policy network and Lyapunov network, the weights and bias terms are iteratively updated using a combination of minibatch stochastic gradient descent (SGD) and the Adam algorithm. The target entropy of the policy network is set based on the results of the soft actor critic (SAC), and the remaining hyperparameters are adjusted according to practical requirements. The application of the digital twin calibration framework construction method based on the Lyapunov strategy is applied to the joint control of industrial robots, including joint angle calibration, torque compensation and real-time optimization of dynamic friction parameters.

[0031] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A method for constructing a digital twin calibration framework based on the Lyapunov strategy, characterized in that The following steps are involved: S1. Construct a digital twin model and a dynamic parameter real-time calibration module: The dynamic behavior of the physical system is simulated by the digital twin model. The model includes a proxy network of the dynamic model and uses a multilayer perceptron to approximate the system dynamics. The total number of layers is 4, including 3 hidden layers and 1 output layer. Each hidden layer contains 100 neurons. The ReLU activation function is uniformly used in the hidden layer. The dimension of the output layer is consistent with the sensor reading vector, which is recorded as , the dimension is ,Right now , the output layer does not use other compression or scaling operations, uses the identity mapping as the "activation function" of the output layer, and directly outputs the predicted value of the sensor reading; S2. Define the Markov decision process: the state space is output by the current model , actual sensor measurement value and environmental conditions Composition, that is ; The action space is the dynamic parameters to be calibrated ,Right now ; Reward function Based on model output With measured value Dynamic calculation of matching error; S3. Design an ActorCritic algorithm based on Lyapunov constraints: Parameters and network initialization: Input for Lyapunov network Learning rate , policy network Learning rate and Lagrange multipliers 、 、 Parameters; Randomly initialize the parameters of the Lyapunov network and the policy network, respectively and ; Set the target network parameters corresponding to the Lyapunov network to be the same as the current network, that is ; Main loop: distributed from the environment state Initial state of sampling , execute subsequent time step cycles and update cycles for this initial state; Time step loop: At each time step , using the current policy network From the status Medium sampling action , apply it to the environment and obtain the state of the next moment and rewards , and the quadruple Recorded to experience replay pool middle; Update loop: From the experience replay pool Randomly sample small batches of transfer samples in the Lyapunov network , policy network , Lagrange multipliers 、 and Perform gradient updates; adopt soft update strategy for target network parameters ,in is the soft update coefficient; in CLAC, the objective function has an additional term that aims to make the policy network in a given similar or close state ( ) to obtain similar optimal actions, the objective function is defined as follows: , in, is a positive Lagrange multiplier; Repeat process: Repeat the above steps until the preset training termination condition is reached; S4. Real-time parameter optimization and training: Real-time parameter inference through a single neural network forward propagation ; An experience replay pool is used to store interaction data, and the network weights are updated by combining mini-batch stochastic gradient descent with the Adam optimizer. The learning rate and soft update factor are configured according to the preset hyperparameters. The Xavier initialization method is used to initialize the weights of the policy network and the Lyapunov network, and the bias term is initialized to zero or a minimum value.

2. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1 is characterized in that: The policy network: The input layer receives the current environment state vector , the dimension of the state vector depends on the specific application; Hidden layer 1 uses a fully connected multilayer perceptron structure with 256 neurons. It uses learnable weights and bias terms to perform linear transformation and nonlinear activation on the input. The activation function can be Leaky ReLU or other compatible functions. Hidden layer 2 is a fully connected structure with 256 neurons, using Leaky ReLU or other activation functions, and the output is passed to the output layer; The output layer outputs the mean of the Gaussian distribution and standard deviation , and post-processed with reversible compression function, through the formula Constrain the output to range, where is Gaussian noise; the objective function of the policy network is ,in is the action smoothness constraint weight, is the dynamic Lagrange multiplier.

3. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1 is characterized in that: The Lyapunov network: The input layer receives the state vector and motion vector , and concatenate the two into a comprehensive input vector; Hidden layer 1 is a fully connected structure with 256 neurons and the activation function is Leaky ReLU; Hidden layer 2 is a fully connected structure with 256 neurons and uses Leaky ReLU or similar activation function; The output layer outputs the Lyapunov function value , used to evaluate strategy stability; Based on the maximum entropy action critic framework, Lyapunov function As the critic in the policy gradient formulation, The objective function is defined as ,in yes The approaching goal, The structure and Same, through hyperparameters The exponential weighted average of the control is used to update its parameters, and its target parameters are updated through the soft update strategy renew, is the soft update coefficient.

4. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1 is characterized in that: The output layer dimension of the proxy network of the dynamic model is consistent with the sensor reading vector, and the physical quantity prediction value is directly generated using identity mapping. The training data is generated through real-time interaction with the physical system through online simulation, without the need for pre-labeled data sets.

5. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1, characterized in that: The hidden layer of the policy network uses the Leaky ReLU activation function with a negative slope of 0.

01.

6. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1, characterized in that: The experience replay pool sampling mechanism is priority experience replay, and the sample weight is dynamically adjusted according to the temporal difference error.

7. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 2, characterized in that: The Gaussian noise input The variance of is dynamically updated through an adaptive adjustment mechanism to match the ambient noise level.

8. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1, characterized in that: The objective function: Dynamic Lagrange multipliers in Online adjustments are made through gradient descent to balance policy performance and constraint satisfaction.

9. The method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1, characterized in that: When optimizing the policy network and Lyapunov network, the weights and bias terms are iteratively updated using a combination of mini-batch stochastic gradient descent and the Adam algorithm. The target entropy of the policy network is set based on the soft actor-critic results, and the remaining hyperparameters are adjusted according to practical requirements.

10. The application of the method for constructing a digital twin calibration framework based on the Lyapunov strategy according to claim 1, characterized in that: This digital twin calibration framework is applied to industrial robot joint control, including joint angle calibration, torque compensation, and real-time optimization of dynamic friction parameters.

Citation Information

Patent Citations

  • Information physical system safety control method based on deep reinforcement learning

    CN113885330A

  • Floating base space manipulator tail end position control method based on model-free reinforcement learning

    CN116442235A

  • Apparatus for indicating the power state of a control box

    KR1020240164482A

Cited By

  • Self-adaptive batch processing maximum correlation entropy orbit determination method and device

    CN121881687A

  • Adaptive batch maximum correlation entropy orbit determination method and device

    CN121881687B

  • Humanoid robot motion control method and system based on stability constraint

    CN122185248A

  • A method and system for motion control of humanoid robots based on stability constraints

    CN122185248B