A dynamic optimization method, system, device and medium for a digital twin system

By combining autoencoder network and deep reinforcement learning in the digital twin system, the Koopman operator theory and MBRL algorithm are used to realize dynamic optimization of complex equipment, solving the problem of insufficient self-evolution capabilities of models in the existing technology, and improving the model accuracy and intelligence level.

CN118586279BActive Publication Date: 2025-08-22BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410707551.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-03
Publication Date
2025-08-22
Estimated Expiration
2044-06-03

AI Technical Summary

Technical Problem

The existing digital twin technology lacks self-evolution ability in the dynamic optimization of complex equipment, is difficult to effectively resist uncertain disturbances, poor model accuracy and interpretability, and is difficult to build a robust adaptive controller, time-consuming and labor-intensive.

Method used

The digital twin model based on the autoencoder network and the decision-making intelligent model based on the deep reinforcement learning framework are adopted to interact virtually and real through the edge-side computing platform, and combined with the data-driven Koopman operator theory and MBRL algorithm, the dynamic evolution of the model and parameter optimization are achieved.

Benefits of technology

It improves the interpretability and generalization capabilities of the model, realizes real-time consistency between the digital twin model and the equipment entity, and enhances the intelligent control and optimization capabilities of complex equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118586279B_ABST
    Figure CN118586279B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of digital twin technology and discloses a dynamic optimization method, system, equipment and medium for a digital twin system, including: constructing a digital twin model and a decision-making intelligent body model and deploying them, performing virtual-reality interaction with the equipment entity through the digital twin model to obtain interaction data; performing data interaction between the decision-making intelligent body model and the digital twin model to obtain real-time simulation data; and storing the real-time simulation data and the interaction data in a database; performing credibility evaluation and dynamic evolution on the digital twin model, extracting real-time simulation data and interaction data from the database according to a set ratio, and optimizing the parameters of the digital twin system based on the extracted data to obtain a digital twin system with updated parameters. The technical solution of the present invention can continuously and automatically update the control optimization system model, thereby improving the intelligent control and optimization capabilities of complex unmanned equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital twin technology, and in particular relates to a dynamic optimization method, system, device and medium for a digital twin system. Background Art

[0002] Equipment operation, maintenance, and control based on digital twins are highly relevant, not least because the equipment itself and its operating environment are inherently highly dynamic and full of uncertainty. For example, wear and corrosion can alter the dynamic behavior of internal components, while changes in electromagnetic field, temperature, and airflow in the operating environment can also cause fluctuations in equipment performance. Traditional models of equipment or their environments fail to fully account for these dynamics and uncertainties, negatively impacting the prediction and control of equipment status. Digital twins, however, integrate model construction and usage, evolving in real time and offering the potential for new performance improvements in equipment operation, maintenance, control, and optimization.

[0003] The existing digital twin-based unmanned equipment decision-making optimization solutions in the field can be roughly divided into two types. One is to design a deep neural network model to predict the future behavior pattern or health status of the equipment, and then manually design the corresponding optimization solution and feed it back to the physical equipment to complete the closed loop from virtual space to the physical world; the other is to build a reduced-order mechanism model of complex equipment, generally various linear models, and then design a more robust and adaptive optimization control solution, which is then deployed to the equipment entity. Using deep reinforcement learning methods, the reduced-order mechanism model is used as the environment for intelligent agent interaction, and then a more robust optimization control solution is designed, which also falls into the category of the second technical solution.

[0004] In summary, the current field lacks a systematic theory of dynamic optimization methods and complete technical solutions based on self-evolution of equipment digital twins. Specifically, according to existing technical solutions in the field, there are the following shortcomings:

[0005] (1) The current digital twin modeling methods designed using various types of time-series deep neural networks for complex equipment behavior monitoring and health monitoring do not have the ability to use real-time data to achieve model self-evolution, and have poor resistance to various uncertain disturbances, which seriously affects the accuracy of the model. In addition, the model has poor interpretability and ultimately must achieve feedback through artificial loops, with a low degree of intelligence.

[0006] (2) There are also bottlenecks in the scheme of using reduced-order mechanism models to construct robust or adaptive controllers. As the complexity of equipment increases, the difficulty of constructing its mechanism model increases greatly. The reduced-order model will inevitably lead to a loss of accuracy. In addition, the design of schemes based on mechanism models requires the uncertainty of model parameters to be pre-determined in advance, and then through complex manual mathematical derivation. However, the uncertainty of complex equipment and working conditions is high and difficult to predict in advance, and redesigning the optimized control scheme is time-consuming and labor-intensive. Summary of the Invention

[0007] The purpose of the present invention is to provide a dynamic optimization method, system, device and medium for a digital twin system to solve the problems existing in the above-mentioned prior art.

[0008] To achieve the above objectives, the present invention provides a dynamic optimization method for a digital twin system, comprising:

[0009] Constructing a digital twin system in the cloud, wherein the digital twin system includes a digital twin model and a decision-making agent model, wherein the digital twin model is constructed based on a neural network with an autoencoder network structure, and the decision-making agent model is constructed based on a deep reinforcement learning framework;

[0010] Deploy the digital twin system on the edge computing platform, perform virtual-reality interaction between the digital twin model and the equipment entity to obtain interaction data; perform data interaction between the decision-making intelligent agent model and the digital twin model to obtain real-time simulation data; and store the real-time simulation data and the interaction data in a database;

[0011] Comparing the interaction data of the digital twin model with the measured data of the equipment entity according to a set period to obtain the credibility data of the digital twin model;

[0012] The credibility data is compared with a preset evolution threshold, dynamic evolution is performed based on the comparison result, the real-time simulation data and the interaction data are extracted from the database according to a set ratio, parameters of the digital twin system are optimized based on the extracted data, and a digital twin system with updated parameters is obtained. The digital twin system with updated parameters is used to perform virtual-reality interaction with the equipment entity.

[0013] Optionally, the constraints of the digital twin model include:

[0014] The first type of loss function L1:

[0015]

[0016] The second type of loss function L2:

[0017]

[0018] The third type of loss function L3:

[0019]

[0020] Where x k is the state vector, z k is a high-dimensional hidden state; is the encoder network, is the decoder network, K x K is the high-dimensional space Koopman linear operator for the system’s own state evolution. u is the high-dimensional space Koopman linear operator for external control input, u k For external input.

[0021] Optionally, the reward function of the decision agent model includes:

[0022] r=-||x k -x r ||2

[0023] Where r is the reward function of the decision-making agent model, x k is the state vector, x r The state that complex equipment is expected to achieve.

[0024] Optionally, the optimization objective function of the decision agent model includes:

[0025]

[0026] In the formula, π represents the strategy function learned by the agent, s k is the state vector, a k is the action vector, E() is the expected function, α is the set weight, H is the entropy function of the strategy, a is the distribution that the agent obeys in the current strategy, ρ π is the action-state distribution under the current strategy.

[0027] Optionally, performing parameter optimization on the digital twin system based on the extracted data to obtain a digital twin system with updated parameters specifically includes:

[0028] Dynamically evolving the digital twin model in the cloud based on the extracted data to obtain an updated digital twin model;

[0029] In the cloud, the decision-making agent model is trained and updated based on the updated digital twin model combined with the MBRL algorithm and continuous learning theory to obtain an updated decision-making agent model;

[0030] A parameter-updated digital twin system is constructed based on the model parameters of the updated digital twin model and the updated decision-making intelligent agent model, and the parameter-updated digital twin system is deployed in the edge computing platform.

[0031] A dynamic optimization system for a digital twin system, comprising:

[0032] A model construction module for constructing a digital twin system in the cloud, wherein the digital twin system includes a digital twin model and a decision-making agent model, the digital twin model is constructed based on a neural network with an autoencoder network structure, and the decision-making agent model is constructed based on a deep reinforcement learning framework;

[0033] A virtual-reality interaction module is used to deploy the digital twin system on the edge computing platform, perform virtual-reality interaction between the digital twin model and the equipment entity to obtain interaction data; perform data interaction between the decision-making intelligent agent model and the digital twin model to obtain real-time simulation data; and store the real-time simulation data and the interaction data in a database;

[0034] A credibility evaluation module is used to compare the interaction data of the digital twin model with the measured data of the equipment entity according to a set period to obtain credibility data of the digital twin model;

[0035] A model optimization module is used to compare the credibility data with a preset evolution threshold, perform dynamic evolution based on the comparison result, extract the real-time simulation data and the interaction data from the database according to a set ratio, optimize the parameters of the digital twin system based on the extracted data, obtain the digital twin system with updated parameters, and perform virtual-reality interaction with the equipment entity based on the digital twin system with updated parameters.

[0036] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the dynamic optimization method of a digital twin system.

[0037] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the dynamic optimization method of a digital twin system.

[0038] The technical effects of the present invention are:

[0039] The present invention adopts the Koopman operator theory method of continuous learning and data-driven to construct a dynamic behavior model of complex equipment. On the one hand, it alleviates the dilemma that the existing mechanism model cannot accurately represent complex behaviors. On the other hand, compared with the traditional deep neural network model, it greatly enhances the interpretability and generalization. On the other hand, through continuous learning, the constructed digital twin model can use real-time sensor data to continuously learn the dynamic changes of the equipment itself, so that the digital twin model remains consistent with the entity throughout its life cycle.

[0040] The present invention combines deep reinforcement learning technology with the constructed equipment digital twin model to form a complete equipment digital twin system, realizing a complete closed loop of online evolution of the digital twin model (from real to virtual) - online intelligent decision-making of the intelligent agent (from virtual to real), which can greatly improve the intelligent control and optimization capabilities of complex unmanned equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0043] Figure 1 Schematic diagram of the Koopman characteristic function and operator identification structure based on the autoencoder deep network in an embodiment of the present invention;

[0044] Figure 2 This is an optimization flow chart in an embodiment of the present invention;

[0045] Figure 3 Schematic diagram of the structure of the hardware components and software components in the embodiment of the present invention. DETAILED DESCRIPTION

[0046] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0047] It should be understood that the terms described herein are intended only to describe particular embodiments and are not intended to limit the present invention. In addition, for numerical ranges herein, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each smaller range between any intermediate value within a stated value or stated range and any other stated value or intermediate value within the stated range is also encompassed by the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded within the scope.

[0048] Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. Although only preferred methods are described herein, any method similar or equivalent to that described herein may also be used in the practice or testing of the present invention. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods associated with the documents. In the event of any conflict with any incorporated document, the contents of this specification shall prevail.

[0049] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present invention without departing from the scope or spirit of the invention. Other embodiments will be apparent to those skilled in the art from the present invention. The present description and examples are intended to be illustrative only.

[0050] The words “include,” “including,” “have,” “contain,” etc. used in this article are open-ended terms, meaning including but not limited to.

[0051] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0052] Example 1

[0053] like Figure 1 - Figure 3As shown, this embodiment provides a dynamic optimization method for a digital twin system, including: constructing a digital twin system in the cloud, wherein the digital twin system includes a digital twin model and a decision-making agent model, the digital twin model is constructed based on a neural network with an autoencoder network structure, and the decision-making agent model is constructed based on a deep reinforcement learning framework; deploying the digital twin system on an edge computing platform, performing virtual-real interaction with an equipment entity through the digital twin model to obtain interaction data; performing data interaction between the decision-making agent model and the digital twin model to obtain real-time simulation data; and The real-time simulation data and the interaction data are stored in a database; the interaction data of the digital twin model is compared with the measured data of the equipment entity according to a set period to obtain the credibility data of the digital twin model; the credibility data is compared with a preset evolution threshold, and dynamic evolution is performed based on the comparison result. The real-time simulation data and the interaction data are extracted from the database according to a set ratio, and the parameters of the digital twin system are optimized based on the extracted data to obtain the digital twin system with updated parameters, and virtual-reality interaction is performed with the equipment entity based on the digital twin system with updated parameters.

[0054] Credibility refers to the simulation credibility of the digital twin model. In the Koopman data-driven model described in this embodiment, its predicted data is compared with the measured data. If the error exceeds a given threshold, an evolution is triggered. The credibility calculation method and threshold are set according to the specific task. The dynamic evolution mainly includes two update links: one is the update of the digital twin model, and the other is the update of the decision-making intelligent agent based on the updated digital twin model.

[0055] When the evolution condition is triggered, a Koopman model based on an autoencoder network is incrementally trained online based on measured sensor data. The number of iterations is set to E, completing the dynamic evolution of the digital twin model and maintaining the model's credibility. Then, based on the more credible model after evolution, the agent interacts with it in virtual space, enabling the agent to adapt to changes in the physical equipment and improve its decision-making optimization capabilities through continuous learning. Specifically, n samples are randomly selected from the actual interaction dataset Dr as the initial state of the equipment. The current decision-making agent then performs control optimization on the equipment twin in virtual space, interacting with the equipment digital twin model (the Koopman autoencoder network) to achieve a j-step prediction of the future state of the equipment under the decision-making agent's control. These virtual simulated interaction data are then stored in Dp. Finally, the interaction data from the real dataset Dr and the simulated dataset Dp are mixed in a certain ratio and used to continuously train the new agent policy. The number of training iterations is set to F. The number of E and F should be determined based on the specific problem and is positively correlated with the complexity of the task.

[0056] This technical embodiment is applicable to remote action optimization and decision-making of complex equipment (such as industrial robotic arms, unmanned vehicles, unmanned boats, drones, etc.), and is aimed at macro-decision-making and control problems of equipment, such as robotic arm motion trajectory control, unmanned equipment trajectory planning, cluster formation control and other scenarios. It is not yet suitable for underlying field-level control scenarios with high real-time requirements.

[0057] This embodiment designs a method and system for constructing a digital twin system for complex unmanned equipment based on data-driven Koopman operator theory and model-based reinforcement learning methods. The digital twin system has the ability to improve its own credibility by continuously and dynamically evolving according to real-time sensor data from equipment entities, and can continuously train deep reinforcement learning agents based on more reliable digital twin models, continuously improving the decision-making optimization capabilities of the agents.

[0058] Composition description: This embodiment consists of two parts, hardware and software. The hardware part includes the equipment physical entity, sensors for collecting physical equipment status data, edge communication equipment, data caching devices, and high-performance computing platforms. The software part includes equipment digital twins, decision-making agents, and supporting software tools. The twin is composed of a neural network with an autoencoder network structure. The data-driven model is used to approximate the Koopman linear operator and the corresponding observation function, thereby characterizing the behavior pattern of the physical equipment. The decision-making agent is composed of an agent supported by the deep reinforcement learning algorithm - the SAC algorithm. The equipment digital twin, the decision-making agent, and the actual physical equipment together constitute a complete dynamically self-evolving and dynamically optimized equipment digital twin system.

[0059] Structural Description: As shown in the accompanying drawings, the hardware components of this embodiment are mainly composed of a cloud-based high-performance computing platform and an edge-side computing platform. Various sensors and actuators deployed on complex equipment entities are responsible for generating virtual-reality interactions with the digital twin. The edge-side computing platform is close to the complex equipment entity and is used to deploy the digital twin system and support its operation. Its software part mainly consists of a data cache and preprocessing module component for collecting real-time data streams from sensors and performing basic data cleaning, time alignment, noise reduction and other preprocessing; an online evaluation and evolution trigger module for performing a trustworthy evaluation of the simulation performance of the digital twin and determining whether to trigger dynamic evolution operation instructions; a deployment environment for the digital twin model constructed based on the data-driven Koopman operator for the normal operation and solution of the digital twin model; and a twin decision-making intelligent agent deployment environment based on deep reinforcement learning for the normal operation and solution of the designed deep reinforcement learning model. It consists of four major parts.

[0060] The cloud-based high-performance computing platform is mainly responsible for the development and training of the digital twin model based on the data-driven Koopman operator theory and the corresponding deep reinforcement learning model based on MBRL theory, as well as the realization of the dynamic evolution function of the digital twin system. Its software part is mainly composed of three parts: a database system for the storage and call of historical data, real-time batch data, and streaming data; a deep learning development and training support environment; and a twin evolution algorithm engine based on continuous learning.

[0061] Specifically, the implementation process and logic of this embodiment are as follows: a neural network with an autoencoder network structure is used to approximate the Koopman linear operator and the corresponding observation function. The data-driven model is used to characterize and predict the dynamic behavior of complex equipment. For example, under a given input, the dynamic output response of the system is predicted. The data-driven model adopts a combination of online and offline training methods. In the offline stage, the neural network is trained based on a large amount of historical data of the equipment. In the online stage, the neural network is incrementally updated using real-time sensor data from the equipment, so that the digital twin model can achieve online dynamic evolution and make more accurate state and behavior predictions, generating a large amount of simulation data for driving subsequent feedback optimization and control based on the MBRL algorithm. The MBRL algorithm is used for feedback control of complex nonlinear systems in the equipment digital twin scenario of this embodiment. The constructed equipment digital twin model is equivalent to the environmental model in the MBRL context. The intelligent agent interacts with the digital twin model in the virtual space to generate a large amount of simulation data. These real-time simulation data and the interaction data between the intelligent agent and the equipment are stored in the data buffer together. At fixed intervals, the intelligent agent extracts these two types of data in a certain proportion for self-learning, so that the intelligent agent can quickly adapt to changes in equipment status or performance, realize continuous evolution of control performance, and thus obtain better feedback control effects in the actual physical space.

[0062] Among them, the data-driven Koopman operator theory proposed in this embodiment constructs a digital twin black box model. Based on the deep autoencoder network model, it considers the external control input of the equipment and adopts the following method: Figure 1 The neural network structure shown.

[0063] in, The encoder network is used to approximate the observation function in the Koopman operator and convert the low-dimensional state vector x k Mapping to high-dimensional space into z k , the nonlinear dynamic process is approximately converted into a linear process, which is approximated by a deep neural network of appropriate size in the present invention, For the decoder network, the hidden state vector z in the high-dimensional space is k+1 Restored to a low-dimensional space state vector x k+1 , which is also approximated by a deep neural network and has the same size as the encoder network. K x K is the high-dimensional space Koopman linear operator for the system’s own state evolution. u is the high-dimensional space Koopman linear operator for external control input, K x With K u are approximated by a linear layer network of appropriate size. The encoded x k , that is, z k, and the external input u k , in high-dimensional space, each is driven by its own Koopman linear operator and linearly transferred to the next state, K x ·z k +K u ·u k , which passes through the decoder network That is, inverse transform the observation function and convert it into a finite-dimensional space, that is, we get x k+1 .

[0064]

[0065] According to Koopman operator theory, deep neural networks should try to meet the following constraints

[0066] (1) First-class loss function - reconstruction error

[0067]

[0068] This type of loss function is designed to guide the autoencoder network to identify an inherent coordinate transformation, which is used to transform the low-dimensional state vector x k Map to high-dimensional space, and then transform z through the corresponding inverse transformation k Restore to x k .

[0069] (2) Second type of loss function - prediction error

[0070]

[0071] This type of loss function is designed to guide the neural network so that it drives the high-dimensional hidden state to evolve forward under the stimulation of external input.

[0072] (3) The third type of loss function - linearization error

[0073]

[0074] This type of loss is derived from the high-dimensional hidden state z k angle, guiding the approximation of a linear Koopman operator.

[0075] Without loss of generality, this embodiment describes the decision optimization problem of complex equipment as follows:

[0076]

[0077] stx k+1 =f physical (x k ,u k , k, γ)

[0078]

[0079] Among them, x k is the state of the complex equipment system, x r is a state that the complex equipment is expected to reach, P and Q are positive definite matrices, K is the time required for control and decision-making, and u- is the threshold of decision-making control input.

[0080] This embodiment uses a model-based reinforcement learning problem to solve the above dynamic decision optimization problem, and uses a SAC-based deep reinforcement learning framework to train the intelligent agent;

[0081] The agent’s rewards are as follows:

[0082] r=-||x k -x r ||2

[0083] The optimization goal of the agent is:

[0084]

[0085] In the formula, represents the strategy function learned by the agent, s k is the state vector, a k is the action vector, E() is the expected function, α is the set weight, H is the entropy function of the strategy, a is the distribution that the agent obeys in the current strategy, ρ π is the action-state distribution under the current strategy.

[0086] The digital twin and decision-making agent designed in this embodiment mainly include two stages: offline training and online training. The offline stage refers to the process of model fitting training using historical data or a small amount of real data before the model is officially deployed online. The digital twin model and decision-making agent model after offline training will be deployed and used, connected to the equipment entity, and enter the online training or online evolution stage. The online evolution stage means that the digital twin uses real-time sensor data streams to update the model itself and keep it consistent with the physical equipment entity. In this embodiment, the dynamic evolution of the digital twin model constructed based on the data-driven Koopman operator theory is realized through the means of continuous learning in deep learning theory. Afterwards, the updated digital twin model and the model-based reinforcement learning theoretical method are used to continuously learn and train the decision-making agent based on deep reinforcement learning in the virtual space, so that it can adapt to the changes of complex equipment entities, and also achieve self-evolution, and thus make more optimized decisions.

[0087] A dynamic optimization system for a digital twin system, comprising:

[0088] A model construction module for constructing a digital twin system in the cloud, wherein the digital twin system includes a digital twin model and a decision-making agent model, the digital twin model is constructed based on a neural network with an autoencoder network structure, and the decision-making agent model is constructed based on a deep reinforcement learning framework;

[0089] A virtual-reality interaction module is used to deploy the digital twin system on the edge computing platform, perform virtual-reality interaction between the digital twin model and the equipment entity to obtain interaction data; perform data interaction between the decision-making intelligent agent model and the digital twin model to obtain real-time simulation data; and store the real-time simulation data and the interaction data in a database;

[0090] A credibility evaluation module is used to compare the interaction data of the digital twin model with the measured data of the equipment entity according to a set period to obtain credibility data of the digital twin model;

[0091] A model optimization module is used to compare the credibility data with a preset evolution threshold, perform dynamic evolution based on the comparison result, extract the real-time simulation data and the interaction data from the database according to a set ratio, optimize the parameters of the digital twin system based on the extracted data, obtain the digital twin system with updated parameters, and perform virtual-reality interaction with the equipment entity based on the digital twin system with updated parameters.

[0092] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the dynamic optimization method of a digital twin system.

[0093] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the dynamic optimization method of a digital twin system.

[0094] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A dynamic optimization method for a digital twin system, characterized in that: include: Constructing a digital twin system in the cloud, wherein the digital twin system includes a digital twin model and a decision-making agent model, wherein the digital twin model is constructed based on a neural network with an autoencoder network structure, and the decision-making agent model is constructed based on a deep reinforcement learning framework; The constraints of the digital twin model include: The first type of loss function L1: The second type of loss function L2: The third type of loss function L3: Where x k is the state vector, z k is a high-dimensional hidden state; is the encoder network, is the decoder network, K x K is the high-dimensional space Koopman linear operator for the system’s own state evolution. u is the high-dimensional space Koopman linear operator for external control input, u k For external input; Deploy the digital twin system on the edge computing platform, perform virtual-reality interaction between the digital twin model and the equipment entity to obtain interaction data; perform data interaction between the decision-making intelligent agent model and the digital twin model to obtain real-time simulation data; and store the real-time simulation data and the interaction data in a database; Comparing the interaction data of the digital twin model with the measured data of the equipment entity according to a set period to obtain the credibility data of the digital twin model; The credibility data is compared with a preset evolution threshold, dynamic evolution is performed based on the comparison result, the real-time simulation data and the interaction data are extracted from the database according to a set ratio, parameters of the digital twin system are optimized based on the extracted data, and a digital twin system with updated parameters is obtained. The digital twin system with updated parameters is used to perform virtual-reality interaction with the equipment entity.

2. The dynamic optimization method of a digital twin system according to claim 1, characterized in that: The reward function of the decision-making agent model includes: r=||x k -x r ||2 Where r is the reward function of the decision-making agent model, x k is the state vector, x r The state that complex equipment is expected to achieve.

3. The dynamic optimization method of a digital twin system according to claim 1, characterized in that: The optimization objective function of the decision-making agent model includes: In the formula, π represents the strategy function learned by the agent, s k is the state vector, a k is the action vector, E() is the expected function, α is the set weight, H is the entropy function of the strategy, a is the distribution that the agent obeys in the current strategy, ρ π is the action-state distribution under the current strategy.

4. The dynamic optimization method of a digital twin system according to claim 1, characterized in that: Optimizing the parameters of the digital twin system based on the extracted data to obtain a digital twin system with updated parameters, specifically including: Dynamically evolving the digital twin model in the cloud based on the extracted data to obtain an updated digital twin model; In the cloud, the decision-making agent model is trained and updated based on the updated digital twin model combined with the MBRL algorithm and continuous learning theory to obtain an updated decision-making agent model; A parameter-updated digital twin system is constructed based on the model parameters of the updated digital twin model and the updated decision-making intelligent agent model, and the parameter-updated digital twin system is deployed in the edge computing platform.

5. A dynamic optimization system for a digital twin system, applied to a dynamic optimization method for a digital twin system according to any one of claims 1 to 4, characterized in that: include: A model construction module for constructing a digital twin system in the cloud, wherein the digital twin system includes a digital twin model and a decision-making agent model, the digital twin model is constructed based on a neural network with an autoencoder network structure, and the decision-making agent model is constructed based on a deep reinforcement learning framework; A virtual-reality interaction module is used to deploy the digital twin system on the edge computing platform, perform virtual-reality interaction between the digital twin model and the equipment entity to obtain interaction data; perform data interaction between the decision-making intelligent agent model and the digital twin model to obtain real-time simulation data; and store the real-time simulation data and the interaction data in a database; A credibility evaluation module is used to compare the interaction data of the digital twin model with the measured data of the equipment entity according to a set period to obtain credibility data of the digital twin model; A model optimization module is used to compare the credibility data with a preset evolution threshold, perform dynamic evolution based on the comparison result, extract the real-time simulation data and the interaction data from the database according to a set ratio, optimize the parameters of the digital twin system based on the extracted data, obtain the digital twin system with updated parameters, and perform virtual-reality interaction with the equipment entity based on the digital twin system with updated parameters.

6. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a dynamic optimization method of a digital twin system according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements a dynamic optimization method for a digital twin system as described in any one of claims 1 to 4.