Cluster intelligent decision model training method based on digital twinning
By adopting digital twin technology and parallel training methods in the multi-agent deep reinforcement learning cluster intelligent decision model, the problem of low training speed and effectiveness of the cluster intelligent decision model is solved, and efficient and fast model training and decision-making process is achieved.
Patent Information
- Application Number
- CN202510146984.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The training speed and effectiveness of the intelligent decision-making model of the multi-agent deep reinforcement learning cluster is low, mainly due to the low data acquisition efficiency and insufficient fidelity of learning samples.
Using a cluster intelligent decision model training method based on digital twins, a system architecture consisting of physical entities, twin models, twin models, smart decision model and twin simulation middleware is built, and a high-fidelity state information is obtained using digital twin technology, and the training process is accelerated by running multiple twin models in parallel.
The training speed and effectiveness of the cluster intelligent decision-making model are significantly improved, and high-fidelity data is obtained through digital twin technology, which reduces the time for learning samples acquisition, and improves the training efficiency of the model and the accuracy of decision-making.
Smart Images

Figure CN120068987A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of training of swarm intelligence decision-making models, and particularly relates to a method for training an intelligent decision-making model based on multi-agent deep reinforcement learning with digital twin. Background Art
[0002] In recent years, with the development of technologies such as big data, machine learning, and deep learning, significant progress has been made in the research on decision-making models based on deep reinforcement learning. These technologies provide powerful computing and analysis capabilities for decision-making models based on deep reinforcement learning (DRL), significantly improving the intelligent decision-making capabilities in various application fields. The great success of DRL has prompted researchers to turn their attention to the multi-agent field, which has given rise to the development of multi-agent deep reinforcement learning (MADRL). After several years of development and innovation, numerous algorithms, rules, and frameworks have emerged in MADRL, and the MADRL swarm intelligence decision-making model has shown broad application prospects and huge value space.
[0003] MADRL is a type of interactive learning method for policy design. Its goal is to learn and accumulate experience through continuous interaction and feedback between multiple agents and a dynamic environment, construct an iteratively optimized behavior policy for multiple agents, and the output of its policy does not depend on a specific environment. Compared with decision-making algorithms for optimization solving and policy design, the multi-agent deep reinforcement learning algorithm for policy optimization has more prominent advantages in solving swarm intelligence decision-making problems in a dynamically uncertain environment.
[0004] MADRL is trained by the method of multi-agent exploring and obtaining data simultaneously. Compared with traditional supervised learning, it does not require processes such as manual data annotation and data cleaning, reducing the difficulty of data acquisition. However, in essence, it is still a data-driven method. The higher the efficiency of multi-agent interacting with the environment to obtain learning samples and the stronger the computing power, the faster the training speed of the MADRL swarm intelligence decision-making model. The training of the MADRL swarm intelligence decision-making model depends on a large amount of learning sample data. Using the method of actual testing to learn and explore the MADRL swarm intelligence decision-making strategy not only has high experimental costs and large risks but also has extremely low learning sample acquisition efficiency, resulting in an extremely slow training speed of the swarm intelligence decision-making model. Although using computer simulation to replace actual testing can solve the above problems, due to the strong nonlinear characteristics and uncertainties of real agent parameters and environmental data, there are large errors between the learning environment constructed only by computer simulation and the real environment, resulting in low fidelity of learning samples and ultimately affecting the effectiveness of the MADRL swarm intelligence decision-making model.
[0005] Digital Twin (DT) is a key technology for realizing the mapping from physical domain entities to digital models in the information domain. It collects data on physical entities through sensors deployed in various parts of the system and transmits the data to the information domain through the virtual-real interaction interface. The information domain integrates data of multi-dimensional physical entities and, through modeling and analysis, finally forms a simulation of multiple disciplines, multiple physical quantities, multiple time scales, and multiple probabilities, so as to nearly real-time present the full life cycle process of physical entities in the real scenario and control the physical entities through the virtual-real interaction interface. With the help of Digital Twin, MADRL can easily obtain high-fidelity state information of the real world for model training. Therefore, a Digital Twin training environment can be established and the MADRL algorithm can be used in data processing to provide intelligent cognitive decisions for collaborative control applications in the real world. By perceiving the changes in the real environment, the intelligent decision-making model running in the Digital Twin can be updated in real time to adapt to the changes in the dynamic environment. In other words, Digital Twin can solve the problem of low data acquisition efficiency when MADRL is applied to collaborative control applications. However, how to combine Digital Twin in the framework of MADRL to solve the training problem of the swarm intelligence decision-making model has not been fully studied. Summary of the Invention
[0006] The object of the present invention is to propose a Digital Twin training method for a swarm intelligence decision-making model aiming at the training problem of the multi-agent deep reinforcement learning decision-making model, aiming to improve the speed and effectiveness of the training of the swarm intelligence decision-making model. To achieve this object, the steps adopted by the present invention are as follows:
[0007] Step 1: Build a system architecture of the swarm intelligence decision-making model training method based on Digital Twin. This system architecture consists of five parts: physical entities, twin models, twin model replicas, intelligent decision-making models, and twin simulation middleware. Among them, the physical entities are the intelligent agents in the real world and their surrounding environment. The twin model is a high-precision mapping of the physical entities in the digital domain. The twin model replica is a replica with the same twin model and simulation parameter configuration obtained by copying and correcting the twin model at the code level. The swarm intelligence decision-making model is implemented by the MADRL algorithm and is used as the calculation and control center for completing the training of the swarm intelligence collaborative control decision-making model. The twin simulation middleware is the bridge for data interaction between the physical entities and the twin models, and between the twin model replicas and the decision-making models.
[0008] Step 2: Before the training starts, the physical entities collect twin data of the environmental information and their own state information through various high-precision sensors and transmit them to their twin simulation models through the twin simulation middleware to correct the model parameters of the twin simulation models and improve the modeling accuracy of the collaborative control task scenarios.
[0009] Step 3: At the start of training, the central server runs N corrected twin model copies simultaneously in parallel. Each agent in each twin model copy takes its own observed state information as the state input of its policy network. The policy network of multiple agents performs decision-making calculations through a deep neural network and directly outputs the collaborative behavior at the next moment. At each moment, each twin model copy calculates the reward obtained by each agent at the current moment, and combines the joint observations, joint actions of each agent at the previous moment, and the joint observations of each agent at the current moment to generate a sample data and stores it in the experience replay unit.
[0010] Step 4: After repeating Step 3 several times, randomly extract some sample data from the experience replay unit at regular intervals to calculate the loss functions of the evaluation network and the policy network, and use the gradient descent method to update the parameters of the evaluation network and the policy network to realize the training of the swarm intelligence decision-making model until the parameters of the evaluation network and the policy network converge.
[0011] Step 5: After training is completed, the central server deploys the trained swarm intelligence decision-making model to the physical entity through the twin simulation middleware. The physical entity distributes and executes specific tasks according to the decision results of the swarm intelligence policy network. At the same time, the swarm intelligence decision-making model can continue to perform continuous training of the swarm intelligence decision-making model according to the state data in the physical entity, and regularly update the better trained model to the physical entity through the twin simulation middleware to realize the continuous evolution of the swarm intelligence decision-making model.
[0012] With the help of the swarm intelligence decision-making model training system based on digital twin, the swarm intelligence decision-making model realizes the rapid training and deployment of the decision-making model in the way of "twin training, distributed execution, and continuous evolution". Brief Description of the Drawings
[0013] Figure 1 is the overall scheme of the swarm intelligence decision-making model training based on digital twin proposed by the present invention
[0014] Figure 2 is the framework of the swarm intelligence decision-making model training system based on digital twin proposed by the present invention
[0015] Figure 3 is the comparison of the swarm intelligence decision-making model training method based on digital twin proposed by the present invention with the non-twin training results Detailed Embodiment
[0016] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0017] Taking the swarm intelligence decision-making model of unmanned aerial vehicles as an example, Figure 1This is the overall solution for training the decision-making model of cluster intelligent collaborative deep reinforcement learning based on digital twin proposed by the present invention, which consists of three parts: "centralized training, distributed decision-making, and autonomous evolution". During the training of the decision-making model, the twin model uses the high-fidelity characteristics of digital twin and the powerful computing resources of the central server to conduct extensive behavior exploration, improve the efficiency of obtaining reinforcement learning samples, and evaluate and test the decision-making model through physical entities to drive the optimization of the decision-making model and ensure the effectiveness of the decision-making model. During the execution process, each agent can perform distributed intelligent collaborative control decisions based on its respective partial observation results, and at the same time send the observation data obtained in the real task environment to the central server to continuously drive the autonomous evolution of the decision-making model.
[0018] Taking the classic actor-critic deep reinforcement learning network structure as an example, Figure 2 The specific method for training the intelligent decision-making model of the UAV cluster in the twin model is also given.
[0019] Step 1: Construct a training system for the intelligent collaborative decision-making model of the UAV cluster based on digital twin. This training system consists of the following five parts:
[0020] (1) Physical entity: The whole composed of the UAV cluster in the real world and the task environment is called the physical entity. Among them, each UAV is restricted by resources in terms of computing and storage and cannot complete the training of the deep reinforcement learning decision-making model efficiently. Each agent is equipped with multiple sensors, which can sense the environmental state and its own state information in real time and execute the behavior control commands from the central server.
[0021] (2) Twin simulation model: The central server collects the observation data of the physical entity from the real world through sensors, and through simulation and modeling, establishes a high-fidelity twin simulation model of the physical entity. At the same time, the central server uses the perception data from the real-world sensors to update the twin simulation model in real time at each time step. The twin simulation model can obtain global state information, which is used to improve the training effectiveness of the deep reinforcement learning decision-making model.
[0022] (3) Twin model replicas: At the code level, the central server generates multiple high-fidelity twin model replicas by copying the twin simulation model description file and simulation parameters, so as to ensure that the training scenarios and simulation parameter configurations of each high-fidelity twin model replica are the same. During the training process, multiple twin model replicas can provide more sample data for the training of the decision-making model within one time step, taking into account improving the training speed and effectiveness of the deep reinforcement learning decision-making model.
[0023] (4) Swarm Intelligence Decision Model: The MADRL algorithm is used to implement the swarm intelligence decision model, which provides decision-making services for swarm collaborative control applications. The swarm intelligence decision model takes the observed state information of the agents as input and utilizes the powerful computing performance of the central server to calculate the loss function of the policy network and optimize the decision model parameters. The physical entities remain connected to the twin simulation model during the execution phase. The decision model can be continuously trained during the execution phase and updated to the real-world agents regularly through the twin simulation middleware, thus realizing the continuous evolution of the decision model.
[0024] (5) Twin Simulation Middleware: During the training phase, the twin simulation middleware establishes a two-way connection channel for the swarm intelligence decision model and multiple twin model replicas. The agent behavior control strategies output by the swarm intelligence decision model can be accurately transmitted to the corresponding twin model replicas through the twin simulation middleware. The observed state data of the agents in the multiple twin model replicas can be transmitted to the experience replay pool of the swarm intelligence decision model through the twin simulation middleware. The twin simulation middleware is also the bridge connecting the physical entities and the twin simulation model, establishing a communication connection channel for the two and providing powerful data interaction and processing capabilities. This communication connection channel is two-way. On the one hand, before the training starts, the observed data of the physical entities can be transmitted to the twin simulation model through the twin simulation middleware to correct the state of the twin simulation model and improve the model accuracy. On the other hand, during the execution phase, the continuously updated model parameters of the twin decision model can also be deployed to the physical entities in the real world through the twin simulation middleware to update their decision model parameters.
[0025] Step 2: Dynamically correct the information domain twin model using the physical domain observation data to reduce the error between the twin model and the physical entity and further improve the accuracy of the twin model.
[0026] Before the training starts, each agent collects the twin data of the environmental information and its own state information through various high-precision sensors, transmits it to its twin simulation model through the twin simulation middleware, and uses the cross-domain data fusion technology to correct the model parameters of the information domain twin simulation model with the physical domain observation data to improve the modeling accuracy of the collaborative control task scenario.
[0027] Step 3: Run multiple high-fidelity twin model derivatives in a parallel working mode to provide more sample data for the efficient training of the decision model and improve the training efficiency.
[0028] At the start of training, the central server runs N corrected twin model copies simultaneously in parallel. Each agent in each twin model copy uses its own observed state information as the state input to its policy network. The policy network of the agent performs decision-making calculations through a deep neural network and directly outputs the collaborative behavior for the next moment. At each moment, each twin model copy calculates the reward obtained by each agent at the current moment, and combines the joint observations, joint actions of each agent at the previous moment, and the joint observations of each agent at the current moment to generate a sample data and stores it in the experience replay unit.
[0029] Step 4: Periodically randomly extract some sample data from the experience replay unit for model training.
[0030] After repeating Step 3 several times, randomly extract some sample data from the experience replay unit at regular intervals to calculate the loss functions of the evaluation network and the policy network, and use the gradient descent method to update the parameters of the evaluation network and the policy network to achieve the training of the deep reinforcement learning decision model until the parameters of the evaluation network and the policy network converge.
[0031] Step 5: During the execution process, the real-world agent sends the observation data in the real task environment to the central server through the twin simulation middleware, and can continue to train the decision model, thus realizing the continuous evolution of the decision model.
[0032] After the training is completed, the central server deploys the trained deep reinforcement learning decision model to the physical entity through the twin simulation middleware. The physical entity distributes and executes specific tasks according to the decision results of the deep reinforcement learning network model. At the same time, the twin decision model continues to perform continuous training on the deep reinforcement learning network model according to the state data in the physical entity, and regularly updates the better training results to the physical entity through the twin simulation middleware to realize the continuous evolution of the deep reinforcement learning method.
[0033] Appendix Figure 3 Presents the preliminary results of the training method for the swarm intelligence decision model based on digital twin proposed in the present invention. When using twin training, the convergence speed of the average episode reward is faster than that of non-twin training. The use of multiple parallel-running twin model copies increases the sampling data and the number of training times per unit time, thus improving the training speed. Twin training and non-twin training both finally converge to almost the same value, further verifying the effectiveness of the proposed training method for the swarm intelligence decision model based on digital twin.
[0034] The content not described in detail in this application of the present invention belongs to the prior art well-known to those skilled in the art.
Claims
1. A cluster intelligent decision-making model training method based on digital twins, characterized in that: The steps taken are: Step 1: Build a system architecture for the cluster intelligent decision-making model training method based on digital twins. The system architecture consists of five parts: physical entity, twin model, twin model replica, intelligent decision model and twin simulation middleware; (1) Physical entity: The real-world drone swarm and mission environment are called physical entities. Each drone is limited by computing and storage resources and cannot efficiently complete the training of deep reinforcement learning decision models. Each intelligent agent is equipped with multiple sensors, which can perceive the environment and its own status information in real time and execute behavior control commands from the central server. (2) Twin simulation model: The central server collects observation data of physical entities from the real world through sensors, and builds a high-fidelity twin simulation model of the physical entity through simulation and modeling. At the same time, the central server uses the perception data from real-world sensors to update the twin simulation model in real time at each time step. The twin simulation model can obtain global state information, which is used to improve the training effectiveness of the deep reinforcement learning decision model. (3) Twin model replicas: At the code level, the central server generates multiple high-fidelity twin model replicas by copying the twin simulation model description file and simulation parameters, thereby ensuring that the training scenario and simulation parameter configuration of each high-fidelity twin model replica are the same; during the training process, multiple twin model replicas can provide more sample data for the training of the decision model within one time step, thereby improving the training speed and effectiveness of the deep reinforcement learning decision model; (4) Cluster intelligent decision model: The MADRL algorithm is used to implement the cluster intelligent decision model to provide decision services for cluster collaborative control applications. The cluster intelligent decision model takes the observed state information of the intelligent agent as input, and uses the powerful computing performance of the central server to calculate the loss function of the policy network and optimize the decision model parameters. The physical entity remains connected to the twin simulation model during the execution phase. The decision model can be continuously trained during the execution phase and regularly updated to the intelligent agent in the real world using the twin simulation middleware, thereby achieving continuous evolution of the decision model. (5) Twin simulation middleware: During the training phase, the twin simulation middleware establishes a two-way connection channel between the cluster intelligent decision-making model and multiple twin model copies. The intelligent agent behavior control strategy output by the cluster intelligent decision-making model can be accurately transmitted to the corresponding twin model copies through the twin simulation middleware, and the observation state data of the intelligent agents in multiple twin model copies can be transmitted to the experience replay pool of the cluster intelligent decision-making model through the twin simulation middleware. The twin simulation middleware is also a bridge connecting the physical entity and the twin simulation model, establishing a communication connection channel between the two and providing powerful data interaction and processing capabilities. This communication connection channel is bidirectional. On the one hand, before the training starts, the observation data of the physical entity can be transmitted to the twin simulation model through the twin simulation middleware to correct the state of the twin simulation model and thus improve the accuracy of the model. On the other hand, during the execution phase, the model parameters continuously updated by the twin decision model can also be deployed to the physical entity in the real world through the twin simulation middleware to update its decision model parameters. Step 2: Before training begins, the physical entity collects twin data of environmental information and its own state information through various high-precision sensors, transmits it to its twin simulation model through the twin simulation middleware, corrects the model parameters of the twin simulation model, and improves the modeling accuracy of the collaborative control task scenario; Step 3: At the beginning of training, the central server runs N corrected twin model copies simultaneously in parallel; each agent in each twin model copy uses its own observation state information as the state input of its policy network; the multi-agent policy network performs decision calculations through a deep neural network and directly outputs the collaborative behavior of the next moment; at each moment, each twin model copy calculates the reward obtained by each agent at the current moment, and combines the joint observation and joint action of each agent at the previous moment with the joint observation of each agent at the current moment to generate a sample data and store it in the experience playback unit; Step 4: After repeating step 3 several times, randomly extract some sample data from the experience playback unit at regular intervals to calculate the loss function of the evaluation network and the policy network, and use the gradient descent method to update the parameters of the evaluation network and the policy network to train the swarm intelligent decision-making model until the parameters of the evaluation network and the policy network converge; Step 5: After the training is completed, the central server deploys the trained cluster intelligent decision-making model to the physical entity through the twin simulation middleware. The physical entity distributes and executes specific tasks according to the decision results of the cluster intelligent strategy network. At the same time, the cluster intelligent decision-making model can continue to train the cluster intelligent decision-making model according to the status data in the physical entity, and regularly update the better training model to the physical entity through the twin simulation middleware to realize the continuous evolution of the cluster intelligent decision-making model.
Citation Information
Cited By
Drying scheme decision-making method and system for improving quality of agricultural products
CN121541614A