Adversarial environment reinforcement learning model training method and system based on digital twinning

By building a digital twin model and generating a virtual adversarial environment, combining the reinforcement learning decision model architecture and optimization algorithm, the problem of insufficient generalization ability of the existing reinforcement learning model in the adversarial environment is solved, and higher decision accuracy and adaptability are achieved.

CN120068991APending Publication Date: 2025-05-30NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510194073.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing reinforcement learning models lack generalization capabilities and lack of training resources in the adversarial environment, making it difficult to effectively capture the dynamic game process, resulting in the failure of decision-making when facing a new adversarial situation.

Method used

By building a digital twin model, a variety of virtual adversarial environment and target behavior variant data are generated, reinforcement learning decision model architecture is designed, strategy gradient algorithm and PPO algorithm are used for iterative training and parameter optimization, and environment dynamic simulation and data generation are combined with Bayesian network and Monte Carlo tree search algorithm.

Benefits of technology

It significantly increases the diversity and scale of the training data, avoids model overfitting, improves decision accuracy and adaptability, and improves the performance and efficiency of the model in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068991A_ABST
    Figure CN120068991A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to an adversarial environment reinforcement learning model training method and system based on digital twinning. Comprising the following steps: constructing a digital twinborn model based on existing confrontation environment data, and generating various virtual confrontation environments and target behavior variant data; designing a reinforcement learning decision model architecture, defining a multi-dimensional state space, constructing an action-reward function and initializing network parameters; pre-training the reinforcement learning model by using the existing data to form a basic model; inputting a plurality of virtual environments generated by digital twinning and target data into a decision model for iterative training, and optimizing network parameters by adopting a strategy gradient algorithm; and testing the iterated model and the basic model on the basis of the original confrontation environment data, if the decision performance of the new model is improved, continuing iteration until reaching the standard, otherwise, reversely correcting the parameters of the digital twinborn model, and continuing iteration. The problems of reinforcement learning model data shortage and overfitting in a confrontation environment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a method and system for training an adversarial environment reinforcement learning model based on digital twin. Background Art

[0002] In the prior art, the training methods of reinforcement learning models generally suffer from the dual bottlenecks of insufficient generalization ability and lack of training resources. Traditional methods mainly rely on original adversarial environment data for model training, and their training effects are limited by the scale and diversity of real-scene data. Specifically, in the process of dynamic adversarial environment modeling, the acquisition of real data requires the deployment of a large number of physical sensors and adversarial equipment, which is costly and difficult to cover all potential adversarial scenarios, resulting in the decision-making failure of the trained model when facing new adversarial situations. At the same time, existing training methods mostly use fixed data sets for batch training, and cannot effectively capture the non-linear time-varying dynamic game process in the adversarial environment. Their parameter update strategies are limited by the static snapshot data at discrete time points, lacking the ability to accurately model the non-linear dynamic evolution process in complex adversarial environments, resulting in the model being difficult to adapt to the real-time changing battlefield situation.

[0003] Although advanced reinforcement learning training techniques can alleviate problems such as data scarcity and overfitting through various virtual environments generated by digital twin, there are still critical defects in practical applications. In the virtual environment construction stage of existing digital twin-driven methods, a single-mode parameter perturbation strategy is often adopted, resulting in the generated adversarial environment variant data showing homogeneous characteristics, and it is difficult to effectively simulate the multi-dimensional dynamic coupling effect in real adversarial scenarios. At the same time, the reinforcement learning model is prone to overfitting to the virtual environment during the iterative training process, manifested as a falsely high decision accuracy in the digital twin environment, but a significant decline in performance when migrated to the real adversarial environment. For complex adversarial scenarios with dynamic evolution, virtual training data that is consistent with the real environment in terms of features cannot be provided, resulting in decision-making feature mismatch problems during the virtual-real migration process, limiting the applicability of digital twin-driven methods in real environments. Summary of the Invention

[0004] In order to solve the problems of data scarcity and overfitting of the reinforcement learning decision model in the adversarial environment, the present invention provides a method and system for training an adversarial environment reinforcement learning model based on digital twin.

[0005] In a first aspect, a method for constructing a knowledge graph related to the situation and action of maritime equipment provided by the present invention adopts the following technical solution: A method for training an adversarial environment reinforcement learning model based on digital twin includes: Constructing a digital twin model based on existing adversarial environment data, and generating a variety of virtual adversarial environments and target behavior variant data; Design a reinforcement learning decision model architecture, define a multi-dimensional state space, construct an action-reward function, and initialize network parameters; Pre-train the reinforcement learning model using existing data to form a basic model; Input multiple virtual environments generated by digital twins and target data into the decision model for iterative training, and use the policy gradient algorithm to optimize network parameters; Test the iterated model and the basic model on the original adversarial environment data. If the decision-making performance of the new model improves, continue iterating until it meets the standard. Otherwise, reverse-correct the parameters of the digital twin model and then continue iterating.

[0006] Furthermore, constructing a digital twin model based on existing adversarial environment data includes: Extract the situation element features in the adversarial environment data, including three-dimensional data of spatial layout, equipment status, and confrontation rules; Construct a digital twin topology structure that includes the dynamic correlation relationship of environmental parameters. The topology structure includes environmental entity nodes, behavior rule edges, and situation attribute tensors; Use a Bayesian network to model the probability dependence relationship between environmental parameters, establish an environmental state transition equation, and effectively simulate the dynamic changes of the environment.

[0007] Furthermore, generating multiple virtual adversarial environments and target behavior variant data includes; Use the Monte Carlo tree search algorithm to generate combinations of environmental parameter perturbations and produce multiple virtual adversarial environment variant data; Use a conditional generative adversarial network to decouple and recombine the features of target behavior data to generate multiple target behavior variant data; Construct continuous transition samples between the original data and the generated data through parameter interpolation to form a progressive training dataset.

[0008] Furthermore, the design of the reinforcement learning decision model architecture includes: Construct a hybrid action space that includes discrete action selection and continuous parameter adjustment; Define a multi-dimensional state space that includes environmental information, target state, and self-state, and quantize and encode it to form an input vector that the model can process; Construct an action-reward function, design a triple reward function based on task completion rate, resource consumption rate, and confrontation advantage degree to motivate the model to learn the optimal decision-making strategy; Set up a three-branch network structure that includes a policy network, a value network, and an environmental modeling network and initialize network parameters to provide initial conditions for model training.

[0009] Further, the pre-training of the reinforcement learning model using existing data to form a basic model includes: Divide the existing adversarial environment data into a training set and a validation set according to a ratio of 7:3 to ensure the rationality and effectiveness of data division; Adopt a double-buffer experience pool mechanism for batch data sampling to improve data utilization efficiency and training stability; Set a dynamically decaying ε-greedy exploration strategy with an initial exploration rate ε = 0.6 to balance the exploration and exploitation capabilities of the model; Terminate the training in a timely manner when the accuracy on the validation set drops through an early stopping mechanism, save the optimal parameters, and prevent overfitting of the model.

[0010] Further, the iterative training of the multiple virtual environments generated by digital twin and the target data input into the decision model, and the optimization of network parameters using the policy gradient algorithm include: Clearly divide the virtual environment data into three training stages: basic, medium, and difficult according to complexity to achieve curriculum learning; Adopt the PPO algorithm for policy update, calculate the gradient of the model at each decision step, and evaluate the influence degree of the model's decision-making strategy on the reward through gradient information; Update the network parameters according to the calculated gradient, adjust the decision-making strategy of the model in the direction of obtaining higher rewards, and gradually optimize the performance of the model; Increase the proportion of high-difficulty samples by 10% after each round of iteration to gradually improve the adaptability of the model.

[0011] Further, the testing of the iterated model and the basic model on the original adversarial environment data includes: Construct a test benchmark set containing 80% of the scenes of the original environment data to ensure the comprehensiveness and representativeness of the test; Run the basic model and the iterated model in parallel and record the differences in the decision-making trajectories of the two models; Calculate the dual evaluation indicators of the change in policy entropy value and the deviation degree of Q-value distribution to comprehensively evaluate the model performance.

[0012] Further, if the decision-making performance of the new model is improved, continue to iterate until it reaches the standard; otherwise, reverse-correct the parameters of the digital twin model and then continue to iterate, including: When the accuracy is improved but not up to the standard, continue to generate multiple virtual adversarial environments through the digital twin model and iteratively train the reinforcement learning model; When the accuracy rate drops, quickly activate the parameter correction loop of the digital twin model: reduce the generator learning rate to slow down the learning speed of the generator; enhance the spectral normalization constraint of the discriminator to improve the stability of the discriminator; inject 10% of the original data features into the generation process to ensure the authenticity of the generated data.

[0013] In a second aspect, a training system for an adversarial environment reinforcement learning model based on digital twins includes: A digital twin model construction module for constructing a digital twin model based on existing adversarial environment data and providing a basic model architecture for subsequent modules; A virtual adversarial environment generation module connected to the digital twin model construction module, generating multiple virtual adversarial environments based on the digital twin model to simulate real adversarial scenarios; A reinforcement learning model construction module for constructing the framework of the reinforcement learning model and providing a basic model architecture for subsequent training; A model pre-training module connected to the reinforcement learning model construction module, pre-training the reinforcement learning model using existing adversarial environment data to form a basic reinforcement learning model; A reinforcement learning model training iteration module connected to the model pre-training module and the virtual adversarial environment generation module, training and iterating the basic model through multiple virtual adversarial environments to continuously improve the model performance; also connected to the test and evaluation module, and will continue iterative training when the model performance does not meet the standard; also connected to the parameter correction module, and will continue the iterative training of the decision-making model after passing through the parameter correction module when the model performance drops; A test and evaluation module connected to the reinforcement learning model training iteration module, testing and evaluating the model after training and iteration to judge the quality of the model; A parameter correction module connected to the test and evaluation module, reversely correcting the parameters of the digital twin model when the performance of the reinforcement learning model drops; A final model output module connected to the test and evaluation module, outputting the qualified model after multiple rounds of iterative training, evaluation, and parameter correction to complete the entire training process.

[0014] In a third aspect, the present invention provides a computer-readable storage medium storing multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the method for constructing a knowledge graph related to the situation and action of maritime equipment.

[0015] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the method for constructing a knowledge graph related to the situation and action of maritime equipment.

[0016] In summary, the present invention has the following beneficial technical effects: 1. A method and system for training an adversarial environment reinforcement learning model based on digital twin proposed by the present invention solves the problem of lack of data for the reinforcement learning decision-making model in the adversarial environment by constructing a digital twin model and generating various virtual adversarial environments and target behavior variant data. Compared with the traditional training method that relies on limited original data, the diversity and scale of training data are greatly increased. At the same time, through the iterative training and parameter correction mechanism, overfitting of the model is avoided, and the decision-making accuracy of the model in practical applications is improved, and both the training efficiency and effect are significantly improved.

[0017] 2. The method for constructing a digital twin model and generating virtual data adopted by the present invention can effectively simulate the dynamic changes of the environment. When constructing the digital twin model, the situation element features are extracted, the topological structure and state transition equation are established, and when generating virtual data, various algorithms are used to generate variant data and progressive training sets. This not only solves the problem that traditional methods are difficult to accurately simulate complex adversarial environments, but also provides a more realistic and rich training scenario for the reinforcement learning model, enhancing the generalization ability and adaptability of the model.

[0018] 3. The reinforcement learning decision model architecture and pre-training method designed by the present invention have good scientificity and effectiveness. A hybrid action space, multi-dimensional state space, triple reward function and three-branch network structure are constructed, and at the same time, reasonable data partitioning, experience pool mechanism, exploration strategy and early stopping mechanism are adopted. These measures comprehensively consider various factors in the adversarial environment, improve the data utilization efficiency and training stability, balance the exploration and utilization capabilities of the model, prevent overfitting, and lay a solid foundation for the subsequent training of the model.

[0019] 4. The strategies adopted by the present invention in the model training, testing and adjustment processes ensure the continuous optimization and stability of the model performance. Curriculum learning and policy update algorithms are adopted during iterative training, and comprehensive evaluation indicators are used during testing, and parameters are iterated or corrected according to the change of accuracy. This enables the model to gradually adapt to adversarial scenarios of different difficulties, accurately evaluate its own performance, and adjust in time to adapt to environmental changes, providing reliable support for decision-making in complex adversarial scenarios and enhancing the practicality and reliability of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is the flowchart of the method in Embodiment 1 of the present invention; Figure 2 is the schematic structural diagram of the system in Embodiment 2 of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Example 1 Reference Figure 1 , a method for training an adversarial environment reinforcement learning model based on digital twin in this embodiment, includes: S1. Construct a digital twin model based on existing adversarial environment data, and generate various virtual adversarial environments and target behavior variant data, including: S11. Extract the situation element features in the adversarial environment data, including three-dimensional data of spatial layout, equipment status, and adversarial rules; In adversarial environment modeling, the comprehensive representation of situation elements is the basis for constructing a digital twin model. By analyzing historical adversarial data, geographical information such as battlefield terrain and obstacle distribution is extracted from the spatial layout dimension; dynamic attributes such as radar detection range and weapon system parameters are extracted from the equipment status dimension; logical constraints such as engagement rules and cooperation agreements are extracted from the adversarial rules dimension. The comprehensive extraction of three-dimensional data provides multi-granularity feature support for the digital twin model.

[0023] S12. Construct a digital twin topology structure containing dynamic correlation relationships of environmental parameters, including environmental entity nodes, behavior rule edges, and situation attribute tensors; Abstract the collected information, take each unit participating in the confrontation as an entity node; the action connections between them as behavior rule edges; use situation attribute tensors to characterize the characteristics of entities and relationships through multi-dimensional matrices. This topology structure constructs the logical framework of the adversarial environment, presenting complex environmental information in a structured form for subsequent analysis and simulation.

[0024] S13. Use Bayesian network to model the probability dependence relationship between environmental parameters and establish an environmental state transition equation; By analyzing a large amount of historical data and expert experience, determine the probability correlation between environmental parameters and form a conditional probability table. The Bayesian network quantifies the dependence relationship between environmental parameters through the conditional probability table. Based on these conditional probability tables, establish a state transition equation, and describe the dynamic evolution law between parameters through the state transition equation, so as to predict the change trend of the adversarial environment state under different conditions, enabling the digital twin model to accurately predict the continuous change of the environmental state.

[0025] S14. Use the Monte Carlo tree search algorithm to generate a combination of environmental parameter perturbations and generate various virtual adversarial environment variant data; Randomly combine and perturb various parameters in the environment through the Monte Carlo tree search algorithm, so as to effectively explore among numerous possible combinations, generate rich and diverse virtual adversarial environments, enable the model to be trained in different scenarios, and improve its adaptability and generalization ability. Thus, on the premise of ensuring environmental rationality, the diversity of training data is significantly expanded.

[0026] S15. Use a conditional generative adversarial network to decouple and recombine the features of the target behavior data to generate multiple variants of the target behavior data; The conditional generative adversarial network decouples the target behavior features into independent latent variables such as tactical intentions and maneuver patterns. By recombining the latent variables, target behaviors with new tactical features are generated, generating new composite tactical behavior data and breaking through the limitations of the behavior patterns of historical data. This enables the model to learn diverse target behavior patterns and enhance its ability to handle complex adversarial situations.

[0027] S16. Construct continuous transition samples between the original data and the generated data through parametric interpolation to form a progressive training dataset; Between the adversarial scenario data of different difficulty levels, a series of transition samples are generated using parametric interpolation. This allows the model to gradually learn from simple to complex behavior patterns, avoiding policy oscillations during training, and thus steadily enhancing the model's adaptability.

[0028] S2. Design a reinforcement learning decision model architecture, including: S21. Construct a hybrid action space that includes discrete action selection and continuous parameter adjustment; The design of the hybrid action space meets complex decision-making requirements. The discrete action layer defines tactical-level selections, and the continuous parameter layer controls specific execution parameters. Through hierarchical action design, both the clarity of high-level decisions and the refinement of low-level control are achieved. The construction of this hybrid action space enables the model to make more flexible and diverse decisions according to different adversarial situations.

[0029] S22. Define a multi-dimensional state space that includes environmental information, target state, and self-state, and quantize and encode it to form an input vector that the model can process; Convert heterogeneous data of various types of environmental information, target state, and self-state into normalized vectors of the same dimension through feature engineering, and use the self-attention mechanism for feature weight allocation to form a standardized input that the model can process. This enables the model to accurately understand the current adversarial state and provide an accurate information basis for subsequent decision-making.

[0030] S23. Construct an action-reward function and design a triple reward function based on task completion rate, resource consumption rate, and adversarial advantage degree; In a specific task scenario, through a dynamic weight adjustment mechanism, different types of reward weights are flexibly allocated according to the characteristics of different scenarios. Through this scenario-specific reward allocation method, different reward dimensions are emphasized at different training stages to guide the model to learn the optimal decision-making strategy to achieve the task goal and optimize resource utilization and adversarial effects.

[0031] S24. Set up a three-branch network structure including a policy network, a value network, and an environment modeling network, and initialize the network parameters; The three-branch network adopts an architecture of a shared feature extraction layer and an independent decision-making layer. The policy network is responsible for generating specific action decisions based on the input state information; the value network is used to evaluate the value of each action in the current state and judge the quality of the decisions; the environment modeling network models and analyzes the input environmental information and extracts key features. Xavier initialization is used in combination with batch normalization to ensure that the initial parameter distributions of each branch network are reasonable and avoid gradient explosion.

[0032] S3. Pre-train the reinforcement learning model using existing data to form a basic model, including: S31. Divide the existing adversarial environment data into a training set and a validation set according to a ratio of 7:3; Stratified sampling strategy is adopted for data division to ensure that each adversarial scenario appears in the training set and the validation set according to the ratio.

[0033] S32. Adopt a double-buffer experience pool mechanism for batch data sampling; During the training process, set up two experience pools. One experience pool is used to store newly generated experience data, and the other experience pool is used to store historical high-quality data to provide data samples for model training. When the data in one experience pool is used for training, the other experience pool can continue to collect new experience data. When sampling, mix the new and old data in a ratio of 3:7, which not only retains historical experience but also incorporates the latest adversarial patterns, improving the training stability. Thus, the utilization efficiency of data is improved, and at the same time, the stability of the training process is enhanced, enabling the model to learn more diverse experiences.

[0034] S33. Set a dynamically decaying ε-greedy exploration strategy with an initial exploration rate ε = 0.6; The exploration rate decays according to the cosine annealing law: ε_t = 0.6×cos(πt / 2T), and smoothly decreases to 0.1 within the training period T. In the initial stage of model training, randomly select actions with a probability of 0.6 to explore new strategies and decision-making methods; as the training progresses, gradually reduce the probability of randomly selecting actions, and make more use of the learned strategies to balance the exploration and exploitation capabilities of the model. This strategy encourages extensive exploration in the initial stage and focuses on experience utilization in the later stage, balancing the exploration-exploitation contradiction.

[0035] S34. Terminate the training in a timely manner when the accuracy on the validation set drops through an early stopping mechanism, and save the optimal parameters; During the model training process, continuously monitor the change in accuracy on the validation set. It is triggered when the accuracy drops by more than 5% for three consecutive epochs, indicating that the model may have started overfitting and overlearning the noise and details in the training data. At this time, stop the training in a timely manner, save the current optimal model parameters, and cooperate with L2 regularization to jointly suppress overfitting and ensure the generalization ability of the model in practical applications.

[0036] S4. Input multiple virtual environments generated by the digital twin and target data into the decision model for iterative training. The policy gradient algorithm is used to optimize network parameters, including: S41. Clearly divide the virtual environment data into three training stages: basic, medium, and difficult according to complexity; The complexity grading is based on the environmental perturbation intensity and the target behavior variability. During training, the curriculum learning strategy is adopted, and the next stage is automatically unlocked when the model's winning rate in the basic stage exceeds 75%. In this way, the model gradually adapts to adversarial environments of different difficulties and steadily improves its capabilities.

[0037] S42. Use the PPO algorithm for policy update and calculate the gradient of the model at each decision step; The PPO algorithm limits the policy update amplitude through the importance sampling ratio. When calculating the advantage function, GAE estimation is used to balance bias and variance. The gradient update step size is dynamically adjusted according to the KL divergence to maintain the stability of policy improvement. After each action is executed by the model, the gradient is calculated based on the obtained reward feedback, and the influence degree of the current decision-making policy on the reward is evaluated through the gradient information.

[0038] S43. Update the network parameters according to the calculated gradient to adjust the decision-making policy of the model in the direction of obtaining higher rewards; Based on the calculated gradient information, adjust the parameters of the policy network, value network, etc. By adjusting these parameters, change the decision-making method of the model when facing different states, so that the model can make better choices in subsequent decisions, gradually optimize the performance of the model in the adversarial environment, and obtain higher rewards.

[0039] S44. Increase the proportion of high-difficulty samples by 10% after each round of iteration; During the training process, as the iteration progresses, gradually increase the proportion of training samples in difficult scenarios after each iteration. Let the model continuously contact and learn more complex adversarial environments, gradually improve its ability to handle complex situations, and enhance the generalization and adaptability of the model.

[0040] S5. Test the iterated model and the basic model on the original adversarial environment data. If the decision-making performance of the new model improves, continue the iteration until it meets the standard. Otherwise, reverse correct the parameters of the digital twin model and then continue the iteration, including: S51. Construct a test benchmark set covering 80% of the scenarios of the original environmental data; This test benchmark should cover as comprehensively as possible all kinds of scenarios that may occur in actual applications, ensuring that the performance of the model in actual application scenarios can be comprehensively tested, and evaluating the reliability and practicality of the model.

[0041] S52. Run the basic model and the iterative model in parallel, and record the differences in the decision-making trajectories of the two models; Adopt a dual-model synchronous deduction architecture and execute in parallel under the same initial conditions. Record the key difference indicators: the difference in the first-fire response time, the deviation of resource utilization rate, and the tactical decision consistency coefficient, and quantify the improvement effect of the model. By comparing and analyzing these differences, find out the improvements of the iterative model relative to the basic model and the existing problems, providing a basis for further optimizing the model.

[0042] S53. Calculate the dual evaluation indicators of the change in policy entropy value and the deviation degree of Q-value distribution to comprehensively evaluate the model performance; By calculating the change in policy entropy value, the certainty of the model's decision-making can be understood. The larger the entropy value, the higher the uncertainty of the decision-making; by calculating the deviation degree of Q-value distribution, the accuracy of the model's evaluation of different state values can be evaluated, and it can be judged whether the model deviates too much from the basic model. Comprehensively evaluate the performance of the model from these two perspectives to fully understand the advantages and disadvantages of the model.

[0043] S54. When the accuracy rate increases but does not reach the standard, continue to generate multiple virtual adversarial environments through the digital twin model and iteratively train the reinforcement learning model; At this time, continue to use the digital twin model to generate more complex virtual adversarial environments, and then input the data of these newly generated virtual environments into the reinforcement learning model for iterative training to continuously optimize the performance of the model and make it gradually reach or exceed the set target accuracy rate.

[0044] S55. When the accuracy rate drops, quickly activate the parameter correction loop of the digital twin model: reduce the learning rate of the generator to slow down the learning speed of the generator; enhance the spectral normalization constraint of the discriminator to improve the stability of the discriminator; inject 10% of the original data features into the generation process to ensure the authenticity of the generated data; When it is found that the accuracy of the model decreases due to the decline in the quality of the generated data, immediately take the above parameter correction measures. Reducing the generator learning rate can prevent the generator from learning too fast and producing unstable results; enhancing the spectral normalization constraint of the discriminator can improve the discriminator's judgment accuracy and stability for the generated data; injecting 10% of the original data features into the generation process can make the generated data closer to the real data, thus stabilizing the training process of the generative adversarial network, improving the quality of the generated data, and ensuring that the performance of the model can be effectively improved.

[0045] This embodiment provides a training method for a reinforcement learning model in an adversarial environment based on digital twin. This method is proposed to address the core problems existing in existing reinforcement learning technologies in dynamic adversarial scenarios, such as distorted environment modeling, single training data, and insufficient policy generalization ability. To solve the technical bottleneck of low training efficiency and poor environmental adaptability of intelligent decision-making models in complex and changeable adversarial environments, this method realizes multi-dimensional analysis and dynamic evolution simulation of the adversarial environment by constructing a digital twin dynamic model with virtual-real mapping. On this basis, the Monte Carlo tree search algorithm is used to generate a combination of environmental parameter perturbations, and the conditional generative adversarial network is combined to decouple and recombine the features of the target behavior data. The progressive training data set containing three complexity levels of basic, medium, and difficult is constructed by the parameter interpolation method, breaking through the limitations of the traditional training data mode being single and the scene coverage being insufficient. In terms of the design of the reinforcement learning model architecture, this method constructs a hybrid action space that includes discrete action selection and continuous parameter adjustment, and innovatively designs a triple reward function based on task completion rate, resource consumption rate, and adversarial advantage degree. By setting the three-branch network structure of the policy network, value network, and environmental modeling network, and adopting the double-buffer experience pool mechanism and dynamic decay ε-greedy exploration strategy, stable convergence in the model pre-training stage is achieved. In the iterative optimization stage, the curriculum learning mechanism is introduced to input the virtual environment data by complexity level, and the PPO algorithm is combined for policy update, through importance sampling ratio control and adaptive gradient step adjustment. After each iteration, 10% of the high-difficulty samples are dynamically injected to force the model to continuously break through the ability boundary. This method improves the decision-making accuracy and adaptability of the reinforcement learning decision-making model in complex adversarial environments, provides strong support for intelligent decision-making in practical application scenarios such as maritime confrontation, and enhances the overall effectiveness of related systems in dynamic adversarial environments.

[0046] Embodiment 2 Refer to Figure 2 , this embodiment provides a training system for a reinforcement learning model in an adversarial environment based on digital twin, including: A digital twin model construction module, which is used to construct a digital twin model based on the existing adversarial environment data and provide a basic model architecture for subsequent modules; The virtual adversarial environment generation module is connected to the digital twin model construction module, generates various virtual adversarial environments based on the digital twin model, and simulates real adversarial scenarios; The reinforcement learning model construction module constructs the framework of the reinforcement learning model and provides the basic model architecture for subsequent training; The model pre-training module is connected to the reinforcement learning model construction module, and pre-trains the reinforcement learning model using existing adversarial environment data to form a basic reinforcement learning model; The reinforcement learning model training and iteration module is connected to the model pre-training module and the virtual adversarial environment generation module, trains and iterates the basic model in various virtual adversarial environments to continuously improve the model performance; it is also connected to the test and evaluation module and will continue iterative training when the model performance does not meet the standard; it is also connected to the parameter correction module and will continue the iterative training of the decision-making model after passing through the parameter correction module when the model performance deteriorates; The test and evaluation module is connected to the reinforcement learning model training and iteration module, tests and evaluates the model after training and iteration, and judges the quality of the model; The parameter correction module is connected to the test and evaluation module, and when the performance of the reinforcement learning model deteriorates, it reversely corrects the parameters of the digital twin model; The final model output module is connected to the test and evaluation module, outputs the qualified model after multiple rounds of iterative training, evaluation, and parameter correction, and completes the entire training process.

[0047] A computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor of a terminal device to perform the method for constructing a knowledge graph related to the situation and action of marine equipment as described above.

[0048] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to perform the method for constructing a knowledge graph related to the situation and action of marine equipment as described above.

[0049] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A training method for a reinforcement learning model in an adversarial environment based on digital twins, characterized in that: include: Build a digital twin model based on existing adversarial environment data and generate multiple virtual adversarial environments and target behavior variant data; Design reinforcement learning decision model architecture, define multi-dimensional state space, construct action-reward function and initialize network parameters; Use existing data to pre-train the reinforcement learning model to form a basic model; The various virtual environments and target data generated by the digital twin are input into the decision model for iterative training, and the policy gradient algorithm is used to optimize the network parameters; The iterated model and the basic model are tested on the original adversarial environment data. If the decision-making performance of the new model is improved, the iteration continues until it meets the standard. Otherwise, the digital twin model parameters are reversed and the iteration continues.

2. According to the digital twin-based adversarial environment reinforcement learning model training method of claim 1, it is characterized in that: The digital twin model is constructed based on the existing adversarial environment data, including: Extract situational element features from confrontation environment data, including three-dimensional data of spatial layout, equipment status, and confrontation rules; Constructing a digital twin topological structure including dynamic associations of environmental parameters, wherein the topological structure includes environmental entity nodes, behavioral rule edges, and situation attribute tensors; By using Bayesian networks to model the probabilistic dependency relationship between environmental parameters, the environmental state transition equation is established to achieve effective simulation of dynamic changes in the environment.

3. According to the digital twin-based adversarial environment reinforcement learning model training method of claim 1, it is characterized in that: The generation of multiple virtual confrontation environments and target behavior variant data includes: The Monte Carlo tree search algorithm is used to generate environmental parameter perturbation combinations and produce a variety of virtual confrontation environment variant data; Use conditional generative adversarial networks to decouple and reorganize the target behavior data to generate multiple target behavior variant data; Continuous transition samples are constructed between the original data and the generated data through parameter interpolation to form a progressive training data set.

4. According to the digital twin-based adversarial environment reinforcement learning model training method of claim 1, it is characterized in that: The design of the reinforcement learning decision model architecture includes: Construct a hybrid action space that includes discrete action selection and continuous parameter adjustment; Define a multidimensional state space containing environmental information, target state, and own state, and quantify and encode it to form an input vector that the model can process; Construct an action-reward function and design a triple reward function based on task completion, resource consumption rate, and adversarial advantage to encourage the model to learn the optimal decision-making strategy; Set up a three-branch network structure including a policy network, a value network, and an environment modeling network and initialize the network parameters to provide initial conditions for model training.

5. According to the digital twin-based adversarial environment reinforcement learning model training method of claim 1, it is characterized in that: The method of using existing data to pre-train the reinforcement learning model to form a basic model includes: The existing adversarial environment data is divided into training set and validation set in a ratio of 7:3 to ensure the rationality and effectiveness of data division; A double-buffered experience pool mechanism is used for batch data sampling to improve data utilization efficiency and training stability; Set a dynamically decaying ε-greedy exploration strategy with an initial exploration rate of ε = 0.6 to balance the exploration and exploitation capabilities of the model; The early stopping mechanism can be used to terminate training in time when the accuracy of the validation set decreases, save the optimal parameters, and prevent model overfitting.

6. According to a digital twin-based adversarial environment reinforcement learning model training method according to claim 1, it is characterized in that: The method of inputting multiple virtual environments and target data generated by digital twins into the decision model for iterative training and optimizing network parameters using a policy gradient algorithm includes the following steps: The virtual environment data is clearly divided into three training stages: basic, medium, and difficult according to the complexity, to achieve course-based learning; The PPO algorithm is used to update the strategy, calculate the gradient of the model in each decision step, and evaluate the impact of the model's decision strategy on the reward through the gradient information; Update the network parameters based on the calculated gradients, so that the model adjusts its decision-making strategy in the direction of obtaining higher rewards and gradually optimizes the performance of the model; After each round of iteration, the proportion of high-difficulty samples is increased by 10% to gradually improve the adaptability of the model.

7. According to the digital twin-based adversarial environment reinforcement learning model training method of claim 1, it is characterized in that: The iterated model and the basic model are tested on the original adversarial environment data, including the following steps: Build a test benchmark set that contains 80% of the original environment data scenarios to ensure the comprehensiveness and representativeness of the test; Run the basic model and the iterative model in parallel and record the differences in the decision trajectories of the two models; The dual evaluation indicators of strategy entropy change and Q value distribution deviation are calculated to comprehensively evaluate the model performance.

8. According to the digital twin-based adversarial environment reinforcement learning model training method of claim 1, it is characterized in that: If the decision performance of the new model is improved, the iteration continues until it reaches the target. Otherwise, the parameters of the digital twin model are reversely corrected and the iteration continues, including the following steps: When the accuracy rate improves but does not meet the target, continue to generate multiple virtual adversarial environments through the digital twin model and iteratively train the reinforcement learning model; When the accuracy drops, the parameter correction loop of the digital twin model is quickly started: reduce the learning rate of the generator to slow down the learning speed of the generator; strengthen the spectral normalization constraint of the discriminator to improve the stability of the discriminator; inject 10% of the original data features into the generation process to ensure the authenticity of the generated data.

9. A digital twin-based adversarial environment reinforcement learning model training system, comprising: The digital twin model construction module is used to build a digital twin model based on the existing adversarial environment data and provide a basic model architecture for subsequent modules; The virtual confrontation environment generation module is connected with the digital twin model construction module to generate a variety of virtual confrontation environments based on the digital twin model to simulate real confrontation scenarios; Reinforcement learning model construction module, which builds the framework of the reinforcement learning model and provides the basic model architecture for subsequent training; The model pre-training module is connected to the reinforcement learning model building module, and uses the existing adversarial environment data to pre-train the reinforcement learning model to form a reinforcement learning basic model; The reinforcement learning model training iteration module is connected to the model pre-training module and the virtual adversarial environment generation module. It iterates the basic model through training in a variety of virtual adversarial environments to continuously improve the model performance. It is also connected to the test and evaluation module. When the model performance does not meet the standard, it will continue to iterate the training. It is also connected to the parameter correction module. When the model performance decreases, it will continue to iterate the training after passing the parameter correction module. The test and evaluation module is connected to the reinforcement learning model training iteration module to test and evaluate the model after training iteration to determine the quality of the model; The parameter correction module is connected to the test and evaluation module. When the performance of the reinforcement learning model deteriorates, the parameters of the digital twin model are reversely corrected. The final model output module is connected to the test and evaluation module to output the qualified model after multiple rounds of iterative training, evaluation and parameter correction, completing the entire training process.

Citation Information

Cited By

  • Training modeling development platform based on action tasks and solutions

    CN120255882A

  • Three-dimensional visual editing management method and system based on digital twinning

    CN120279195A

  • A three-dimensional visual editing management method and system based on digital twins

    CN120279195B

  • Power storage dynamic storage location distribution method and system fusing multi-dimensional factors

    CN120450375A

  • Power storage dynamic storage allocation method and system integrating multi-dimensional factors

    CN120450375B