Intelligent war game deduction decision-making method based on sample migration

By using the CycleGAN network framework in wargaming, and training the agent model under soft constraints, the problem of loss of decision-making knowledge during the migration process between different battlefield environments is solved, efficient migration and rapid adaptation are achieved, and the flexibility and efficiency of wargaming are improved.

CN120079113APending Publication Date: 2025-06-03XIDIAN UNIV HANGZHOU RES INST
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510116999.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In wargame deduction, during the migration process of the agent between different battlefield environments, data modification may lead to the loss of decision-making knowledge, reduce the effectiveness of knowledge transfer, and the model optimization ability of non-ideal training data is insufficient, making it easy to fall into local optimal solutions.

Method used

The migration network model based on the CycleGAN network framework is used to migrate sample data, generate migration data, and train the intelligent model under soft constraints to ensure that high-level decision-making knowledge is retained during the data migration process, and improve the effectiveness of knowledge migration and the robustness of algorithms.

Benefits of technology

Through this method, agents can efficiently migrate and adapt quickly in different deduction scenarios, improving the effectiveness of knowledge transfer and the robustness of algorithms, meeting the needs of cross-platform universality and rapid deployment of agents, and significantly improving the flexibility and efficiency of wargame deduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120079113A_ABST
    Figure CN120079113A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent war game deduction decision-making method based on sample migration, and belongs to the technical field of artificial intelligence and military deduction games. The method comprises the following steps: determining environment information of wargame deduction migration, and constructing a reference data set; constructing a migration network model based on a CycleGAN network framework, training the migration network model by using the reference data set, and generating migration data by using the trained migration network model; constructing a soft constraint condition according to the migration data; under the constraint of the soft constraint condition, the migration data and exploration data in the current combat scene state are utilized to train the agent model, and a trained agent model is obtained; the trained agent model is used for intelligent wargame deduction decision making. The method has higher sample migration similarity and knowledge transitivity, meanwhile, the problem that an existing method is poor in training effect under non-ideal data is solved, and the robustness of an algorithm is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and military deduction games, and particularly relates to an intelligent war game deduction decision-making method based on sample migration. Background Art

[0002] In the field of military deduction games, war game deduction, as a tool for simulating wars, originated from military training in the 19th century. Its core lies in reproducing the battlefield environment in a model-based manner to help commanders formulate strategies and conduct tactical analyses. With the development of technology, especially the progress of computer technology and artificial intelligence, war game deduction has gradually evolved into a digital and virtual form, and the introduction of intelligent agents has brought revolutionary changes to this field.

[0003] In military deduction games, intelligent agents are designed as entities that can perceive the environment, make decisions based on perception, and execute actions. Their goal is to simulate the complex decision-making processes in real wars through interaction with the environment, including troop deployment, tactical selection, target attack, etc. The importance of intelligent agent design in military deduction games is self-evident. First of all, they provide decision-making support for commanders. In a complex and ever-changing battlefield environment, commanders are faced with a large amount of information and choices. Intelligent agents can help commanders evaluate the potential consequences of various decisions by simulating different scenarios, improving the scientificity and accuracy of decisions. Secondly, intelligent agents contribute to enhancing the combat capabilities of troops. Through systematic deduction training, troops can become familiar with various combat scenarios and improve their ability to respond to emergencies. In addition, intelligent agents can also provide key support for war preparation. By deducing different war scenarios, the military can identify its own deficiencies and make targeted improvements and adjustments.

[0004] The intelligent agent cross-environment migration technology allows intelligent agents to migrate seamlessly between different battlefield environments or simulation systems and adapt to new battlefield conditions without re-training. This means that the strategies, experiences, and knowledge learned by intelligent agents in a specific battlefield environment can be efficiently applied to other battlefield environments, greatly improving the adaptability and flexibility of intelligent agents. Through the cross-environment migration technology, intelligent agents in military deduction games can more realistically simulate the complexity and uncertainty in wars and provide more comprehensive and in-depth decision-making support for commanders.

[0005] The problem of intelligent agent migration in war game deduction is essentially a problem of reinforcement learning model migration. At present, intelligent agent migration technology has been widely studied in fields such as robot control, motion navigation, and real-time strategy games. The main purpose of the research on intelligent agent transfer learning is to achieve the transfer of the model between the source domain and the target domain, ensuring that the transferred model can still play its expected role at a relatively small re-training cost and quickly adapt to changes in different scenarios.

[0006] According to the different transfer objects in transfer learning, this method is divided into the following two categories: knowledge transfer methods based on sample transfer and knowledge transfer methods based on policy transfer. Among them, the knowledge transfer method based on sample transfer focuses on the transfer and utilization of source domain samples. Specifically, it transfers excellent samples in the source domain to the sample experience pool of the target domain. Due to the inherent correlation between the source domain and the target domain for transfer, a large amount of rich sample data obtained in the source domain contains the dynamic model and reward model information of the target domain to a certain extent. Through the method of sample transfer, the source domain samples greatly expand the training sample pool of the target domain agents, provide rich example information for the target domain policies, and thus improve the training efficiency and reduce the demand for target domain samples. For the method based on policy transfer, the most common way is to directly transfer the entire policy network trained in the source domain to the target domain and fine-tune the agent's policy according to the samples generated by the interaction between the agent and the target domain. However, this method has a certain probability of causing the forgetting of the knowledge learned in the source domain and even causing the agent model to fail to converge. Based on this idea, many works disassemble the original agent policy and transfer the disassembled components.

[0007] In the deduction, according to the actual combat needs, the deduction needs to switch different combat scenarios. The switching of combat scenarios will cause changes in the battlefield information data set input to the agent, such as combat terrain data, combat meteorological data, the deployment of both sides' troops, etc. Due to the differences in the data sets between scenarios, during the cross-scenario transfer process, it is necessary to perform domain adaptation on the different parts of the sample data to conform to the data distribution of the target scenario. However, during the data domain adaptation process, the modification of the data may lead to the loss of decision-making knowledge and reduce the effectiveness of knowledge transfer. In addition, the insufficient model optimization ability for non-ideal training data is a common problem faced in transfer learning. When the transferred data contains noise, incorrect labels, data skew, or incomplete data, the performance of the model often cannot reach the optimal. Using the above transfer methods for knowledge transfer, the model is easily trapped in a local optimal solution during the transfer process. Summary of the Invention

[0008] To solve the above problems existing in the prior art, the present invention provides an intelligent war game deduction decision-making method based on sample transfer. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0009] In the first aspect, the present invention proposes an intelligent war game deduction decision-making method based on sample transfer, including:

[0010] Step 1: Determine the environmental information for war game deduction transfer and construct a reference data set;

[0011] Step 2: Construct a migration network model based on the CycleGAN (Cycle Generative Adversarial Networks) network framework, train the migration network model using the reference dataset, and generate migration data using the trained migration network model;

[0012] Step 3: Construct soft constraint conditions based on the migration data; and under the constraint of the soft constraint conditions, train the agent model using the migration data and the exploration data in the current combat scenario state to obtain a trained agent model; the trained agent model is used for intelligent wargame deduction decision-making.

[0013] In a second aspect, the present invention proposes an intelligent wargame deduction decision-making device based on sample migration, which is used to implement the intelligent wargame deduction decision-making method provided in the first aspect of the present invention. The device includes:

[0014] A data construction module, configured to determine the environmental information for wargame deduction migration and construct a reference dataset;

[0015] A sample migration module, configured to construct a migration network model based on the CycleGAN network framework, train the migration network model using the reference dataset, and generate migration data using the trained migration network model;

[0016] An agent training module, configured to construct soft constraint conditions based on the migration data, and under the constraint of the soft constraint conditions, train the agent model using the migration data and the exploration data in the current combat scenario state to obtain a trained agent model; the trained agent model is used for intelligent wargame deduction decision-making.

[0017] In a third aspect, the present invention proposes an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0018] The memory is used to store a computer program;

[0019] The processor is configured to execute the program stored on the memory to implement the method steps provided in the first aspect of the present invention.

[0020] In a fourth aspect, the present invention proposes a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps provided in the first aspect of the present invention are implemented.

[0021] Advantages of the present invention:

[0022] 1. The intelligent wargame deduction decision-making method based on sample transfer provided by the present invention, on the one hand, constructs a transfer network model based on the CycleGAN network framework, and uses the trained transfer network model to perform wargame deduction sample data transfer to obtain transfer data; on the other hand, constructs soft constraint conditions based on the obtained transfer data, and trains the agent model based on the soft constraint conditions, so as to use the trained agent model to make intelligent wargame deduction decisions. This method enables the data to be as close as possible to the target domain data during the transfer process, while retaining the high-level decision-making knowledge of the original data, has a higher sample transfer similarity and knowledge transferability, and improves the effectiveness of knowledge transfer; at the same time, uses the agent training method based on soft constraints. During agent training, transfer data is used to construct training constraints, overcoming the problem of poor training effect of existing methods under non-ideal data, improving the robustness of the algorithm, and realizing the efficient transfer and rapid adaptation of the agent in different deduction scenarios;

[0023] 2. The intelligent wargame deduction decision-making method based on sample transfer provided by the present invention can meet the requirements of the deduction system for the cross-platform generality and rapid deployment of agents in the field of cross-scenario wargame deduction, and significantly improve the flexibility and efficiency of wargame deduction.

[0024] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0025] Figure 1 is a flowchart of the intelligent wargame deduction decision-making method based on sample transfer provided by the embodiment of the present invention;

[0026] Figure 2 is a schematic diagram of the transfer network model and sample transfer based on the CycleGAN network framework provided by the embodiment of the present invention;

[0027] Figure 3 is a flowchart of agent training based on soft constraints provided by the embodiment of the present invention;

[0028] Figure 4 is a structural block diagram of the intelligent wargame deduction decision-making device based on sample transfer provided by the embodiment of the present invention. Detailed Embodiments

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] The first aspect of the present invention provides an intelligent wargame deduction decision-making method based on sample migration. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the intelligent wargame deduction decision-making method based on sample migration provided by an embodiment of the present invention. The method mainly includes the following steps:

[0031] Step 1: Determine the environmental information for wargame deduction migration and construct a reference data set.

[0032] 11) Determine the source domain and target domain for wargame deduction migration.

[0033] Specifically, for the source domain and target domain of wargame deduction migration, it is required that the two scenarios have a certain degree of similarity, and there is common knowledge in the combat decision-making of the intelligent agent, which generally refers to combat factors such as combat type, combat arms, and combat objectives.

[0034] 12) Prepare a large amount of source domain data and part of the target domain data according to the migration background.

[0035] Specifically, according to the migration background, a large amount of source domain data and part of the target domain data need to be prepared. The data can be data reflecting battlefield information such as combat terrain data, combat meteorological data, and the deployment of both sides' forces, and it is required that the prepared data conforms to the basic format of the deduction data, and the source domain data is required to have a high decision-making level to ensure that the migration data can guide the training of the intelligent agent.

[0036] 13) Perform data cleaning and data shaping on the source domain data and the target domain data to obtain a reference data set.

[0037] Specifically, illegal data and invalid data are removed from the prepared large amount of source domain data and part of the target domain data, and then data shaping is performed according to the requirements of subsequent algorithms. All the processed data forms a reference data set to be used as subsequent training data.

[0038] Step 2: Construct a migration network model based on the CycleGAN network framework, train the migration network model using the reference data set, and generate migration data using the trained migration network model.

[0039] 21) Construct a migration network model based on the CycleGAN network framework.

[0040] Optionally, as an implementation method, in this embodiment, two discriminators are respectively constructed to distinguish the real data and the corresponding generated data of the source domain and the target domain; at the same time, two generators are constructed to generate migration data based on the target domain and the source domain data respectively to complete the migration generation of the data.

[0041] Specifically, please refer to Figure 2 , Figure 2It is the migration network model and sample migration schematic diagram based on the CycleGAN network framework provided by the embodiments of the present invention.

[0042] In this embodiment, the migration network model constructed based on the CycleGAN network framework includes a first generator G Y , a second generator G X , a first discriminator D Y and a second discriminator D X ;

[0043] The first generator G Y is used to generate migration data G(x) based on source domain data;

[0044] The second generator G X is used to generate restored data F(G(x)) based on target domain data;

[0045] The first discriminator D Y is used to discriminate between source domain data and the migration data G(x) generated by the first generator G Y ;

[0046] The second discriminator D X is used to discriminate between target domain data and the restored data F(G(x)) generated by the second generator G X .

[0047] 22) Construct a loss function.

[0048] Specifically, based on the constructed migration network model, discriminator constraints, generative adversarial constraints, and cycle consistency constraints are respectively defined, and a loss function of the migration network model is constructed according to the corresponding constraints.

[0049] For the discriminator, it is required that the discriminator can discriminate between real data and migration-generated data as much as possible. Then, the loss function constructed according to the discriminator constraint is expressed as:

[0050]

[0051] In the formula, represents the loss function of the first discriminator D Y , represents the loss function of the second discriminator D X , represents the expectation of the real data y, where y is sampled from the real data distribution Y, Y represents the target domain data set, and y represents the data in Y, represents the expectation of the input data x, where x is sampled from the input data distribution X, X represents the source domain data set, and x represents the data in X, D Y () represents the first discriminator D YThe probability of determining that the data is from Y, D X () represents the second discriminator D X The probability of determining that the data is from X, G Y () represents the first generator G Y The generated target domain data, G X () represents the second generator G X The generated source domain data;

[0052] For the generator network, there are two requirements. On the one hand, the generator and the discriminator need to generate adversarially to deceive the discriminator and migrate to obtain as real data as possible. On the other hand, during the migration process, it is necessary to ensure that the high-quality decision-making knowledge hidden in the migration samples can be retained.

[0053] Taking the first generator G Y as an example, based on the reference data x, the data G(x) is migrated and generated. It is expected that the migrated data can be close to the target data and confuse the discriminator. Based on the discriminator feedback, the generative adversarial constraint is defined, and the loss function is defined as follows:

[0054]

[0055] In the formula, represents the loss function of the first generator G Y .

[0056] Correspondingly, the loss function of the second generator G X is defined as follows:

[0057]

[0058] In the formula, represents the loss function of the second generator G X .

[0059] Furthermore, a cyclic consistency constraint is constructed between the reference samples and the restored samples to ensure the retention of decision-making knowledge during the migration process, and the cyclic consistency constraint loss function is expressed as:

[0060] L cyc =λ 1 L cyc1 (G X ,G Y )+λ 2 L cyc2 (G X ,G Y );

[0061]

[0062] In the formula, L cycRepresents the cyclic consistency constraint loss function, λ 1 , λ 2 respectively represent the weights of the two loss functions, d(,) represents the difference between two samples, and V x represents the value evaluation of sample x in the reinforcement learning value network, represents sample G X (G Y (x)) represents the value evaluation of sample G t (x) in the reinforcement learning value network, and π(a t |x) represents the policy output of sample x in the reinforcement learning policy network, and a t represents the specific decision-making action of sample x in the reinforcement learning value network.

[0063] In the cyclic consistency constraint loss function constructed in this embodiment, L cyc1 (G X , G Y ) is constructed based on the reference data and the restored data to avoid drastic deviation of the data during training and loss of decision-making knowledge. L cyc2 (G X , G Y ) is constructed based on the reinforcement learning value network and the policy network to retain the decision-making knowledge between samples.

[0064] 23) Perform network training and data migration.

[0065] Specifically, use the above-defined network model and the data prepared in step 1 for model training. During the training process, the training parameters can be adjusted according to the training effect.

[0066] After the model training is completed, the model can be used to generate target domain data with the source domain data as the input to achieve cross-scenario data migration.

[0067] This embodiment uses the CycleGAN framework to achieve cross-scenario migration of wargame data, which has higher sample migration similarity and knowledge transferability compared with the existing sample migration methods.

[0068] Step 3: Construct soft constraint conditions based on the migration data; and under the constraints of the soft constraint conditions, use the migration data and the exploration data in the current combat scenario state to train the agent model to obtain a trained agent model; the trained agent model is used for intelligent wargame decision-making.

[0069] In this embodiment, the agent can perform online reinforcement learning in the target scenario. At the same time, based on the migration data, training soft constraints can be constructed to guide the agent training, so as to achieve accelerated training of the agent while ensuring the training effect.

[0070] In this embodiment, based on the migrated data, a soft constraint on the decision-making level of the agent can be constructed. The specific constraint construction can refer to the following formula:

[0071]

[0072] In the formula, θ k+1 represents the model parameters in the (k + 1)-th round during the reinforcement learning iteration process, Based on the model optimization objective under the policy That is, the long-term deduction benefit, respectively represent the reinforcement learning policies in the k-th and (k + 1)-th rounds, D represents the difference between reinforcement learning policies, represents the occupancy measure of the countermeasure y for evaluating the policy difference, represents the occupancy measure of the migrated data, s represents the combat situation, a represents the next action taken by the policy, d k represents the dynamic policy constraint in the k-th round during the iteration process, D KL represents the KL divergence between policies for evaluating the difference between policies, represents the probability of the situation action under the policy δ represents the update step size constraint between the iteration policies.

[0073] Among them, the first constraint conditions the difference between the exploration data used in the reinforcement learning process and the migrated decision-making data, and the second constraint conditions the training step size of the reinforcement learning to avoid the learning oscillation problem caused by large-step learning.

[0074] Under the above soft constraint conditions, according to whether the exploration data satisfies the soft constraint conditions, it is determined whether the training method for the current training round is reinforcement learning or imitation learning. After updating the decision model, new exploration data is collected using the new model, and this behavior is repeated until the training ends.

[0075] Specifically, please refer to Figure 3 , Figure 3 which is the flowchart of the agent training based on soft constraints provided by the embodiment of the present invention.

[0076] a) In the current training round, if the exploration data satisfies the soft constraint conditions, online reinforcement learning is performed on the agent model using the exploration data; otherwise, imitation learning is performed on the agent model using the migrated data;

[0077] b) After completing the training of the current round, the agent model is updated;

[0078] c) Determine whether the maximum training round has been reached. If so, the training ends and the trained agent model is obtained; otherwise, continue to collect the exploration data in the current combat scenario state and return to step a) for the next round of training.

[0079] Based on the migrated samples, this embodiment uses the migrated data as soft constraints for reinforcement learning training, overcoming the problem of local optimal solutions caused by non-ideal data training and improving the robustness of the training algorithm.

[0080] It should be noted that the specific implementation methods of online reinforcement learning and imitation learning are both mature technologies in this field and can be implemented by referring to existing relevant technologies. This embodiment will not introduce them in detail here.

[0081] After the agent training is completed, intelligent wargame deduction decisions can be made according to the current battlefield information data.

[0082] The intelligent wargame deduction decision-making method based on sample migration provided by the present invention, on the one hand, constructs a migration network model based on the CycleGAN network framework, and uses the trained migration network model to migrate the wargame deduction sample data to obtain migrated data; on the other hand, constructs soft constraint conditions according to the obtained migrated data, and trains the agent model based on the soft constraint conditions, so as to use the trained agent model to make intelligent wargame deduction decisions. This method enables the data to be as close as possible to the target domain data during the migration process, while retaining the high-level decision-making knowledge of the original data, having a higher sample migration similarity and knowledge transferability, and improving the effectiveness of knowledge transfer; at the same time, using the agent training method based on soft constraints, in the agent training, using the migrated data to construct training constraints, overcomes the problem of poor training effect of the existing method under non-ideal data, improves the robustness of the algorithm, and realizes the efficient migration and rapid adaptation of the agent in different deduction scenarios. In the field of cross-scenario wargame deduction, this method can meet the requirements of the deduction system for the cross-platform generality and rapid deployment of the agent, and significantly improve the flexibility and efficiency of wargame deduction.

[0083] Based on the same inventive concept, the second aspect of the present invention also provides an intelligent wargame deduction decision-making device based on sample migration, which is used to implement the intelligent wargame deduction decision-making method based on sample migration provided in the first aspect of the present invention. Please refer to Figure 4 , Figure 4 which is the structural block diagram of the intelligent wargame deduction decision-making device based on sample migration provided in the embodiment of the present invention. The device includes:

[0084] A data construction module, configured to determine the environmental information of wargame deduction migration and construct a reference data set;

[0085] A sample migration module, configured to construct a migration network model based on the CycleGAN network framework, train the migration network model using the reference data set, and generate migrated data using the trained migration network model;

[0086] An agent training module is used to construct soft constraint conditions based on migration data, and under the constraints of the soft constraint conditions, train an agent model using the migration data and exploration data in the current combat scenario state to obtain a trained agent model; the trained agent model is used for intelligent military wargame deduction decision-making.

[0087] Based on the same inventive concept, a third aspect of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus.

[0088] The memory is used to store a computer program.

[0089] When the processor is used to execute the program stored on the memory, it implements the method steps provided in the first aspect of the present invention.

[0090] The intelligent military wargame deduction decision-making method based on sample migration provided by the embodiments of the present invention can be applied to an electronic device. Specifically, the electronic device can be: a desktop computer, a portable computer, an intelligent mobile terminal, a server, etc. There is no limitation here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.

[0091] Based on the same inventive concept, a fourth aspect of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed, it can implement the method steps provided in the first aspect of the present invention.

[0092] For the device / electronic device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.

[0093] It should be noted that the device, electronic device, and storage medium of the embodiments of the present invention are respectively the device, electronic device, and storage medium applying the above-mentioned intelligent military wargame deduction decision-making method based on sample migration. Then, all embodiments of the above-mentioned intelligent military wargame deduction decision-making method based on sample migration are applicable to the device, electronic device, and storage medium, and can achieve the same or similar beneficial effects.

[0094] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices (equipment), or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects, which are collectively referred to herein as "modules" or "systems". Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The computer program is stored / distributed in a suitable medium, provided together with other hardware or as part of the hardware, and can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0095] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An intelligent war game deduction decision method based on sample migration, characterized in that: include: Step 1: Determine the environment information for war game simulation migration and build a reference dataset; Step 2: construct a migration network model based on the CycleGAN network framework, train the migration network model using the reference data set, and generate migration data using the trained migration network model; Step 3: construct soft constraints based on the migration data; And under the constraints of the soft constraints, the intelligent agent model is trained using the migration data and the exploration data under the current combat scenario state to obtain a trained intelligent agent model; the trained intelligent agent model is used for intelligent war game simulation decision-making.

2. The intelligent war game deduction decision method based on sample migration according to claim 1 is characterized in that: Step 1 specifically includes: Determine the source and target domains for wargame transfer; Prepare a large amount of source domain data and some target domain data according to the migration context; The source domain data and the target domain data are cleaned and shaped to obtain a reference data set.

3. The intelligent war game deduction decision method based on sample migration according to claim 2 is characterized in that: In step 2, a migration network model constructed based on the CycleGAN network framework includes a first generator, a second generator, a first discriminator, and a second discriminator; The first generator is used to generate migration data based on source domain data; The second generator is used to generate restored data based on the target domain data; The first discriminator is used to discriminate between source domain data and migration data generated by the first generator; The second discriminator is used to discriminate between the target domain data and the restored data generated by the second generator.

4. The intelligent war game deduction decision method based on sample migration according to claim 3 is characterized in that: In step 2, before training the migration network model using the reference data set, it also includes: Define discriminator constraints, generative adversarial constraints, and cycle consistency constraints, and construct the loss function of the migration network model based on the corresponding constraints.

5. The intelligent war game deduction decision method based on sample migration according to claim 4 is characterized in that: The loss function constructed according to the discriminator constraints is expressed as: In the formula, Denotes the first discriminator D Y The loss function is Denotes the second discriminator D X The loss function is represents the expectation of the real data y, Y represents the target domain dataset, y represents the data in Y, represents the expectation of the input data x, X represents the source domain dataset, x represents the data in X, and D Y () represents the first discriminator D Y The probability that the data comes from Y, D X () represents the second discriminator D X The probability that the data comes from X, G Y () represents the first generator G Y The generated migration data, G X () represents the second generator G X Generated restore data; The loss function constructed according to the generative adversarial constraint is expressed as: In the formula, Denotes the second generator G X The loss function is Denotes the first generator G Y The loss function of The loss function constructed according to the cycle consistency constraint is: L cyc =λ1L cyc1 (G X ,G Y )+λ2L cyc2 (G X ,G Y ); Where, L cyc represents the cycle consistency constraint loss function, λ1 and λ2 represent the weights of the two loss functions, d(,) represents the difference between two samples, V x Represents the value evaluation of sample x in the reinforcement learning value network, Represents sample G X (G Y (x)) value evaluation in the reinforcement learning value network, π(a t |x) represents the policy output of sample x in the reinforcement learning policy network, a t Represents the specific decision action of sample x in the reinforcement learning value network.

6. The intelligent war game deduction decision method based on sample migration according to claim 1 is characterized in that: In step 3, the soft constraint condition constructed according to the migration data is expressed as: In the formula, θ k+1 represents the model parameters of the k+1th round in the reinforcement learning iteration process, Strategy-based The model optimization goal under , that is, the long-term deduction benefit, They represent the reinforcement learning strategies of the kth round and the k+1th round respectively, and D represents the difference in reinforcement learning strategies. Express countermeasures Occupancy measure for evaluating strategy differences, represents the occupancy measure of the migration data, s represents the combat situation, a represents the next action adopted by the strategy, and d k represents the dynamic strategy constraint of round k in the iteration process, D KL Represents the KL divergence between strategies, which is used to evaluate the differences between strategies. Indicated in strategy The probability of the situation action under , δ represents the update step size constraint between iterative strategies.

7. The intelligent war game deduction decision method based on sample migration according to claim 1 is characterized in that: In step 3, the agent model is trained using the migration data and the exploration data in the current combat scenario state, specifically including: a) in the current training round, if the exploration data satisfies the soft constraint condition, the exploration data is used to perform online reinforcement learning on the agent model; otherwise, the migration data is used to perform imitation learning on the agent model; b) after completing the current round of training, updating the agent model; c) Determine whether the maximum number of training rounds has been reached. If so, end the training and obtain a trained agent model; otherwise, continue to collect exploration data under the current combat scenario state and return to step a) for the next round of training.

8. An intelligent war game deduction and decision-making device based on sample migration, used to implement the intelligent war game deduction and decision-making method based on sample migration as described in any one of claims 1 to 7, characterized in that: The device includes: The data construction module is used to determine the environmental information for war game migration and build a reference data set; A sample migration module, used to build a migration network model based on the CycleGAN network framework, train the migration network model using the reference data set, and generate migration data using the trained migration network model; The intelligent agent training module is used to construct soft constraints based on the migration data, and under the constraints of the soft constraints, use the migration data and the exploration data under the current combat scenario state to train the intelligent agent model to obtain a trained intelligent agent model; the trained intelligent agent model is used for intelligent war game simulation decision-making.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to execute the program stored in the memory to implement the method steps described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method steps described in any one of claims 1 to 7 can be implemented.

Citation Information

Cited By

  • Generative confrontation deduction method for safety events for industrial production

    CN120373475A

  • Multi-modal cross-domain wargame deduction decision data construction method and device based on trajectory logic migration

    CN121960233A

  • A multi-modal cross-domain war game deduction decision data construction method and device based on trajectory logic migration

    CN121960233B