Multi-task intelligent decision-making computing method and system based on meta-learning and game theory

By employing a multi-task intelligent decision-making computation method based on meta-learning and game theory, the problem of military confrontation game decision-making under incomplete information conditions is solved, thereby optimizing combat strategies and improving robustness, and assisting commanders in generating efficient decisions.

CN116341924BActive Publication Date: 2025-12-16BEIJING MECHANICAL EQUIP INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310152631.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-12-16
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

Existing technologies for military adversarial game decision-making under conditions of incomplete information fail to fully consider the influence of multiple factors, lack effective theoretical basis, and are difficult to predict decision results and battle losses and gains. Furthermore, the design of human-machine intelligent game system modules has not explored the mechanism of adversarial strategy generation in depth.

Method used

A multi-task intelligent decision-making computation method based on meta-learning and game theory is adopted. By initializing basic combat information, the combat decision-making system is trained locally and globally. By using the adversarial settings of generative adversarial networks and discriminators, optimized combat strategies are generated, and human-computer interaction modules are combined to assist in decision-making.

Benefits of technology

It improves the comprehensiveness and accuracy of operational decision-making, enables the optimization of decisions under conditions of incomplete information, generates efficient combat strategies, assists commanders in making optimal decisions, and enhances the effectiveness and robustness of military operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116341924B_ABST
    Figure CN116341924B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a multi-task intelligent decision-making computing method and system based on meta-learning and game theory, an electronic device, and a storage medium. The method comprises: inputting preset combat basic information into a combat decision-making system; classifying the preset combat basic information according to combat tasks, and taking the combat basic information support set as input to locally update and train the output layer and decision layer of the combat decision-making system; taking the query set of the preset combat basic information as input to globally train the combat decision-making system, and update all parameters of the combat decision-making system to complete the training of the combat decision-making system. The present disclosure mines potential combat effectiveness in the way of game confrontation and deep learning, comprehensively considers combat influencing factors and combat costs to assist commanders to generate optimal combat decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of operational decision-making, and more specifically, to a multi-task intelligent decision-making computation method, system, electronic device, and computer-readable storage medium based on meta-learning and game theory. Background Technology

[0002] Modern military warfare is influenced by a multitude of factors, including social factors such as politics, economics, law, and culture; environmental factors such as geographical location, weather conditions, and topography; and human factors such as the different branches of the armed forces, air force, navy, and rocket force, and the different types of weapons and equipment. Commanders must consider these multiple factors and make effective judgments about their primary and secondary influences and their degree of impact in order to make sound decisions and maintain control of the battlefield.

[0003] In the early stages of the war, the preliminary analysis, strategic decisions, and issuance of orders for command largely relied on the subjective judgment of the commanders. Sole subjective judgment is insufficient to fully consider influencing factors, lacks effective theoretical basis, and cannot predict the outcome of the decision, as well as the losses and gains of both sides.

[0004] With the continuous development of information technology, game theory and deep learning have been gradually applied to various fields and have achieved remarkable results. Game theory, initially proposed by mathematician John von Neumann and applied to the economic field, is essentially a decision-making theory describing the interaction and interdependence of opposing sides. Modern military warfare is characterized by certain information, uncertain conditions, constantly changing situations, strong adversarial nature, and poor stability. Using game theory-based decision optimization methods allows for the statistical analysis and simulation of combat decisions under conditions of limited information completeness, calculating the probabilities and costs of different decisions, thereby continuously optimizing the model and ultimately outputting the optimal decision.

[0005] To address the aforementioned problems faced in modern military warfare, this invention proposes an intelligent decision-making computation method based on meta-learning and game theory. It constructs a basic adversarial model for both sides and an adversarial game-theoretic intelligent learning model based on a deep generative adversarial network (GAN), setting initial model parameters. The model is trained using the basic adversarial model, the adversarial game-theoretic intelligent learning model, the initial parameters, and multi-type task data generated during the adversarial process. This trains the model to analyze, predict, and generate scenarios involving both sides and incomplete battlefield situations under adversarial conditions. Relying on the adversarial settings between the generator and discriminator of the GAN, a new, optimized version of the combat strategy is generated and input back into the model for training and cost calculation. This iterative process yields the optimal model parameters and the final combat strategy information.

[0006] In existing technologies, intelligent decision-making methods for military adversarial games under incomplete information conditions identify and predict incomplete information in military adversarial scenarios, transforming it into complete information to obtain the final decision. This method first constructs a basic model of military adversarial game decision dynamics and an intelligent learning model based on deep learning and self-play. These two models predict the battlefield situation with incomplete information and obtain an optimized decision according to an intelligent optimization decision-making mode of "decision-feedback-dynamic optimization." This invention focuses on the generation process of military intelligent decision-making but does not provide related system-level explanations. The proposed human-machine intelligent game system includes a deducer decision module, an intelligent agent framework module, and a deduction room module. The deducer decision module determines a set of actions based on the situation information input by the intelligent agent framework module. The intelligent agent framework module inputs the situation information sent by the deduction room module to the deducer decision module and feeds back the generated action set to the deduction room module. The deduction room module applies the action set to the simulation environment to obtain time-varying situation information provided by the environment. This invention mainly focuses on the design of functional modules at the level of human-machine intelligent game system, but it does not further explore the generation mechanism of adversarial strategies or their impact on other modules.

[0007] Therefore, one or more methods are needed to solve the above problems.

[0008] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0009] The purpose of this disclosure is to provide a multi-task intelligent decision computing method, system, electronic device, and computer-readable storage medium based on meta-learning and game theory, thereby overcoming, at least to some extent, one or more problems caused by the limitations and defects of related technologies.

[0010] According to one aspect of this disclosure, a multi-task intelligent decision-making computation method based on meta-learning and game theory is provided, comprising:

[0011] Step S110: Input preset basic operational information into the operational decision-making system and initialize the operational decision-making system. The preset basic operational information includes a support set and a query set.

[0012] Step S120: Classify the preset basic combat information according to the combat mission, and use the support set of the preset basic combat information corresponding to the classification of the combat mission as input to perform local update training on the output layer and decision layer of the combat decision system, and update the parameters of the output layer and decision layer of the combat decision system.

[0013] Step S130: Based on the output layer and decision layer of the combat decision system that have completed parameter updates after local update training, the combat decision system is globally trained using the preset query set of basic combat information as input, and all parameters of the combat decision system are updated to complete the training of the combat decision system.

[0014] In one exemplary embodiment of this disclosure, the preset basic operational information in the method includes seat information, situational information, facilities, weapons, troop strength and their deployment, operational targets, and operational tasks.

[0015] In one exemplary embodiment of this disclosure, the method further includes:

[0016] The combat decision system is initialized and configured, and scoring conditions and scoring rules are set for the combat decision system.

[0017] In one exemplary embodiment of this disclosure, the method further includes performing local update training on the output layer and decision layer of the combat decision system:

[0018] Extract the input features from the support set and concatenate the input features;

[0019] The input features are calculated based on a pre-defined generative adversarial network to obtain the training loss and gradient;

[0020] The parameters of the output layer and decision layer of the combat decision system are updated based on the training loss and gradient.

[0021] In one exemplary embodiment of this disclosure, the method further includes performing local update training on the output layer and decision layer of the combat decision system based on game theory:

[0022] Step S121: Input the situational information, target or mission information of the preset basic combat information into the generator of the combat decision system;

[0023] Step S122: Based on preset initial parameters, use the generator model to generate decision instruction information;

[0024] Step S123: Input the decision instruction information generated by the generator and the actual label instruction into the discriminator. The discriminator will then distinguish between the decision instruction information generated by the generator and the decision instruction information used as labels, and output the discrimination result.

[0025] Step S124: Calculate the loss function based on the discrimination result and optimize it based on the gradient information;

[0026] Step S125: Repeat steps S121 to S124 for iterative calculation until the Nash equilibrium state is determined based on the discrimination result.

[0027] In one exemplary embodiment of this disclosure, the method further includes performing global training on the combat decision system:

[0028] Based on the pre-defined basic combat information of the full classification, and using the pre-defined historical game data information as the input of the combat decision system, the combat decision system is subjected to multi-task meta-enhanced global training.

[0029] The global training is accelerated by manually adjusting the meta-reinforcement learning process online.

[0030] In one exemplary embodiment of this disclosure, the method further includes performing global training on the combat decision system:

[0031] Based on the pre-defined basic combat information of the full classification, meta-knowledge is extracted for the generation of multi-task strategies, and a query set sample including meta-knowledge is generated.

[0032] The combat decision system is trained globally using the preset query set of basic combat information as input.

[0033] In one aspect of this disclosure, a multi-task intelligent decision-making computing system based on meta-learning and game theory is provided, comprising a situation information input module, a database module, an intelligent decision generation module, and a human-computer interaction module, wherein:

[0034] The situation information input module is used to import or set the situation information of the two sides in the game, including seats, seat relationships, facilities, weapons, troop strength (i.e. troop deployment), objectives, and tasks.

[0035] The database module is used for storing situational information and searching and matching information, including knowledge extraction, knowledge refinement, and knowledge representation, to provide data support for intelligent decision-making based on game-based adversarial competition.

[0036] The intelligent decision generation module is used to calculate and generate decision schemes based on the information provided by the situation input module and the database module, and to recommend and rank the decision schemes according to their probability of winning.

[0037] The human-computer interaction module includes a visualization analysis module, a voice and text interaction module, and a brain signal processing module. The human-computer interaction module is used to provide personnel with information visualization functions and assists personnel in decision-making through human-computer interaction modes such as voice interaction, text interaction, gesture interaction, human brain signal processing, and visual analysis.

[0038] In one aspect of this disclosure, an electronic device is provided, comprising:

[0039] Processor; and

[0040] A memory storing computer-readable instructions that, when executed by the processor, implement the method according to any one of the preceding claims.

[0041] In one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of the preceding claims.

[0042] An exemplary embodiment of this disclosure provides a multi-task intelligent decision-making computation method based on meta-learning and game theory. The method includes: inputting preset basic operational information into a combat decision-making system; classifying the preset basic operational information according to combat tasks and using a support set of the basic operational information as input to perform local update training on the output layer and decision layer of the combat decision-making system; using a query set of the preset basic operational information as input to perform global training on the combat decision-making system and update all parameters of the combat decision-making system, thus completing the training of the combat decision-making system. This disclosure assists commanders in generating optimal combat decisions by mining potential combat effectiveness and comprehensively considering combat influencing factors and costs through game theory and deep learning.

[0043] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0044] The above and other features and advantages of this disclosure will become more apparent from the detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0045] Figure 1 A flowchart is shown for a multi-task intelligent decision computing method based on meta-learning and game theory according to an exemplary embodiment of the present disclosure;

[0046] Figure 2A-2B A model diagram of a multi-task intelligent decision computing method based on meta-learning and game theory according to an exemplary embodiment of the present disclosure is shown.

[0047] Figure 3 A generative adversarial network schematic diagram of a multi-task intelligent decision computing method based on meta-learning and game theory according to an exemplary embodiment of the present disclosure is shown.

[0048] Figure 4 A schematic block diagram of a multi-task intelligent decision computing system based on meta-learning and game theory according to an exemplary embodiment of the present disclosure is shown.

[0049] Figure 5 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is schematically shown; and

[0050] Figure 6 The illustration shows a schematic diagram of a computer-readable storage medium according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0051] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0052] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details described, or other methods, components, materials, apparatuses, steps, etc., can be employed. In other instances, well-known structures, methods, apparatuses, implementations, materials, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0053] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, or in one or more software-hardened modules, or in different network and / or processor devices and / or microcontroller devices.

[0054] In this example embodiment, a multi-task intelligent decision-making computation method based on meta-learning and game theory is first provided; (Refer to...) Figure 1 As shown, this multi-task intelligent decision-making computation method based on meta-learning and game theory may include the following steps:

[0055] Step S110: Input preset basic operational information into the operational decision-making system and initialize the operational decision-making system. The preset basic operational information includes a support set and a query set.

[0056] Step S120: Classify the preset basic combat information according to the combat mission, and use the support set of the preset basic combat information corresponding to the classification of the combat mission as input to perform local update training on the output layer and decision layer of the combat decision system, and update the parameters of the output layer and decision layer of the combat decision system.

[0057] Step S130: Based on the output layer and decision layer of the combat decision system that have completed parameter updates after local update training, the combat decision system is globally trained using the preset query set of basic combat information as input, and all parameters of the combat decision system are updated to complete the training of the combat decision system.

[0058] An exemplary embodiment of this disclosure provides a multi-task intelligent decision-making computation method based on meta-learning and game theory. The method includes: inputting preset basic operational information into a combat decision-making system; classifying the preset basic operational information according to combat tasks and using a support set of the basic operational information as input to perform local update training on the output layer and decision layer of the combat decision-making system; using a query set of the preset basic operational information as input to perform global training on the combat decision-making system and update all parameters of the combat decision-making system, thus completing the training of the combat decision-making system. This disclosure assists commanders in generating optimal combat decisions by mining potential combat effectiveness and comprehensively considering combat influencing factors and costs through game theory and deep learning.

[0059] The following will further explain a multi-task intelligent decision computing method based on meta-learning and game theory in this example embodiment.

[0060] Example 1:

[0061] In step S110, preset basic operational information can be input into the operational decision-making system, and the operational decision-making system can be initialized. The preset basic operational information includes a support set and a query set.

[0062] In this example embodiment, the preset basic operational information in the method includes seat information, situational information, facilities, weapons, troop strength and their deployment, operational targets, and operational tasks.

[0063] In this example embodiment, the method further includes:

[0064] The combat decision system is initialized and configured, and scoring conditions and scoring rules are set for the combat decision system.

[0065] In step S120, the preset basic combat information can be classified according to the combat mission, and the support set of the preset basic combat information corresponding to the classification of the combat mission can be used as input to perform local update training on the output layer and decision layer of the combat decision system, and update the parameters of the output layer and decision layer of the combat decision system.

[0066] In this example embodiment, the method for locally updating and training the output layer and decision layer of the combat decision system further includes:

[0067] Extract the input features from the support set and concatenate the input features;

[0068] The input features are calculated based on a pre-defined generative adversarial network to obtain the training loss and gradient;

[0069] The parameters of the output layer and decision layer of the combat decision system are updated based on the training loss and gradient.

[0070] In this example embodiment, the method further includes performing local update training on the output layer and decision layer of the combat decision-making system based on game theory:

[0071] Step S121: Input the situational information, target or mission information of the preset basic combat information into the generator of the combat decision system;

[0072] Step S122: Based on preset initial parameters, use the generator model to generate decision instruction information;

[0073] Step S123: Input the decision instruction information generated by the generator and the actual label instruction into the discriminator. The discriminator will then distinguish between the decision instruction information generated by the generator and the decision instruction information used as labels, and output the discrimination result.

[0074] Step S124: Calculate the loss function based on the discrimination result and optimize it based on the gradient information;

[0075] Step S125: Repeat steps S121 to S124 for iterative calculation until the Nash equilibrium state is determined based on the discrimination result.

[0076] In step S130, based on the output layer and decision layer of the combat decision system that have completed parameter updates after local update training, the combat decision system can be globally trained using the preset query set of basic combat information as input, and all parameters of the combat decision system can be updated to complete the training of the combat decision system.

[0077] In this example embodiment, the method for globally training the combat decision system further includes:

[0078] Based on the pre-defined basic combat information of the full classification, and using the pre-defined historical game data information as the input of the combat decision system, the combat decision system is subjected to multi-task meta-enhanced global training.

[0079] The global training is accelerated by manually adjusting the meta-reinforcement learning process online.

[0080] In this example embodiment, the method for globally training the combat decision system further includes:

[0081] Based on the pre-defined basic combat information of the full classification, meta-knowledge is extracted for the generation of multi-task strategies, and a query set sample including meta-knowledge is generated.

[0082] The combat decision system is trained globally using the preset query set of basic combat information as input.

[0083] Example 2:

[0084] In this example embodiment, the combat personnel input the combat scenario into the system or set it themselves in the system, including seat information, situational information, facilities, weapons, forces and their deployment, combat objectives and combat missions, etc.; and set scoring conditions and scoring rules.

[0085] In the embodiments of this example, as Figure 2A-2B The diagram illustrates the training of a multi-task decision-making intelligent generative model based on meta-learning. During the learning and training of multi-task strategies, the dataset for each task type is first divided into a support set and a query set. The support set is used for local training and parameter updates for each task type, while the query set is used for global model updates. During the local update phase, only the parameters of the decision layer and output layer are updated. After each task type completes its local update, the query set is used for joint training of the model, at which point all model parameters are updated.

[0086] In this example embodiment, for each type of task, a set of reaction actions is set, and adversarial strategies are generated and trained using a game theory-based generative adversarial network according to the scoring conditions and scoring rules. The network is then iterated and evolved multiple times to optimize the quality and efficiency of the generated strategies.

[0087] In this example embodiment, for multiple types of combat missions, historical game data information is input, and multi-task meta-reinforcement learning training is carried out. The process is accelerated by human online adjustment of the auxiliary meta-reinforcement learning process.

[0088] In this example embodiment, the input layer and embedding layer are treated as a whole, with their corresponding training parameters denoted by θ1. The decision layer and output layer are treated as a whole, with their corresponding training parameters denoted by θ2. The embedding layer extracts useful features from the input of combat actions and situations, learns the association between these features and the decision-making process, and concatenates the extracted feature vectors. Intelligent decision generation is performed by the decision layer and output layer, where the decision layer is a generative adversarial network (GAN), and the output layer follows the decision layer. Borrowing from the two-layer loop concept of MAML, the "inner loop" process, i.e., the local update process, only trains and updates the parameters of the decision layer and output layer, and in this system, it is mainly responsible for simulating the model learning new decision strategies. The "outer loop," corresponding to the aforementioned global update, is responsible for collecting the effects after the "inner loop" learning and using the query set to train the overall parameters.

[0089] Furthermore, the multi-task combat decision learning and training process based on meta-learning can be represented by the following steps:

[0090] 1. Initialize the parameters θ1 of the input layer and embedding layer, and the parameters θ2 of the decision layer and output layer;

[0091] 2. For each type of input task or action, divide it into a support set and a query set from the basic sample set (i.e., the relevant indicator information of the actions taken under the current situation and action conditions);

[0092] 3. Extract the input features obtained from the support set and perform concatenation operation. Use the generative adversarial network to further calculate the output, obtain the training loss and the corresponding gradient, and perform local updates for the parameters θ2 of the decision layer and the output layer (this process simulates the inner loop process of the model, learning decisions for each type of target and action).

[0093] 4. When all local updates are completed, test the joint loss of the multiple models on the query set, calculate the loss gradient of all actions or tasks, and perform global updates for the parameters θ1 of the input layer and embedding layer, and the parameters θ2 of the decision layer and output layer.

[0094] To ensure robustness and stability during model learning, local updates assume that the interaction data remains unchanged—that is, the situation and actions do not change. Only the parameters of the decision and output layers are updated, while the parameters of the input and embedding layers are not. Thus, the learned model is only sensitive to the correlation between decisions and input actions / situational information, improving the robustness of the learned model across different types of tasks.

[0095] Let P be the amount of basic data related to the task, then for task i of this type, the embedding vector A i It can be represented as

[0096] A i =[e i1 c i1 ;…;e ip c ip ] T

[0097] Among them, c ip Represents task-related d p One-dimensional heat vector, e ip d represents the content corresponding to the task. e -d p Embedding matrix, d e and d p These are the dimensions of the embedding and the number of content categories related to the task.

[0098] The representation of environment and situation is similar to that of the task, using embedding vectors S. j express

[0099] S j =≥[e j1 c j1 ;…;e jq c jq ]

[0100] Among them, c jq Represents task-related d q One-dimensional heat vector, e jq d represents the content corresponding to the situation. e -d q Embedding matrix, d e and d q These represent the embedded dimension and the number of content categories related to the situation, respectively. When the input data samples are continuous, the input nodes are directly connected to each other, eliminating the need for a splicing layer.

[0101] Since the task and situation-related dimensions of the embedding layer are not necessarily the same, matrix factorization layers, which have the same requirements for the embedding dimensions, cannot be used. Therefore, generative adversarial networks (GANs) are adopted as the decision network. The output layer follows the decision layer and aims to output the corresponding decision instructions. The decision layer and output layer can be represented as follows:

[0102]

[0103] Among them, W n and b n W is the weight matrix and bias vector of the N-dimensional decision layer. o and b o These are the weight matrix and bias vector of the output layer. Based on the current situation and environmental information S j The output decision instruction, 'a' and 'σ', represent the activation functions of the decision layer and output layer, respectively. Linear activation units (ReLU) are used as the activation function 'a'.

[0104] The loss function can be expressed as

[0105]

[0106] Among them, H i y represents the historical training data. ij These are relevant indicators for real-world decision-making. During the local update process, the calculated gradient can be expressed as...

[0107]

[0108] Then the parameters of the decision layer and the output layer are updated.

[0109]

[0110] After completing multiple local updates for each task, a global parameter update is performed, which can be represented as follows:

[0111]

[0112]

[0113] In this example embodiment, during the learning and training process, meta-knowledge is extracted from the strategy generation for multiple tasks. This knowledge is then used to perform rapid transfer learning and adaptation on combat tasks with few samples, enabling the model to have strong generalization and robustness, thus better addressing the cold start problem with few or zero samples.

[0114] In this example embodiment, in this invention, for each type of combat mission, a generative adversarial network model based on game theory is used to generate combat decisions, wherein the generator and the discriminator can represent the opposing sides respectively.

[0115] The specific game-playing and decision-making process is as follows:

[0116] 1. Input situational information, target, or mission information into the generator;

[0117] 2. Based on the initial parameters, use a generator model to generate decision instruction information;

[0118] 3. Input the decision instructions generated by the generator and the real label instructions into the discriminator. The discriminator will identify which of the decision instructions generated by the generator and the decision instructions used as labels are real, and output the discrimination result.

[0119] 4. Based on the discrimination results, calculate the loss function and optimize it using gradient information;

[0120] 5. Repeat steps 1 to 4, iterating and optimizing until Nash equilibrium is reached.

[0121] The objective function for a generative adversarial network model is shown below:

[0122]

[0123] In this context, E(*) represents the expected value of the distribution function, Pr represents the distribution of the real samples, Px represents the low-dimensional noise distribution, G(x) represents the samples generated by the generator, and D represents the discriminator. In generative adversarial networks (GANs), the generator and discriminator continuously optimize themselves during a game-like process to improve their generation and discrimination abilities, ultimately achieving higher scores. The objective function is reflected in the following way: the discriminator D aims to have a higher probability of classifying D(y) from the real distribution and a lower probability of classifying D(G(x)) from the generated distribution. For D, naturally, a larger objective function value is desirable. The generator G aims for the final generated distribution G(x) to be closer to the original input decision instruction, so that the discriminator classifies the generated distribution G(x) as the real decision, i.e., a larger D(G(x)), which in turn makes the objective function value smaller. The two interact in a game-optimization-game-like dynamic.

[0124] In this example embodiment, a meta-learning-based learning and training method is used to learn and train the computational generation of decisions for various types of tasks, thereby effectively improving the generalization ability and robustness of the learning model for tasks with few samples and multiple types. Based on generative adversarial networks, decision-making for a single task is learned and trained, utilizing the idea of ​​game theory to obtain an output that infinitely approximates the real decision instructions, improving the efficiency and quality of decision-making. The intelligent decision computation and generation system based on meta-learning and game theory includes a situational information input module, a database module, an intelligent decision generation module, and a visualization analysis module. Through these modules, historical data information and the advantages of "human-in-the-loop" can be effectively utilized to improve the quality and efficiency of decision-making for multiple types of tasks.

[0125] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0126] Furthermore, this example embodiment also provides a multi-task intelligent decision-making computing system based on meta-learning and game theory. (Refer to...) Figure 4 As shown, the multi-task intelligent decision-making computing system 400 based on meta-learning and game theory may include: a situation information input module 410, a database module 420, an intelligent decision generation module 430, and a human-computer interaction module 440. Wherein:

[0127] The situation information input module 410 is used to import or set the situation information of the two sides in the game, including seats, seat relationships, facilities, weapons, troop strength (i.e. troop deployment), targets, and tasks.

[0128] Database module 420 is used for storing situational information and searching and matching information, including knowledge extraction, knowledge refinement, and knowledge representation, to provide data support for intelligent decision-making based on game-based confrontation.

[0129] The intelligent decision generation module 430 is used to calculate and generate decision schemes based on the information provided by the situation input module and the database module, and to recommend and rank the decision schemes according to their probability of winning.

[0130] The human-computer interaction module 440 includes a visualization analysis module, a voice and text interaction module, and a brain signal processing module. The human-computer interaction module is used to provide personnel with information visualization functions and assists personnel in decision-making through human-computer interaction modes such as voice interaction, text interaction, gesture interaction, human brain signal processing, and visual analysis.

[0131] The specific details of each of the multi-task intelligent decision computing system modules based on meta-learning and game theory mentioned above have been described in detail in the corresponding multi-task intelligent decision computing method based on meta-learning and game theory, so they will not be repeated here.

[0132] It should be noted that although several modules or units of a multi-task intelligent decision computing system 400 based on meta-learning and game theory have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0133] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0134] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented as entirely hardware embodiments, entirely software embodiments (including firmware, microcode, etc.), or embodiments combining hardware and software aspects, collectively referred to herein as “circuit,” “module,” or “system.”

[0135] The following reference Figure 5 To describe an electronic device 500 according to such an embodiment of the present invention. Figure 5 The electronic device 500 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0136] like Figure 5 As shown, the electronic device 500 is manifested in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including storage unit 520 and processing unit 510), and a display unit 540.

[0137] The storage unit stores program code that can be executed by the processing unit 510, causing the processing unit 510 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 510 can perform actions such as... Figure 1 Steps S110 to S130 are shown in the diagram.

[0138] Storage unit 520 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 5201 and / or cache memory 5202, and may further include a read-only memory (ROM) 5203.

[0139] Storage unit 520 may also include a program / utility 5204 having a set (at least one) program module 5203, such program module 5205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0140] Bus 550 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0141] Electronic device 500 can also communicate with one or more external devices 570 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 500, and / or with any device that enables electronic device 500 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 550. Furthermore, electronic device 500 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 560. As shown, network adapter 560 communicates with other modules of electronic device 500 via bus 550. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0142] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0143] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the invention described in the "Exemplary Methods" section above.

[0144] refer to Figure 6 As shown, a program product 600 for implementing the above-described method according to an embodiment of the present invention is described. It may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0145] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0146] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0147] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0148] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0149] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0150] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0151] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A multi-task intelligent decision-making computation method based on meta-learning and game theory, characterized in that, The method includes: Step S110: Input preset basic operational information into the operational decision-making system and initialize the operational decision-making system. The preset basic operational information includes a support set and a query set. Step S120: Classify the preset basic combat information according to the combat mission, and use the support set of the preset basic combat information corresponding to the classification of the combat mission as input to perform local update training on the output layer and decision layer of the combat decision system, and update the parameters of the output layer and decision layer of the combat decision system. The method further includes performing local update training on the output layer and decision layer of the combat decision system: Extract the input features from the support set and concatenate the input features; The input features are calculated based on a pre-defined generative adversarial network to obtain the training loss and gradient; The parameters of the output layer and decision layer of the combat decision system are updated based on the training loss and gradient. The method also includes performing local update training on the output layer and decision layer of the combat decision-making system based on game theory: Step S121: Input the situational information, target or mission information of the preset basic combat information into the generator of the combat decision system; Step S122: Based on preset initial parameters, use the generator model to generate decision instruction information; Step S123: Input the decision instruction information generated by the generator and the actual label instruction into the discriminator. The discriminator will then distinguish between the decision instruction information generated by the generator and the decision instruction information used as labels, and output the discrimination result. Step S124: Calculate the loss function based on the discrimination result and optimize it based on the gradient information; Step S125: Repeat steps S121 to S124 for iterative calculation until the Nash equilibrium state is determined based on the discrimination result; Step S130: Based on the output layer and decision layer of the combat decision system that have completed parameter updates after local update training, the combat decision system is globally trained using the preset query set of basic combat information as input, and all parameters of the combat decision system are updated to complete the training of the combat decision system.

2. The method as described in claim 1, characterized in that, The preset basic operational information in the method includes seat information, situational information, facilities, weapons, troop strength and deployment, operational objectives, and operational tasks.

3. The method as described in claim 1, characterized in that, The method also includes: The combat decision system is initialized and configured, and scoring conditions and scoring rules are set for the combat decision system.

4. The method as described in claim 1, characterized in that, The method further includes performing global training on the combat decision-making system: Based on the pre-defined basic combat information of the full classification, and using the pre-defined historical game data information as the input of the combat decision system, the combat decision system is subjected to multi-task meta-enhanced global training. The global training is accelerated by manually adjusting the meta-reinforcement learning process online.

5. The method as described in claim 1, characterized in that, The method further includes performing global training on the combat decision-making system: Based on the pre-defined basic combat information of the full classification, meta-knowledge is extracted for the generation of multi-task strategies, and a query set sample including meta-knowledge is generated. The combat decision system is trained globally using the preset query set of basic combat information as input.

6. A multi-task intelligent decision-making computing system based on meta-learning and game theory, characterized in that, The system, employing the method described in any one of claims 1-5, comprises a situational information input module, a database module, an intelligent decision generation module, and a human-computer interaction module, wherein: The situation information input module is used to import or set the situation information of the two sides in the game, including seats, seat relationships, facilities, weapons, troop strength (i.e. troop deployment), objectives, and tasks. The database module is used for storing situational information and searching and matching information, including knowledge extraction, knowledge refinement, and knowledge representation, to provide data support for intelligent decision-making based on game-based adversarial competition. The intelligent decision generation module is used to calculate and generate decision schemes based on the information provided by the situation input module and the database module, and to recommend and rank the decision schemes according to their probability of winning. The human-computer interaction module includes a visualization analysis module, a voice and text interaction module, and a brain signal processing module. The human-computer interaction module is used to provide personnel with information visualization functions and assists personnel in decision-making through human-computer interaction modes such as voice interaction, text interaction, gesture interaction, human brain signal processing, and visual analysis.

7. An electronic device, characterized in that, include Processor; and A memory storing computer-readable instructions that, when executed by the processor, implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent decision-making method for military confrontation games under incomplete information conditions

    CN112329348A

  • Battlefield game strategy reinforcement learning training method based on cluster influence degree

    CN113705828A