Simulation action method and application based on HTN and action feedback
By adopting HTN and action feedback methods in simulation actions, decomposing the composite task into multiple atomic tasks and selecting the optimal meta action in real time, the problems of low execution efficiency and insufficient dynamic adaptability in the prior art are solved, and more efficient and flexible simulation actions are achieved.
Patent Information
- Application Number
- CN202510197336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art fails to effectively consider the consumption and efficiency of performing meta actions in simulation actions, resulting in low task execution efficiency and long time, weak dynamic adaptability, and inability to effectively deal with resource conflicts.
Using a simulation action method based on HTN and action feedback, the composite task is decomposed into multiple atomic tasks, each atomic task has multiple executable meta actions, and the feedback value of the meta action is obtained in real time based on the battlefield simulation data, and the meta action with the optimal feedback value is selected for execution.
It improves the efficiency and dynamic adaptability of simulation actions, can select the optimal solution when there is a possibility conflict between multitasking or methods, saves computing resources, reduces consumption, and improves simulation flexibility and adaptability.
Smart Images

Figure CN120123083A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of simulation operations, and relates to a simulation operation method and application based on HTN and action feedback. Background Art
[0002] With the in-depth research of simulation operations, when business personnel analyze the actions of simulation entities, they widely use hierarchical task networks to sequentially decompose and refine simulation actions layer by layer. The main method is to sequentially decompose a complete action into multiple sub-actions layer by layer, and select one of the actions to execute according to the current state and set conditions, that is, to summarize the complete action as a composite task, and then recursively decompose the composite task into multiple atomic tasks, and select one of the meta-actions to execute through the calculation results of the conditions in the method.
[0003] In the prior art, subsequent action executions are all based on preconditions, and factors such as the consumption and efficiency of executing meta-actions are not considered. For example, without considering the efficiency of executing meta-actions, there often appears a phenomenon that although the task is completely executed, the efficiency is not high, resulting in too long a time to complete the task.
[0004] Summarizing the exposed disadvantages, they are mainly attributed to the following points: First, the dynamic adaptability of the prior art is weak. When dealing with a dynamically changing environment or task, HTN requires additional mechanisms to enhance its dynamic adaptation ability. Second, it is impossible to control the conflict of computing resources: for large-scale and complex task networks, if there are no optional meta-actions to execute, once resource conflicts occur, the HTN decomposition and execution process may consume a large amount of computing resources.
[0005] Therefore, in view of the above technical problems, it is necessary to provide a method for simulation operations.
[0006] The information disclosed in this background art section is only intended to increase the understanding of the overall background of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention
[0007] The purpose of the present invention is to provide a simulation operation method and application based on HTN and action feedback, which can increase real-time feedback, dynamically adjust task planning according to the actual situation during the task execution process, and ensure that the optimal solution can be selected when there are possible conflicts between multiple tasks or methods.
[0008] In order to achieve the above purpose, the technical solution provided by a specific embodiment of the present invention is as follows:
[0009] In the first aspect, the present invention provides a simulation operation method based on HTN and action feedback, which includes:
[0010] Decompose the composite task for performing the simulation action into multiple atomic tasks, each atomic task having multiple executable meta-actions, and each meta-action can complete the corresponding atomic task;
[0011] Traverse the decision record entries. If the atomic task is included in the decision record entry, execute the meta-action corresponding to the atomic task in the decision record entry;
[0012] If the atomic task is not included in the decision record entry, based on the battlefield simulation data, obtain the feedback value of the meta-action corresponding to the atomic task in the current state. Based on the state requirements of the current simulation action, compare the feedback values of the meta-actions corresponding to the atomic task, and select the meta-action with the optimal feedback value for execution. The feedback value is used to reflect the impact of the corresponding meta-action on the state requirements of the simulation action.
[0013] In one or more embodiments of the present invention, the obtaining the feedback value of the meta-action corresponding to the atomic task in the current state based on the battlefield simulation data includes:
[0014] Obtain the influencing factors corresponding to the state requirements. The influencing factors include the first influencing factor obtained based on the battlefield simulation data and the second influencing factors corresponding to different meta-actions;
[0015] Based on the mapping relationship between the influencing factors and the feedback value, obtain the feedback value of the combined action of the second influencing factor corresponding to the meta-action and the first influencing factor.
[0016] In one or more embodiments of the present invention, decomposing the composite task for performing the simulation action into multiple atomic tasks includes:
[0017] Index the task decomposition strategy corresponding to the composite task for performing the simulation action;
[0018] Based on the battlefield simulation data, execute the judgment logic in the task decomposition strategy for determining whether the composite task can be decomposed in the current state;
[0019] Based on the step-by-step execution of the judgment logic, decompose the composite task into multiple atomic tasks.
[0020] In one or more embodiments of the present invention, the method further includes:
[0021] After executing the meta-action to complete the corresponding atomic task, write the executed meta-action, atomic task information, and the time of execution of the meta-action into the decision record entry.
[0022] In one or more embodiments of the present invention, the method further includes:
[0023] Obtain the difference between the current moment and the execution time of each primitive action in the decision record entry;
[0024] If the difference is greater than the preset elimination threshold, delete the execution time of the primitive action, the corresponding primitive action executed, and the atomic task information.
[0025] In a second aspect, the present invention provides a simulation action system based on HTN and action feedback, which is applied to the simulation action method based on HTN and action feedback, and includes:
[0026] A decomposition module for decomposing a composite task for executing a simulation action into multiple atomic tasks, each atomic task having multiple executable primitive actions, and each primitive action can complete its corresponding atomic task;
[0027] An acquisition module for obtaining, based on the battlefield simulation data, the feedback value of the primitive action corresponding to the atomic task in the current state when executing the atomic task, and the feedback value is used to reflect the influence of the corresponding primitive action on the state requirements of the simulation action;
[0028] A selection module for comparing the feedback values of the primitive actions corresponding to the atomic task based on the state requirements of the current simulation action, and selecting the primitive action with the optimal feedback value for execution.
[0029] In one or more embodiments of the present invention, obtaining the feedback value of the primitive action corresponding to the atomic task in the current state based on the battlefield simulation data includes:
[0030] Obtain the influencing factors corresponding to the state requirements, and the influencing factors include a first influencing factor obtained based on the battlefield simulation data and a second influencing factor corresponding to different primitive actions;
[0031] Based on the mapping relationship between the influencing factors and the feedback value, obtain the feedback value jointly affected by the second influencing factor corresponding to the primitive action and the first influencing factor.
[0032] In a third aspect, the present invention provides a simulation action architecture for implementing the simulation action system based on HTN and action feedback, and includes:
[0033] A resource layer for providing the battlefield simulation resources and action decision resources required for simulation action planning, where the battlefield simulation resources include terrain model resources, weather environment resources, equipment data / model resources, and the action decision resources include rule data resources and primitive action resources;
[0034] A service layer for providing the core services required for simulation action to respond to the decision request of the simulation action, call the battlefield simulation resources and action decision resources of the resource layer, and perform the escape and adaptation of communication data;
[0035] The application layer is provided with a plurality of functional modules for performing business logic processing on the data resources called based on the simulation action decision request, and the functional modules include a project file management module, a graphical editing module, and a resource library management module.
[0036] In a fourth aspect, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the simulation action method based on HTN and action feedback by executing the computer instructions.
[0037] In a fifth aspect, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the simulation action method based on HTN and action feedback.
[0038] Compared with the prior art, the simulation action method based on HTN and action feedback provided by the present invention reconstructs the meta-actions corresponding to the atomic tasks, increases the selection surface of each entity operation during simulation runtime, and overcomes the problem that the prior art cannot make multi-dimensional judgments on battlefield changes through a single execution action plan. At the same time, non-single action execution can save computing resources and reduce consumption when facing large-scale and complex task networks.
[0039] On the other hand, considering that the same meta-action may produce different effects in different simulation environments, and that the meta-action itself has factors such as consumption and efficiency, the present invention selects meta-actions by feeding back the simulation data of the battlefield in real time. This improves the flexibility and dynamic adaptability of the simulation, and can dynamically adjust the task planning according to the actual situation during the task execution, making full use of existing resources to make the best choice.
[0040] Furthermore, by accessing the decision record table, the meta-action information adopted by the recent atomic task can be obtained; the same atomic task can be used over a period of time, which improves the efficiency of the solution implementation. The data in the decision record table will be eliminated regularly, ensuring the timeliness of the meta-action selection strategy, and the elimination time can be freely configured, which improves the flexibility and adaptability of the solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 It is a schematic diagram of a simulation action scenario based on HTN and action feedback in an embodiment of the present invention;
[0043] Figure 2 It is a schematic flow diagram of a simulation action method based on HTN and action feedback in an embodiment of the present invention;
[0044] Figure 3 It is a structural block diagram of a simulation action system based on HTN and action feedback in an embodiment of the present invention;
[0045] Figure 4 It is a structural block diagram of an electronic device in an embodiment of the present invention;
[0046] Figure 5 It is a schematic diagram of a composite task decomposition strategy in a specific embodiment of the present invention;
[0047] Figure 6 It is the correspondence between two-dimensional influence factors and feedback values in a specific embodiment of the present invention. Detailed implementation manners
[0048] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0049] Unless otherwise clearly stated, in the whole specification and claims, the term "comprising" or its variations such as "including" or "having" etc. will be understood to include the stated elements or components, without excluding other elements or other components.
[0050] In a technical solution, there is a method for action execution based on a hierarchical task network, which specifically includes: decomposing a composite task into multiple atomic tasks, and each atomic task is set with a corresponding meta-action. When a target task needs to be executed, if the target task is used as a root node in a preset hierarchical task network, the target task is decomposed based on the above hierarchical neural network, and then the corresponding meta-action is directly executed.
[0051] However, due to the relatively special environment of battlefield simulation, the environmental factors and the combination methods of the primitive actions themselves will have different impacts on the execution of tasks. That is, in different battlefield environments, the same primitive action may have completely different effects. Therefore, when this technical solution is applied to the combat simulation scenario, there will inevitably be problems such as poor dynamic adaptability and the inability to accurately adjust the task execution strategy according to the actual usage environment. Further, for the large-scale and complex task network of battlefield simulation, if there are no alternative primitive actions to execute or the target task does not correspond to any root node in the hierarchical task network, once a resource conflict occurs, the HTN decomposition and execution process may consume a large amount of computing resources.
[0052] The inventors of the present invention found the main drawbacks of the prior art and proposed a new technical implementation idea based on the drawbacks of the prior art: First, reconstruct the primitive actions. Based on the battlefield simulation scenario, diversify the primitive actions corresponding to each atomic task, so that the same atomic task corresponds to multiple atomic tasks, providing more choices for the execution of atomic tasks in the battlefield environment. On the other hand, in order to avoid the situation that the target task does not exist in the task decomposition library in the form of a root node, resulting in the inability to decompose tasks or excessive consumption of computing resources. When indexing the corresponding task decomposition tree in the present invention, the branch task nodes and atomic task nodes are also included in the scope of indexing. When the task corresponding to the branch node is the target task, the task decomposition strategy in the out-edge direction of the branch node is intercepted. This enhances the overall adaptability and flexibility of the solution and improves the hit rate of indexing. At the same time, in order to cooperate with the setting of multiple primitive actions, the present invention will collect battlefield simulation data in real time when executing the corresponding atomic task, and give action feedback values respectively in combination with different primitive actions. Based on the above feedback values, the optimal primitive action can be intuitively given, which cooperates and supports each other with the above-mentioned multiple technical features to achieve the purpose of improving the adaptability of the solution.
[0053] Please refer to Figure 1 shown in the following figure, which is a schematic diagram of the simulation action architecture in an embodiment of the present invention. The simulation action architecture is divided into a resource layer 101, a service layer 102, and an application layer 103 from bottom to top.
[0054] Among them, the resource layer 101 covers various resources required in the simulation, including terrain, actions, rules, etc., integrating the existing and practical resources into the system to minimize the limitation of domain resources during HTN action planning. Specifically, the resource layer 101 is used to provide battlefield simulation resources and action decision-making resources required for simulation action planning. The battlefield simulation resources include but are not limited to terrain model resources, weather environment resources, equipment data resources, equipment model resources, etc.; the action decision-making resources include but are not limited to rule data resources and primitive action resources;
[0055] The service layer 102 is used to provide the core services required for simulation operations. In response to the decision request of the simulation operation, it calls the battlefield simulation resources and action decision resources of the resource layer, performs the escape and adaptation of communication data. In an exemplary embodiment, the services that the service layer 102 can provide include, but are not limited to, integrated management services, data acquisition services, communication interconnection services, HTN domain management services, parsing and verification services, and operation control services.
[0056] Multiple functional modules and / or service instances are provided in the application layer 103, which are used to perform business logic processing on the data resources called based on the simulation operation decision request. The functional modules include an engineering file management module, a graphical editing module, and a resource library management module.
[0057] Specifically, the engineering file management module mainly includes functions such as the creation, editing, and deletion of meta-action or other behavior model engineering files, engineering file structure management, and software help; the graphical editing management module mainly provides functions such as meta-action or other behavior model editing domain management, node management, behavior logic design, node attribute setting, and domain parameter management, to implement the graphical editing function of the behavior models of equipment and personnel; the action fusion parsing module can parse the graphical design results according to the preset fusion rules, can judge and sense the action conditions, and can plan a series of behavior actions of the entity according to the action rules defined by the user; the behavior model export module mainly provides functions such as exporting the behavior model editing domain diagram and encapsulating and exporting the behavior model engineering file, and can generate engineering files recognizable by general coding tools; the resource library management module mainly provides the management functions of the meta-action resource library, operation condition resource library, and sensor resource library required for building the behavior model, and can realize the expansion of the resource library.
[0058] It should be noted that communication connections are provided between the resource layer 101, the service layer 102, and the application layer 103. The communication network extended by the above communication connections can include various connection types, including but not limited to: wired connection, wireless connection, or fiber optic cable connection, etc. At the same time, this communication network can be a local area network, a metropolitan area network, a wide area network, or any combination of the three.
[0059] It should also be noted that the simulation operation architecture of the present invention further includes a user terminal. The user can perform operations that need to be manually set based on the user terminal interface. For example, configuring the decomposition strategy of the composite task, reconstructing the meta-action corresponding to the atomic task, etc. At the same time, the user terminal is installed with a computer software program that matches the adapter development method provided by this method; the user terminal can include, but is not limited to, portable electronic devices such as desktop computers (PC terminals), desktop computers, smart phones, handheld computers, tablet computers, personal digital assistants (PDAs), etc., or wearable electronic devices. The embodiments of the present invention do not limit the above content.
[0060] It should also be noted that the adapter development method in the embodiments of the present invention can be applied to the adapter development system in the embodiments of the present invention. This adapter development system can be configured in a terminal. The terminal may include, but is not limited to, a PC (Personal Computer), a PDA (tablet computer), a smart phone, a smart wearable device, and the like.
[0061] Please refer to Figure 2 As shown, it is a schematic flowchart of a simulation action method based on HTN and action feedback in an embodiment of the present invention. This simulation action method based on HTN and action feedback specifically includes the following steps:
[0062] S201: Decompose the composite task for performing the simulation action into multiple atomic tasks. Each atomic task has multiple executable primitive actions, and each primitive action can complete its corresponding atomic task;
[0063] It should be noted that HTN (Hierarchical Task Network), that is, a hierarchical task network, is a hierarchical structure method for decomposing complex tasks into simpler and more manageable subtasks. These subtasks can be atomic tasks or composite tasks. Among them, an atomic task refers to a task that cannot be further decomposed, and a composite task refers to a task that can be further decomposed. A primitive action, also known as an atomic action, is the basic building block for constructing a complex task. In a hierarchical task network, an atomic action corresponds to an action in the STRIPS (Stanford Research Institute Problem Solver) planning system and is the basic element for constructing a plan.
[0064] In an exemplary embodiment, decomposing the composite task for performing the simulation action into multiple atomic tasks includes: indexing the task decomposition strategy corresponding to the composite task for performing the simulation action; based on the battlefield simulation data, executing the judgment logic in the task decomposition strategy for determining whether the composite task can be decomposed in the current state; based on gradually executing the judgment logic, decomposing the composite task into multiple atomic tasks.
[0065] It should be noted that the above task decomposition strategy refers to a strategy for evaluating based on the current battlefield simulation data and then decomposing the composite task into multiple atomic tasks. The evaluation process depends on the preset judgment logic in the task decomposition strategy. The judgment logic determines how the current composite task should be decomposed based on the current battlefield situation. Optionally, the above task decomposition strategy decomposes the composite task in a layer-by-layer manner, that is, after executing the judgment logic of the nth layer, if the decomposed result is still a composite task, then continue to execute the judgment logic of the n + 1th layer until the composite task is decomposed into atomic tasks.
[0066] For example, as Figure 5 shown in the schematic diagram of the composite task decomposition strategy in an embodiment of the present invention. The corresponding composite task to be decomposed is the "personnel maneuver task". After collecting the current battlefield simulation data, the first-level judgment logic is "whether the personnel encounter the enemy". If the judgment result is no, the "personnel maneuver task" is decomposed into the atomic task P; on the other hand, if the first-level judgment logic determines yes, the "personnel maneuver task" is decomposed into another composite task, the "encountering the enemy and firing task". Since it has not been completely decomposed into atomic tasks at this time, the second-level judgment logic "whether the enemy army is aware" is executed to further decompose the "encountering the enemy and firing task". If the result of the second-level judgment logic is yes, the composite task is decomposed into the atomic task P1; if the result of the second-level judgment logic is no, the composite task is decomposed into the atomic task P2.
[0067] It should also be noted that before the task starts, multiple corresponding meta-actions are configured for each atomic task based on the user terminal. Each meta-action can complete the above atomic task, but different meta-actions have different impacts on the expected requirements in combination with the current simulation combat scenario.
[0068] As Figure 5 shown, the meta-actions corresponding to the atomic task P are normal maneuver, low-profile maneuver, and crawling maneuver; the meta-actions corresponding to the atomic task P1 are passing through the bunker and passing through quickly; the meta-actions corresponding to the atomic task P2 are attacking the enemy army and retreating. Taking the atomic task P as an example, it can be understood that in the case where the personnel do not encounter the enemy, normal maneuver, low-profile maneuver, or crawling maneuver can all achieve the goal of the atomic task, that is, complete the "personnel maneuver".
[0069] S202: Traverse the decision record table entries. If the atomic task is included in the decision record table entries, execute the meta-action corresponding to the atomic task in the decision record table entries;
[0070] It should be noted that the present invention can select and execute the optimal meta-action to complete the current atomic task by calculating the feedback value; however, for tasks that are repeated N times within a short period, N identical task decomposition and feedback value calculation processes are required, and a large amount of redundant calculation processes will increase the load pressure of the system and waste computing resources.
[0071] It can be understood that within a short period, the simulation combat environment changes little. Therefore, for the execution of the same atomic task, the meta-action selected based on the expected state requirements usually does not change within a short period. Based on this, in the embodiment of the present invention, after each execution of the meta-action to complete the corresponding atomic task, the executed meta-action, atomic task information, and the time when the meta-action is executed can be written into the preset decision record table entries.
[0072] In order to maintain the timeliness of the data in the table entry, the present invention introduces an elimination mechanism for deleting the entry information with a relatively long existence time in the table entry. This is to avoid the situation where, as time goes by and the environment changes, the optimal meta-action selected based on the desired state requirements also changes. In one implementation, the time difference between the current moment and the execution time of each meta-action on the decision record table entry can be obtained by using the execution time of each meta-action recorded in the table entry; if the difference is greater than the preset elimination threshold, the execution time of the meta-action, the meta-action corresponding to the execution time, and the atomic task information are deleted.
[0073] Based on different simulation scenarios, the setting of the elimination threshold can be adjusted adaptively. For example, in a short-term weather environment such as a shower, since the environment may change at any time, the size of the elimination threshold should be reduced. In the embodiments of the present invention, there is no limitation on the setting of the size of the elimination threshold.
[0074] S203: If the decision record table entry does not contain the atomic task, based on the battlefield simulation data, obtain the feedback value of the meta-action corresponding to the atomic task in the current state. Based on the state requirements of the current simulation action, compare the feedback values of the meta-actions corresponding to the atomic task, and select the meta-action with the optimal feedback value for execution. The feedback value is used to reflect the impact of the corresponding meta-action on the state requirements of the simulation action.
[0075] It can be understood that when the atomic task that needs to be executed currently is not included in the decision record table entry, it means that the simulation action has not executed the current atomic task within the preset time period. Therefore, in order to ensure that the meta-action for completing the current atomic task is optimal, the corresponding meta-action should be reselected for execution.
[0076] It should be noted that although the meta-actions corresponding to the same atomic task can all achieve the atomic task. However, with the change of the battlefield simulation environment, the same action will show different characteristics in different scenarios. The degree of influence on the state requirements of the simulation action will change.
[0077] In an exemplary embodiment of the present invention, obtaining the feedback value of the meta-action corresponding to the atomic task in the current state based on the battlefield simulation data includes: obtaining the influencing factors corresponding to the state requirements, where the influencing factors include a first influencing factor obtained based on the battlefield simulation data and a second influencing factor corresponding to different meta-actions; based on the mapping relationship between the influencing factors and the feedback value, obtain the feedback value jointly affected by the second influencing factor corresponding to the meta-action and the first influencing factor. Among them, a perceptron is configured in the hierarchical task network, and the data in the perceptron comes from the real-time feedback in the deduction. The real-time feedback mentioned in the present invention is a key part of realizing the military simulation action deduction by using the hierarchical task network planning.
[0078] It can be understood that different state requirements are restricted by various factors. For simulated combat, battlefield simulation data and upcoming meta-actions will all have varying degrees of impact on the completion of tasks. Among them, battlefield simulation data can include, but is not limited to: terrain model data, enemy and friendly equipment data, weather environment data, etc.
[0079] For example, in Figure 5 the specific embodiment shown, if the personnel do not encounter the enemy, the atomic task P is executed. Normal maneuvering, low-profile maneuvering, or crawling maneuvering can all achieve the goal of the atomic task, that is, the above three meta-actions can all complete "personnel maneuvering". However, with the change of battlefield environment data, the degree of influence of different maneuvering methods on the simulation action state requirements will change. In clear weather, the maneuvering speed of normal maneuvering is significantly faster than that of low-profile maneuvering and crawling maneuvering; under the influence of heavy snow and haze weather, the maneuvering speed of low-profile maneuvering may be higher than that of normal maneuvering and crawling maneuvering, that is, the feedback value will change accordingly.
[0080] It should also be noted that in an exemplary embodiment of the present invention, a corresponding relationship between multi-dimensional influence factors and corresponding feedback values can be preset; after determining the atomic task to be executed, based on the battlefield simulation data, obtain the feedback values corresponding to different meta-actions. At the same time, it can be understood that the feedback value is used to reflect the influence of the corresponding meta-action on the simulation action state requirements, and it can be a numerical value, a vector, or a self-defined degree of magnitude, and the embodiments of the present invention do not limit this.
[0081] For example, in Figure 6 the present invention shows the corresponding relationship between two-dimensional influence factors and feedback values in a specific embodiment. As can be seen from the figure, its feedback is represented by a self-defined degree, and the influence factors of the simulation state "maneuvering speed" are the environment and the maneuvering method. The maneuvering method is the meta-action to be taken, and the environment is obtained from the battlefield simulation data. If the current battlefield simulation data shows that the current environment is heavy snow, the maneuvering speed of normal maneuvering is the slowest, the speed of low-profile maneuvering is the fastest, and the maneuvering speed of crawling maneuvering is medium.
[0082] In another embodiment of the present invention, corresponding algorithms can also be set based on dynamic indicators, damage calculation formulas, influence factors related to each of the atomic tasks, and their weights, etc. Obtain the current battlefield parameters required for the corresponding algorithm based on the battlefield simulation data, and calculate the corresponding feedback value through calculation. The calculation method of the feedback value comes from the quantitative processing of the situations involved in the task scenario by military experts and military personnel according to experience values or military standards.
[0083] For example, flat terrain refers to terrain where the ratio of elevation difference to horizontal distance is less than or equal to 2%; open terrain refers to terrain without obvious obstacles and with a line-of-sight distance greater than 500 meters. The above examples can be used to plan maneuvering behaviors that conform to military logic. Under satisfied conditions, there are three ways for the executor to maneuver forward (primitive actions). During the forward movement, the consumption of the current path (such as time, fuel consumption, etc.) is sensed. The above evaluations and calculations of time and fuel consumption can be derived from the classification of the flatness of the maneuvering route.
[0084] The state requirements for the simulation operation can be dynamically adjusted according to the state of the entire operation during the execution of the entire operation, thereby changing the method of calculating the feedback value.
[0085] It can be understood that under different battlefield environments and rapidly changing battlefield situations, the state requirements for similar simulation operations may be very different. After obtaining the feedback values corresponding to different primitive actions, the primitive action with the optimal feedback value should be executed based on the current state requirements.
[0086] The present invention provides multiple primitive actions that can complete tasks in atomic tasks, and selects an optimal primitive action based on the current situation from them according to the real-time feedback value of executing the primitive action (the value calculated by factors such as the consumption and efficiency of all primitive actions in the current atomic task when executing the atomic task). For example, if high efficiency is required to complete the primitive action in the current state, the primitive action with the shortest execution time can be selected to complete the task; if resources are tight in the current state, the primitive action with the lowest consumption can be selected to complete the task.
[0087] For example, in clear weather, the maneuvering speed of normal maneuvering is significantly faster than that of low-profile maneuvering and crawling maneuvering; while under the influence of heavy snow and foggy weather, the maneuvering speed of low-profile maneuvering may be higher than that of normal maneuvering and crawling maneuvering. Then, at this time, when the personnel have not encountered the enemy, if the state requirement of the current simulation operation is "conduct maneuvering operations at the fastest speed", the low-profile maneuvering will be preferentially selected as the primitive action to be executed.
[0088] From the perspective of state requirements, under the condition that the enemy has detected, attacking the enemy will cause more losses to the enemy than retreating, but our side will also suffer more losses. Therefore, if the state requirement at this time is "strike the enemy", attacking the enemy will be preferentially selected as the primitive action to be executed; if the state requirement at this time is "preserve our strength", retreating will be preferentially selected as the primitive action to be executed.
[0089] Please refer to Figure 3 As shown, based on the same inventive concept as the foregoing simulation operation method based on HTN and action feedback, an embodiment of the present invention provides a simulation operation system 300 based on HTN and action feedback, which includes: a decomposition module 301, a traversal module 302, and a selection module 303.
[0090] Specifically, the decomposition module 301 is used to decompose the composite task for performing the simulation action into multiple atomic tasks. Each atomic task has multiple executable meta-actions, and each meta-action can complete its corresponding atomic task; the traversal module 302 is used to traverse the decision record table entries. When the atomic task is included in the decision record table entry, execute the meta-action corresponding to the atomic task in the decision record table entry; the selection module 303 is used to when the atomic task is not included in the decision record table entry, based on the battlefield simulation data, obtain the feedback value of the meta-action corresponding to the atomic task in the current state, and based on the state requirements of the current simulation action, compare the feedback values of the meta-actions corresponding to the atomic task, and select the meta-action with the optimal feedback value for execution. The feedback value is used to reflect the influence of the corresponding meta-action on the state requirements of the simulation action.
[0091] In the current simulation action system based on HTN and action feedback, in order to avoid the lack of domain resources and enable business personnel to use the existing system to make behavior plans, the resource library has been expanded, and rich action rules, simulation action models, etc. are built-in. Secondly, based on the original HTN action planning, the present invention implements a sensor, which can add a feedback value to the meta-action on the execution sequence through the world state value, and judge which meta-action to execute according to the feedback value, so as to ensure the execution of the optimal action plan in the current state, such as the lowest consumption and the highest efficiency.
[0092] Currently, the action planning based on HTN in the military field only allows each method to be executed in a single line according to conditions. The action model planned by this method can only complete military actions, and theoretical and practical experience values are not added to the model. This results in the program execution body only completing one task, rather than making a "thinking" choice for a better solution. On this basis, the present invention adds a sensor with real-time feedback to the planning of the behavior model, which is equivalent to the execution body being able to make a choice according to the state changes during the execution process in the execution logic divided by the behavior model, making military simulations more realistic in actual combat.
[0093] Please refer to Figure 4 As shown, the embodiment of the present invention also provides an electronic device 400, which includes at least one processor 401, a memory 402 (such as a non-volatile memory), a memory 403, and a communication interface 404, and at least one processor 401, the memory 402, the memory 403, and the communication interface 404 are connected together via a bus 405. At least one processor 401 is used to call at least one program instruction stored or encoded in the memory 402, so that at least one processor 401 executes various operations and functions of the simulation action method based on HTN and action feedback described in each embodiment of this specification.
[0094] In the embodiments of this specification, the electronic device 400 may include but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile electronic devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable electronic devices, consumer electronic devices, and so on.
[0095] An embodiment of the present invention also provides a computer-readable medium, on which computer-executable instructions are carried. When the computer-executable instructions are executed by a processor, they can be used to implement various operations and functions of the simulation action method based on HTN and action feedback described in the embodiments of this specification.
[0096] The computer-readable medium in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.
[0097] In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, on which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0098] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0099] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, systems, and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 means for realizing the functions specified in one block or multiple blocks.
[0100] The foregoing description of the specific exemplary embodiments of the present invention is for the purposes of illustration and exemplification. These descriptions are not intended to limit the present invention to the precise forms disclosed, and it is obvious that many changes and variations can be made in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical applications, so that those skilled in the art can implement and utilize various different exemplary embodiments of the present invention, as well as various different selections and changes. The scope of the present invention is intended to be defined by the claims and their equivalents.
[0101] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0102] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A simulation action method based on HTN and action feedback, characterized in that: include: Decomposing the composite task for executing the simulation action into multiple atomic tasks, each atomic task has multiple executable meta-actions, and each meta-action can complete the atomic task corresponding to it; Traversing the decision record table entry, if the decision record table entry contains the atomic task, executing the meta-action corresponding to the atomic task in the decision record table entry; If the decision record table entry does not include the atomic task, then based on the battlefield simulation data, the feedback value of the meta-action corresponding to the atomic task in the current state is obtained; based on the state requirements of the current simulation action, the feedback values of the meta-actions corresponding to the atomic task are compared, and the meta-action with the best feedback value is selected for execution. The feedback value is used to reflect the impact of the corresponding meta-action on the state requirements of the simulation action.
2. The method for simulating action based on HTN and action feedback according to claim 1, characterized in that: Based on the battlefield simulation data, obtaining the feedback value of the meta-action corresponding to the atomic task in the current state includes: Acquire influencing factors corresponding to the state requirement, wherein the influencing factors include a first influencing factor acquired based on battlefield simulation data and a second influencing factor corresponding to different meta-actions; Based on the mapping relationship between the influencing factors and the feedback values, a feedback value of the second influencing factor corresponding to the meta-action and the first influencing factor acting together is obtained.
3. The method for simulating action based on HTN and action feedback according to claim 1, characterized in that: Decompose the composite task used to perform simulation actions into multiple atomic tasks, including: Indexing the task decomposition strategy corresponding to the composite task for executing the simulation action; Based on the battlefield simulation data, executing the judgment logic in the task decomposition strategy for judging whether the composite task can be decomposed in the current state; The judgment logic is executed step by step to decompose the complex task into multiple atomic tasks.
4. The method for simulating action based on HTN and action feedback according to claim 1, characterized in that: The method further comprises: After executing the meta-action to complete the corresponding atomic task, the executed meta-action, atomic task information and the time when the meta-action was executed are written into the decision record entry.
5. The method for simulating action based on HTN and action feedback according to claim 1, characterized in that: The method further comprises: Get the difference between the current time and the execution time of each element action in the decision record table; If the difference is greater than a preset elimination threshold, the meta-action execution time and the meta-action and atomic task information executed corresponding to the meta-action execution time are deleted.
6. A simulation action system based on HTN and action feedback, applied to the simulation action method based on HTN and action feedback as described in any one of claims 1 to 5, characterized in that: include: A decomposition module, used for decomposing a composite task for executing a simulation action into a plurality of atomic tasks, each atomic task having a plurality of executable meta-actions, each meta-action being capable of completing the atomic task corresponding thereto; A traversal module, used for traversing the decision record table entries, and when the decision record table entry contains the atomic task, executing the meta-action corresponding to the atomic task in the decision record table entry; A selection module is used to obtain the feedback value of the meta-action corresponding to the atomic task in the current state based on the battlefield simulation data when the decision record table entry does not contain the atomic task; based on the state requirements of the current simulation action, compare the feedback values of the meta-actions corresponding to the atomic task, and select the meta-action with the best feedback value for execution. The feedback value is used to reflect the influence of the corresponding meta-action on the state requirements of the simulation action.
7. The simulation action system based on HTN and action feedback according to claim 6, characterized in that: The obtaining, based on the battlefield simulation data, a feedback value of the meta-action corresponding to the atomic task in the current state includes: Acquire influencing factors corresponding to the state requirement, wherein the influencing factors include a first influencing factor acquired based on battlefield simulation data and a second influencing factor corresponding to different meta-actions; Based on the mapping relationship between the influencing factors and the feedback values, a feedback value of the second influencing factor corresponding to the meta-action and the first influencing factor acting together is obtained.
8. A simulation action architecture, used to implement a simulation action system based on HTN and action feedback as shown in any one of claims 6-7, characterized in that: include: The resource layer is used to provide battlefield simulation resources and action decision resources required for simulation action planning. The battlefield simulation resources include terrain model resources, weather environment resources, and equipment data / model resources. The action decision resources include rule data resources and meta-action resources. The service layer is used to provide the core services required for the simulation action, respond to the decision request of the simulation action, call the battlefield simulation resources and action decision resources of the resource layer, and perform the escape and adaptation of the communication data; The application layer is provided with a plurality of functional modules for performing business logic processing on the data resources called based on the simulation action decision request, and the functional modules include a project file management module, a graphical editing module, and a resource library management module.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the simulation action method based on HTN and action feedback as described in any one of claims 1 to 5 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the simulation action method based on HTN and action feedback according to any one of claims 1 to 5.