An equipment simulation training task scheduling method for command and control cooperation
Patent Information
- Application Number
- CN202611294679.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-29
AI Technical Summary
传统系统通常要求在执行前唯一确定一条路径,而缺乏对多种潜在意图进行并行保留、概率评估和动态收敛的能力,因而无法适应真实指挥过程中先按意图方向推进、再随态势反馈逐步明确的协同模式
1.本申请提供了一种面向指控协同的装备模拟训练任务编排方法,通过同步获取来自指控端的非结构化指令文本、交互行为特征数据以及当前仿真态势数据进行多模态联合解析,映射生成带有节点概率分布的意图概率任务图,并利用影子仿真副本对条件模糊节点的多维参数搜索区间进行离散采样与独立推演,进而通过轨迹比对提取低风险的公共任务节点,形成试探性任务指令集并入主仿真推演流执行;从而使系统能够在指令语义模糊或战术参数缺失的情况下,不依赖阻断式的强制交互确认,而是通过隐式行为特征与实时态势的自适应匹配来动态构建并验证任务演化路径;有效避免了传统模拟训练系统因过度依赖绝对精确的结构化输入而导致的指控流程割裂与推演停滞问题,使得低风险的前置试探动作能够安全地提前介入主推演流,极大提升了非精确对抗环境下意图解析的容错能力和实战化训练的连续沉浸感;
Smart Images

Figure CN122839679A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mission orchestration technology, and in particular to a method for orchestrating equipment simulation training missions for command and control coordination. Background Technology
[0002] In current equipment simulation training systems, task orchestration mechanisms often employ a unidirectional linear mapping logic from instructions to task actions. Upon receiving an instruction, the system attempts to immediately parse it into defined task nodes, execution objects, and action sequences. If the instruction contains semantic ambiguity, incomplete conditional triggers, or unclear object referencing, the system often cannot continue forming an executable task chain and must resort to blocking clarification through pop-up windows, voice prompts, or manual parameter completion. While this approach can improve parameter certainty to some extent, it directly interrupts the simulation process, forcing commanders to detach from the continuously changing tactical situation and resort to tabular, menu-based, or question-and-answer-style data entry, resulting in a disconnect between the command and control instruction flow and the simulation flow. Especially in highly dynamic combat scenarios, situational data such as enemy-friendly distance, target formation, firepower availability, reconnaissance windows, and areas of interest in the field of view are constantly changing. If the system cannot perform real-time parsing and dynamic orchestration of ambiguous instructions while the simulation continues, it will lead to delayed task response, decreased training immersion, and reduced human-machine collaboration efficiency.
[0003] In related technologies, the technical bottleneck brought about by unstructured command and control instructions is not merely the lack of a certain input field, but rather the need for the system to establish a sustainable coupling relationship between uncertain semantics, dynamic situations, and continuous simulation. For the same fuzzy instruction, multiple reasonable task interpretation paths may exist simultaneously. For example, different target objects, different task areas, different triggering times, or different handling intensities may all match the current situation. Traditional systems typically require a unique path to be determined before execution, lacking the ability to retain, probabilistically evaluate, and dynamically converge multiple potential intentions in parallel. Therefore, they cannot adapt to the collaborative mode of advancing according to the intention direction first and then gradually clarifying with situational feedback in real command processes. In other words, existing technologies lack a dynamic task orchestration mechanism that can simultaneously maintain multiple candidate task branches in the background and continuously adjust the credibility of each branch based on situational changes, entity selection status, field of view focus positions, and subsequent interactive feedback. This reduces the efficiency of equipment simulation training task orchestration and requires improvement. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this application provides a method for arranging equipment simulation training tasks for command and control coordination.
[0005] This application provides a method for arranging equipment simulation training tasks for command and control coordination, comprising the following steps: Simultaneously acquire unstructured command text, interactive behavior feature data, and current simulation situation data from the command and control terminal. The interactive behavior feature data includes the coordinates of the field of view center and the entity selection status. Unstructured instruction text is parsed into action instruction feature words, tactical entity nouns and condition trigger words. Combined with interactive behavior feature data and current simulation situation data, an intent probability task graph containing multiple task nodes and node probabilities is generated. A multi-dimensional parameter search interval is formed for the conditionally ambiguous task nodes in the intent probability task graph. Discrete sampling of the intent probability task map is performed according to the multidimensional parameter search interval to generate multiple candidate task decomposition schemes. Each candidate task decomposition scheme is then injected into the shadow simulation copy formed by copying the current simulation situation data to obtain multiple sets of situation evolution trajectory data. Path comparison is performed on multiple sets of situation evolution trajectory data to extract common task nodes that meet logical consistency and have risk values below the preset risk threshold, forming a set of tentative task instructions, and then the set of tentative task instructions is incorporated into the main simulation training data stream for execution. Collect feedback feature vectors after the execution of the exploratory task instruction set, correct the multidimensional parameter search interval and node probability based on the feedback feature vectors, and calculate the distribution entropy value of the intent probability task graph; When the distribution entropy value is lower than the preset convergence threshold, the task path with the highest node probability is extracted and converted into an equipment-level task operation sequence; when the distribution entropy value is not lower than the preset convergence threshold within a preset number of steps, the candidate task decomposition scheme is regenerated based on the corrected multidimensional parameter search interval until the equipment-level task operation sequence is obtained, and the equipment-level task operation sequence is output to complete the task arrangement.
[0006] In summary, this application includes at least one of the following beneficial technical effects: 1. This application provides a method for orchestrating equipment simulation training tasks for command and control coordination. It performs multimodal joint analysis by synchronously acquiring unstructured command text, interactive behavior feature data, and current simulation situation data from the command and control end. This analysis maps and generates an intent probability task graph with node probability distributions. Furthermore, it uses shadow simulation copies to discretely sample and independently extrapolate the multidimensional parameter search range of conditionally ambiguous nodes. Finally, it extracts low-risk common task nodes through trajectory comparison, forming a tentative task instruction set which is then incorporated into the main simulation extrapolation flow for execution. This allows the system to dynamically construct and verify task evolution paths through adaptive matching of implicit behavioral features and real-time situation, even when command semantics are ambiguous or tactical parameters are missing, without relying on blocking-style forced interactive confirmation. This effectively avoids the command and control process fragmentation and extrapolation stagnation problems caused by the over-reliance on absolutely precise structured input in traditional simulation training systems. It allows low-risk preliminary tentative actions to safely intervene in the main extrapolation flow in advance, greatly improving the fault tolerance of intent analysis in imprecise adversarial environments and the continuous immersion of realistic training. 2. This application utilizes the feedback feature vectors from the collected exploratory task instruction set execution as a closed-loop driving force to correct the boundary range and node probability of the multi-dimensional parameter search interval in real time. It also combines the distribution entropy value of the intent probability task graph as a dynamic quantitative evaluation basis for intent convergence. When the distribution entropy value does not meet the convergence requirement, the adaptive driving system regenerates candidate schemes in the corrected parameter space and iteratively deduces until it accurately locks the task path with the highest probability and automatically transforms it into the underlying equipment-level task operation sequence. This not only helps the system gradually and accurately converge the commander's true tactical intent through silent human-machine implicit game in complex dynamic confrontation situations, but also achieves smooth dimensionality reduction and intelligent arrangement from macroscopic vague command and control commands to microscopic equipment-level precise control commands without interrupting the subjective command rhythm. This significantly improves the intelligent evaluation accuracy, intent capture efficiency, and adaptive evolution level of the command and control collaborative training system to highly dynamic and complex battlefields. Attached Figure Description
[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a flowchart illustrating a method for arranging equipment simulation training tasks for command and control coordination, as described in this application.
[0009] Figure 2 This is a schematic diagram illustrating the implementation logic connection of the intent probability task graph in the embodiments of this application.
[0010] Figure 3 This is a schematic diagram illustrating the implementation logic of the candidate task decomposition scheme in an embodiment of this application.
[0011] Figure 4 This is a schematic diagram illustrating the logical connection of the feedback feature vector implementation in an embodiment of this application. Detailed Implementation
[0012] The following description, in conjunction with the implementation of the present invention, is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention, and all such modifications and additions should fall within the protection scope of the present invention.
[0013] Example
[0014] This implementation method is applied to the command and control coordination task orchestration scenario in equipment simulation training. Its core lies in interpreting the unstructured command text input from the command and control terminal, interactive behavior feature data, and current simulation situation data within the same data processing link. This allows the system to avoid directly interrupting the training process when faced with incomplete or uncertain command and control intentions. Instead, it first deduces multiple possible task paths in a shadow simulation copy, then extracts a set of low-risk and logically consistent exploratory task commands and inputs them into the main simulation training data stream. Subsequently, it uses the feedback feature vectors from the command and control terminal during this exploratory execution process to correct the multi-dimensional parameter search interval and node probabilities, gradually converging to an executable equipment-level task operation sequence. After this processing, the unstructured command text is no longer simply understood as a single command input, but becomes a task data source that can be continuously corrected according to situational evolution and interactive feedback.
[0015] This application discloses a method for arranging equipment simulation training tasks for command and control coordination.
[0016] Reference Figure 1-4 A method for orchestrating equipment simulation training tasks for command and control coordination includes the following steps: Simultaneously acquire unstructured command text, interactive behavior feature data, and current simulation situation data from the command and control terminal. The interactive behavior feature data includes the coordinates of the field of view center and the entity selection status. In a specific embodiment, the above method is executed by a task orchestration processing module deployed in the simulation training system. The task orchestration processing module is connected to the human-computer interaction recording interface, simulation situation data interface, simulation training task library and main simulation training data stream of the command and control terminal. When the command and control terminal generates unstructured instruction text, the task orchestration processing module synchronously acquires the unstructured instruction text, interaction behavior feature data and current simulation situation data at the same timestamp.
[0017] The unstructured command text can be a text string obtained from speech recognition or natural language text input from the command and control terminal. Its data format typically includes the command content, issuance time, and command source identifier. Interactive behavior feature data includes at least the field-of-view center coordinates and entity selection status. The field-of-view center coordinates use two-dimensional coordinates from the current electronic situation view with map projection labels. The entity selection status records the entity identifier of the selected entity and the time of selection. The current simulation situation data comes from the situation snapshot output by the simulation engine at this timestamp, including entity status data, environmental status data, resource status data, enemy-friendly distance changes, firepower availability status, reconnaissance availability status, and target attributes corresponding to each entity identifier. To avoid time discrepancies between different data sources, the task orchestration processing module uses the issuance time of the unstructured command text as a benchmark, reads interactive behavior feature data and current simulation situation data with a time difference not exceeding one simulation step from the cache queue, and encapsulates these three types of data into the same task parsing record for use in subsequent steps.
[0018] Unstructured instruction text is parsed into action instruction feature words, tactical entity nouns and condition trigger words. Combined with interactive behavior feature data and current simulation situation data, an intent probability task graph containing multiple task nodes and node probabilities is generated. A multi-dimensional parameter search interval is formed for the conditionally ambiguous task nodes in the intent probability task graph. Furthermore, the mapping generates an intent probability task graph containing multiple task nodes and node probabilities, including: Unstructured command text is segmented and syntactically ordered to extract action command feature words, tactical entity nouns, and conditional trigger words. Read the task template table from the simulation training task library, match the action instruction feature words with the standard task actions in the task template table, and generate candidate task nodes; By associating tactical entity names with entity identifiers in the current simulation situation data, the execution object data corresponding to the candidate task nodes can be obtained; Based on the order of action instruction feature words in the unstructured instruction text and the positional relationship of the execution object data in the current simulation situation data, the execution order between candidate task nodes is determined and node connection relationships are generated. Mark the candidate task nodes corresponding to the conditional trigger words as conditional fuzzy task nodes, and generate node branches based on the trigger categories corresponding to the conditional trigger words in the task template table; Construct an intent probability task graph based on node connection relationships, conditional fuzzy task nodes, and node branches, and assign initial node probabilities to node branches according to the number of trigger categories corresponding to each node branch.
[0019] In a specific embodiment, the task orchestration processing module parses the unstructured instruction text, categorizing action-related words into action instruction feature words, words corresponding to equipment, targets, areas, enemy units, and friendly units into tactical entity nouns, and words indicating timing, conditions, degree, and triggering relationships into conditional trigger words. The parsing results do not directly form the final task; instead, they are first jointly verified with interactive behavior feature data and current simulation situation data. For example, when action instruction feature words point to standard task actions such as reconnaissance, tracking, suppression, and strike, the module needs to combine the entity distribution near the field of view center coordinates and the entity identifiers corresponding to the entity selection state to determine which target area and execution object the action is more likely to act upon. Based on this, the task orchestration processing module generates an intent probability task graph containing multiple task nodes and node probabilities. Task nodes represent executable task actions, and node probabilities represent the relative likelihood of different node branches being selected. If a conditional trigger word causes a task node to lack a clear triggering timing, target category, task area, or handling intensity, the task node is marked as a conditionally fuzzy task node, and a multi-dimensional parameter search interval is formed around this conditionally fuzzy task node. Once the multidimensional parameter search interval is formed, it is written into the intention probability task graph, so that subsequent discrete sampling, shadow inference and feedback correction are all carried out around the same data structure.
[0020] Specifically, the construction of the intent probability task graph can adopt a relatively stable hierarchical parsing approach. The task orchestration processing module first performs word segmentation and syntactic order annotation on the unstructured instruction text. During word segmentation, the original word order, part-of-speech tags, and intra-sentence positions are preserved. Syntactic order annotation records the sequential relationships between action instruction feature words, tactical entity nouns, and conditional trigger words. For an unstructured instruction text containing multiple actions, the task orchestration processing module does not simply select actions based on word frequency. Instead, it matches the action instruction feature words against the task template table in the simulation training task library item by item. The task template table stores standard task actions, task types, available execution objects, trigger categories, timing constraints, disposal levels, and the weights corresponding to task types. The relationship between action instruction feature words and standard task actions can be achieved using a combination of dictionary mapping and edit distance verification. When the mapping results fall into the same set of standard task actions, candidate task nodes are generated.
[0021] After candidate task nodes are generated, tactical entity names need to be associated with entity identifiers in the current simulation situation data. The task orchestration processing module reads the entity identifier, entity name, owner, location coordinates, and target attributes from the current simulation situation data, performs text matching between the tactical entity names and entity names, and then narrows down the matching range by utilizing the entity distribution near the field of view center coordinates. If the entity selection state already points to a certain entity identifier, and the target attribute of that entity identifier matches the task type of the candidate task node, the task orchestration processing module prioritizes using that entity identifier as the execution object data corresponding to the candidate task node. The execution object data obtained in this way is not only used for constructing the intent probability task graph, but also continues to be used in subsequent multi-dimensional parameter search interval formation, candidate task decomposition scheme generation, and equipment-level task operation sequence conversion, avoiding the parsing results remaining only at the text level.
[0022] The generation of node connections relies on two types of sequence information. One type comes from the order of action command feature words in the unstructured command text, and the other comes from the positional relationship of the execution object data in the current simulation situation data. The task orchestration processing module first establishes initial connections according to the action order in the unstructured command text, and then uses the distance between the execution object data, the region they belong to, and the accessibility of their actions to verify whether the initial connections conform to the simulation logic. If the execution result of the previous task node can provide the subsequent task node with detection result data, position constraints, or resource status, the node connection relationship between the two is retained; if the connection would cause the same simulated equipment entity to execute conflicting tasks in the same time period, the connection relationship is deleted. Candidate task nodes corresponding to conditional trigger words are marked as conditionally fuzzy task nodes. For example, the trigger category can include types such as target appearance, threat escalation, entry into a region, and resource satisfaction. The task orchestration processing module generates node branches according to the trigger category corresponding to the conditional trigger word in the task template table, and constructs an intent probability task graph according to the node connection relationship, conditionally fuzzy task nodes, and node branches. The initial node probability of a node branch can be evenly distributed according to the number of trigger categories, and appropriate corrections can be made when there is an entity selection state, the field of view center coordinates are concentrated and pointing, and the current simulation situation data has strong constraints. However, after correction, the sum of the node probabilities under the same condition fuzzy task node should still be kept to one.
[0023] In this implementation, although the system acquires interactive behavior feature data and current simulation situation data simultaneously at the beginning of the analysis, it still chooses to initially distribute the data equally according to the number of trigger categories. The core purpose is to establish an unbiased logical baseline and prevent premature convergence of intent deduction. Since the core scenario of this method is to handle incomplete and uncertain accusation intents and rely on shadow simulation copies to deduce multiple possibilities, if the system relies too much on the instantaneous field of view center coordinates or entity selection state to absolutely allocate or even truncate probabilities in the initial stage, it is very easy for the system to miss those logically reasonable but potential real intents that lack significant interactive features at the moment the command is issued. Therefore, the system adopts a strategy of first preserving diversity and then adding weights: first, it relies solely on the static "task template table" to generate logically complete node branches and distribute the initial probabilities equally, ensuring that all possible task evolution directions can obtain a legitimate survival space in the multi-dimensional parameter search range; then, it uses the real-time entity selection state, field of view center, and strong situation constraints as appropriate correction factors to tilt the distribution baseline. The above design not only utilizes existing multimodal interaction data to guide high-probability directions, but also preserves the diversity of candidate task paths to the maximum extent in the underlying data structure. This lays the necessary hypothesis verification space for subsequent gradual and accurate convergence of distribution entropy values through shadow inference, trial execution, and continuous feedback feature vectors.
[0024] Furthermore, a multi-dimensional parameter search interval is formed for the conditionally fuzzy task nodes in the intent probability task graph, including: Read the trigger category corresponding to the fuzzy task node and break down the trigger category into task area dimension, target category dimension, trigger timing dimension and handling intensity dimension; Based on the projection position of the field of view center coordinates in the current simulation situation data, the central region of the task region dimension is determined, and the task region search interval is formed by combining the distribution range of entities around the central region. Based on the entity identifier corresponding to the selected entity state, read the target attribute of the entity identifier in the current simulation situation data to form a target category search range; Based on the timing constraints corresponding to the conditional trigger words in the task template table, and combined with the enemy-friendly distance change data in the current simulation situation data, a trigger timing search interval is formed; Based on the response level corresponding to the conditional trigger words, read the firepower availability status and reconnaissance availability status in the current simulation situation data to form a response intensity search range; The task area search range, target category search range, trigger timing search range, and treatment intensity search range are merged into a multi-dimensional parameter search range and written back to the corresponding conditional fuzzy task node.
[0025] Specifically, after forming the intent probability task graph, the conditionally fuzzy task nodes need to be further expanded into multi-dimensional parameter search intervals. The task orchestration processing module reads the trigger category corresponding to the conditionally fuzzy task node and breaks it down into task region dimension, target category dimension, trigger timing dimension, and action intensity dimension. The task region dimension is used to limit the spatial range of the task occurrence, the target category dimension is used to limit the range of target attributes, the trigger timing dimension is used to limit when the conditionally fuzzy task node is executed, and the action intensity dimension is used to limit the strength level of action methods such as reconnaissance, tracking, suppression, and attack. These dimensions are not independent; they all originate from the same conditionally fuzzy task node and will be combined into parameter combination data during subsequent discrete sampling.
[0026] The task area search interval is jointly determined by the field of view center coordinates and the current simulation situation data. The task orchestration processing module projects the field of view center coordinates onto the map coordinate system used by the current simulation situation data to obtain the projected position, and reads the entity distribution range of surrounding entities with the projected position as the center. The entity distribution range can be determined by the bounding rectangle of visible entities, the radius of the task area, and the terrain boundary. If there is an entity identifier matching the tactical entity name within the area corresponding to the field of view center coordinates, the task area search interval will prioritize covering the actionable area around the entity identifier; if there are multiple suspicious targets around the center area, the task area search interval will retain the continuous area where these targets are located, so as to generate different task area candidates during subsequent discrete sampling. The target category search interval is derived from the entity identifier corresponding to the entity selection state. The task orchestration processing module reads the target attributes of the entity identifier in the current simulation situation data, such as target category, maneuver status, threat level, and allied side, and writes the target attribute range that meets the constraints of the task template table into the target category search interval. If the entity selection state is empty, the target category search interval can still be formed based on the association result between the tactical entity name and the current simulation situation data, but its range is usually wider.
[0027] The trigger timing search interval is formed by the timing constraints corresponding to the conditional trigger words in the task template table and the enemy-friendly distance change data in the current simulation situation data. The task orchestration processing module reads the current distance, direction of distance change, and magnitude of distance change within the most recent simulation steps from the enemy-friendly distance change data, and matches them with the timing constraints. That is, the task orchestration processing module first reads the current enemy-friendly distance, direction of distance change, and magnitude of distance change within the most recent simulation steps from the situation data; then, it uses these dynamically changing parameters to perform trend extrapolation, matches and calculates them with the timing constraints in the task template table, and derives the expected time point that satisfies the constraint condition; finally, it expands a certain time span in both directions forward and backward with the expected time point of the constraint condition as the center, thereby dynamically constructing the trigger timing search interval. For example, if the timing constraint requires the target to enter the effective reconnaissance range, the trigger timing search interval expands forward and backward with the simulation step length of the expected entry into the range as the center. The disposal intensity search interval needs to combine the disposal level corresponding to the conditional trigger words and the firepower availability status and reconnaissance availability status in the current simulation situation data. If reconnaissance availability is sufficient but firepower availability is limited, the response intensity search range will be biased towards reconnaissance, tracking, and calibration. If firepower availability meets the resource conditions of the corresponding mission template table, the response intensity search range can cover suppression and strike levels. The mission area search range, target category search range, trigger timing search range, and response intensity search range are merged to form a multi-dimensional parameter search range, which is then written back to the corresponding conditional fuzzy mission node for direct use by discrete sampling.
[0028] Discrete sampling of the intent probability task map is performed according to the multidimensional parameter search interval to generate multiple candidate task decomposition schemes. Each candidate task decomposition scheme is then injected into the shadow simulation copy formed by copying the current simulation situation data to obtain multiple sets of situation evolution trajectory data. Furthermore, the intent probability task graph is discretely sampled according to the multidimensional parameter search interval to generate multiple candidate task decomposition schemes, including: Encode each dimension in the multidimensional parameter search interval with a predetermined number of bits to obtain the discrete value sequence corresponding to each dimension; The discrete value sequences are combined according to the node connection relationship in the intention probability task graph to generate multiple parameter combination data; Parameter combinations that do not meet the requirements of entity availability, task distance limits, and resource occupancy in the current simulation situation data are removed to obtain valid parameter combinations. Fill the conditional fuzzy task node in the intent probability task graph with the effective parameter combination data to generate a candidate task path with definite parameters. Based on the task nodes, execution object data, and node connection relationships in the candidate task path, multiple candidate task decomposition schemes are obtained.
[0029] In one specific embodiment, the task orchestration processing module discretely samples the intent probability task graph according to the multi-dimensional parameter search interval. Each dimension is first converted into an enumerable discrete value sequence, and then combined into parameter combination data according to the node connection relationships in the intent probability task graph. After the parameter combination data is verified by the current simulation situation data, multiple candidate task decomposition schemes are generated. Each candidate task decomposition scheme includes task nodes, execution object data, node connection relationships, and definite parameters for filling in fuzzy task nodes with conditions. In order to evaluate these candidate task decomposition schemes without affecting the main simulation training data stream, the task orchestration processing module copies the current simulation situation data to form a shadow simulation copy, and injects each candidate task decomposition scheme into the corresponding shadow simulation copy. The shadow simulation copy runs at the same simulation step size, and records entity position data, detection result data, resource consumption data, and threat change data at the end of each simulation step size, thereby obtaining multiple sets of situation evolution trajectory data.
[0030] Specifically, after the multi-dimensional parameter search interval has been written back to the intent probability task graph, the task orchestration processing module encodes each dimension according to a predetermined number of bits. The predetermined number of bits can be set to three to six bits depending on the computational power of the training system and the task precision; in this embodiment, four bits can be used, so that each continuous interval is divided into sixteen discrete levels. For discrete attributes such as the target category dimension, the encoding is performed according to the target category index in the task template table. After encoding, the task region search interval yields several region grid points; after encoding, the target category search interval yields the target category value; after encoding, the trigger timing search interval yields the selectable simulation step size; and after encoding, the treatment intensity search interval yields the treatment level sequence. These results together constitute the discrete value sequence corresponding to each dimension.
[0031] After each discrete value sequence is generated, the task orchestration processing module combines them according to the node connection relationships in the intent probability task graph to obtain multiple parameter combination data. The combination is not an unconstrained Cartesian expansion of all values, but is constrained by the current simulation situation data. The task orchestration processing module reads each parameter combination data one by one, determining whether the execution object is in an entity available state, whether the task area and the current position of the execution object meet the task distance limit, and whether there is a resource occupancy status in the resource status data. If the parameter combination data requires a simulated equipment entity to execute two mutually exclusive tasks simultaneously, or requires the task action to be completed outside the task distance limit, the parameter combination data is discarded. The remaining valid parameter combination data is filled into the conditional fuzzy task nodes in the intent probability task graph, generating candidate task paths with definite parameters. Each candidate task path is then converted into a candidate task decomposition scheme according to the task nodes, execution object data, and node connection relationships, recording the execution order, execution object data, task area, triggering time, and handling intensity of each task node.
[0032] Furthermore, each candidate task decomposition scheme is injected into a shadow simulation copy formed by replicating the current simulation situation data, resulting in multiple sets of situation evolution trajectory data, including: Copy the entity status data, environmental status data, and resource status data from the current simulation situation data to generate a shadow simulation copy with the same number of candidate task decomposition schemes. Write the task nodes in each candidate task decomposition scheme into the corresponding shadow simulation copy in the execution order, and drive the shadow simulation copy to perform the simulation step by the same simulation step. At the end of each simulation step, record entity location data, detection result data, resource consumption data, and threat change data in the shadow simulation copy; According to the candidate task decomposition scheme number, the entity location data, detection result data, resource consumption data, and threat change data are sequentially spliced together to form situational evolution trajectory data corresponding to the candidate task decomposition scheme.
[0033] Specifically, after the candidate task decomposition schemes are formed, the task orchestration processing module copies the entity state data, environmental state data, and resource state data from the current simulation situation data to generate shadow simulation copies with the same number of candidate task decomposition schemes. Entity state data includes entity identifiers, locations, velocities, status flags, and payload status; environmental state data includes terrain, weather, electromagnetic environment, and visibility; and resource state data includes available firepower, reconnaissance payloads, communication links, and task occupancy. During copying, the shadow simulation copies maintain the same initial timestamp and random seed configuration as the current simulation situation data, ensuring that the differences in the extrapolation between different candidate task decomposition schemes mainly stem from task parameters rather than fluctuations in the simulation environment.
[0034] The task nodes in each candidate task decomposition scheme are written into the corresponding shadow simulation copy in execution order. The shadow simulation copy is simulated according to the same simulation step size. The simulation step size can be set to one second, two seconds, or the basic step size specified by the training system; in this embodiment, one second is used as an example. At the end of each simulation step, the task orchestration processing module records the entity position data, detection result data, resource consumption data, and threat change data in the shadow simulation copy. Among them, the entity position data is used to describe the position changes of the simulated equipment entity and the target entity; the detection result data is used to record whether the target has been detected, the identification confidence level, and the detection source; the resource consumption data is used to record ammunition, payload time, communication occupation, and energy consumption; and the threat change data is used to record changes in the enemy threat level, exposure risk, and suppression status.
[0035] The task orchestration module sequentially concatenates entity location data, detection result data, resource consumption data, and threat change data according to the candidate task decomposition scheme number, forming situation evolution trajectory data corresponding to the candidate task decomposition scheme. This situation evolution trajectory data retains the time sequence, enabling subsequent path comparison to accurately identify common task nodes in different candidate task paths.
[0036] Path comparison is performed on multiple sets of situation evolution trajectory data to extract common task nodes that meet logical consistency and have risk values below the preset risk threshold, forming a set of tentative task instructions, and then the set of tentative task instructions is incorporated into the main simulation training data stream for execution. Furthermore, path comparison is performed on multiple sets of situational evolution trajectory data to extract common task nodes that meet logical consistency and have risk values below a preset risk threshold, forming a set of exploratory task instructions, including: According to the execution order of the task nodes, prefix comparison is performed on the candidate task paths corresponding to multiple sets of situation evolution trajectory data to extract the common task nodes that appear in each candidate task path. Read the resource consumption data, threat change data, and detection result data of the common task nodes in the corresponding situational evolution trajectory data; The resource consumption data, threat change data, and detection result data are normalized and then weighted and summed according to the weights of the corresponding task types in the task template table to obtain the risk value of the common task node. Retain public task nodes with risk values below the preset risk threshold, and delete public task nodes that have conflicting execution objects; Based on the execution order of the retained common task nodes in the intent probability task graph, generate a set of exploratory task instructions.
[0037] In one specific embodiment, after multiple sets of situation evolution trajectory data are generated, the task orchestration processing module does not immediately select the complete solution with the highest score. Instead, it first performs path comparison on the multiple sets of situation evolution trajectory data. During path comparison, the execution order of task nodes in the candidate task paths is used as the main line to find common task nodes that appear in different candidate task paths and whose execution relationships do not conflict. For these common task nodes, the task orchestration processing module reads their resource consumption data, threat change data, and detection result data in the situation evolution trajectory data and calculates the risk value of the common task nodes.
[0038] The preset risk threshold is determined by the task orchestration module before path comparison, taking into account the specific training environment and the dynamic real-time battle situation. When determining the preset risk threshold, the system first extracts the set difficulty level of the current training subject and the pre-configured fault tolerance benchmark. Simultaneously, it reads the overall resource sufficiency index and the average threat level of the global environment from the current simulation situation data. Then, using the fault tolerance benchmark as the initial baseline, it applies a decay penalty coefficient as a negative constraint and performs appropriate positive compensation mapping according to the resource sufficiency index. This comprehensive calculation determines the critical risk limit that the training system can withstand under the current time window, which serves as the preset risk threshold for this comparison. This adaptive value selection strategy based on situation and subject setting can ensure that the conditions for allowing probing actions are strictly tightened during the confrontation exercise phase when resources are scarce or high pressure is tight, while the tolerance is appropriately relaxed when the situation is stable or the reconnaissance is in the early stage. In this way, under the premise of absolutely ensuring the security of the main simulation training data stream and preventing tactical collapse, the most reasonable and realistic probing boundary is provided for the generation and issuance of subsequent probing mission instruction sets.
[0039] Common task nodes with risk values below a preset risk threshold and satisfying logical consistency are retained and formed into a set of exploratory task instructions according to the execution order in the intent probability task graph. The set of exploratory task instructions is then incorporated into the main simulation training data stream for execution. Its role is not to complete all combat intentions, but to first promote observable changes in the simulation situation within a low-risk range, so that the system can further collect feedback from the command and control end on these changes.
[0040] Specifically, after obtaining multiple sets of situation evolution trajectory data, the task orchestration module performs prefix comparison on the candidate task paths corresponding to the multiple sets of situation evolution trajectory data according to the execution order of the task nodes. Prefix comparison starts from the starting task node of each candidate task path. Task nodes that appear in all candidate task paths and whose execution object data is consistent are temporarily designated as common task nodes. When different task nodes appear at the same position in a candidate task path, or when the execution object data conflicts, the prefix comparison stops at that position to avoid mistakenly identifying subsequent tasks with significant discrepancies as exploratory content. Through this process, common task nodes usually fall on the pre-execution actions that all candidate task decomposition schemes need to perform, exhibiting good reversibility and low decision-making risk.
[0041] When calculating the risk value of common task nodes, the task orchestration module reads the resource consumption data, threat change data, and detection result data of the common task node from the corresponding situational evolution trajectory data. Since the three types of data have different dimensions, resource consumption data can be normalized according to the resource occupancy ratio, threat change data can be normalized according to the magnitude of threat level changes, and detection result data can be normalized according to changes in the probability of being detected and the confidence level of identification. The task template table stores the weights of different task types; for example, reconnaissance tasks focus more on detection result data, while strike tasks focus more on resource consumption data and threat change data. The task orchestration module performs a weighted sum according to the weights of the corresponding task types in the task template table to obtain the risk value of the common task node. If the risk value of a common task node is lower than a preset risk threshold, the common task node is retained; if multiple common task nodes contend for the same execution object data, the task orchestration module deletes the common task nodes with conflicting execution objects. The retained common task nodes are arranged according to the execution order in the intent probability task graph, forming a tentative task instruction set. Since the exploratory task instruction set comes from the intersection of multiple candidate task paths and has been filtered by risk values, it is more suitable as the initial action in the main simulation training data stream.
[0042] Specifically, when calculating the risk value of common task nodes, the task orchestration module first needs to address the issues of inconsistent data dimensions and inconsistent directions of risk contribution. Specifically, resource consumption data and increased threats typically directly increase execution risk, while detection result data showing no target detected or being detected in reverse will translate into detection failure risk. The task orchestration module performs max-min normalization on these three types of data, mapping them to a standard range of zero to one, and calculates the final risk value using a specific formula, which is as follows: ; In this calculation formula, This represents the final risk value of the common task node. (The right side of the formula...) This represents the normalized resource consumption evaluation score. The larger the value, the more ammunition or reconnaissance payload resources the node consumes. The normalized threat change assessment reflects the increase in threat caused by enemy fire against the target of this node in the shadow simulation. This represents the risk level of detection failure or entity exposure after the detection result data has been transformed and normalized. The three normalized variables are multiplied by their corresponding weights in the formula. This represents the resource consumption weight read from the task template table that matches the current task type. Represents the weight of threat changes. The weights representing the detection results are always equal to one, and their specific values are dynamically adjusted depending on the mission type. For example, when evaluating a forward reconnaissance mission, The value of will be significantly higher than the other two, thus ensuring that multi-dimensional business data can be reasonably integrated into a risk value that accurately reflects the degree of potential danger.
[0043] Specifically, regarding the normalized resource consumption evaluation score... The task orchestration module uses a saturation truncation mapping method based on the scarcity of multiple resource types for calculation. The specific formula is as follows: In this computational logic, because common task nodes consume various resources (such as fuel, ammunition, computing power, or communication bandwidth) during shadow simulation, the subscript... Used to iterate through the types of resources consumed. Representing the The specific absolute amount of resource consumption. Indicates the first The scarcity weight of each resource type (the scarcer the resource, the higher the weight); the weighted consumption of each resource type is summed and then divided by the maximum allowable overall resource consumption tolerance for the current entity. This yields a resource consumption evaluation score. Furthermore, to meet the requirements of subsequent standard normalization calculations for the total risk value, a cutoff function of min(1.0,X) is nested outside the formula to ensure that when abnormally high energy consumption causes the result inside the formula to be greater than 1, [the calculation is incomplete]. The value will be forcibly capped at 1.0, thus accurately mapping it to a resource consumption evaluation level between 0 and 1, objectively reflecting the logistical or support pressure and risk caused by task execution.
[0044] Threat change assessment after normalization The specific conversion formula is as follows: .in, and These represent the comprehensive threat index of the friendly entity before and after the execution of the shadow simulation situation evolution trajectory of the public task node; This represents the maximum threat threshold that the system can withstand. The smoothing adjustment coefficient (ranging from 0 to 1) can be obtained by fitting historical data. Through the above calculations, complex threat change data can be converted into a threat change assessment level of 0 to 1, which both punishes the placement of troops in high-risk areas and strictly controls reckless probing actions that could lead to a rapid deterioration of the situation.
[0045] Regarding the risk level of detection failure or entity exposure after conversion and normalization of detection result data. The task orchestration module employs a two-factor weighted mapping method based on detection effectiveness and exposure probability for transformation calculation. The specific transformation formula is as follows: .in, Represents the target discovery probability or detection coverage effectiveness extracted from the detection results data (value between 0 and 1), (1- This is used to quantify the risk of mission failure due to undetected targets or loss of vision; This represents the probability (between 0 and 1) of a friendly mission entity being detected by the enemy's sensor network during shadow simulation due to activating active radar or breaking silence. and These are the failure penalty coefficient and exposure penalty coefficient set by the system, respectively. Since location exposure in actual combat often means a survival crisis, typically... The value will be significantly greater than The above calculation formula allows for the unified fusion and compression of complex, multi-dimensional, non-standard detection results data, mapping them to a standard range of zero to one, ensuring the final generated data... It can accurately balance the hidden risks of failing to achieve tactical objectives with the fatal risk of being targeted in reverse.
[0046] Collect feedback feature vectors after the execution of the exploratory task instruction set, correct the multidimensional parameter search interval and node probability based on the feedback feature vectors, and calculate the distribution entropy value of the intent probability task graph; Furthermore, the feedback feature vectors after the execution of the exploratory task instruction set are collected, including: After the exploratory task instruction set is incorporated into the main simulation training data stream, the field of view center coordinate change data, entity selection state change data, and supplementary instruction text generated by the command and control terminal within the preset feedback time window are recorded. The coordinate change data of the field of view center is matched with the search interval of the task area corresponding to the exploratory task instruction set to generate regional interest features; The entity's selected state change data is matched with the execution object data corresponding to the exploratory task instruction set to generate object confirmation features; New action instruction feature words and new conditional trigger words are extracted from the supplementary instruction text to generate semantic correction features; The region-of-interest features, object-confirmation features, and semantic correction features are concatenated in chronological order of collection to form a feedback feature vector.
[0047] Furthermore, the multi-dimensional parameter search interval and node probability are corrected based on the feedback feature vector, including: When the region of interest feature falls within the task region search interval, the boundary of the task region search interval is compressed to the coverage area corresponding to the region of interest feature, and the node probability of the node branch corresponding to the coverage area is increased. When the object confirmation features are consistent with the execution object data, the target category search interval is compressed to the target attribute range corresponding to the execution object data, and the node probability of the node branch corresponding to the target attribute range is increased; When the semantic correction feature contains new conditional trigger words, the new conditional trigger words are matched with the trigger categories in the task template table, the original trigger categories in the conditionally ambiguous task nodes are replaced, and the trigger timing search range and the handling intensity search range are reconstructed accordingly. The node probabilities of each branch after correction are normalized so that the sum of the node probabilities under the same condition fuzzy task node remains one. Write the corrected multidimensional parameter search range and node probabilities back into the intent probability task graph.
[0048] In one specific embodiment, after the exploratory task instruction set enters the main simulation training data stream, the task orchestration processing module begins to collect feedback feature vectors. The feedback data comes from the human-computer interaction recording interface of the same control terminal and mainly includes field-of-view center coordinate change data, entity selection state change data, and supplementary instruction text. The preset feedback time window can be set according to the rhythm of the training subjects, for example, collecting feedback data within twenty seconds after the exploratory task instruction set begins execution. Field-of-view center coordinate change data is used to determine whether the control terminal is paying attention to the task area search interval corresponding to the exploratory task instruction set; entity selection state change data is used to determine whether the control terminal has confirmed the execution object data corresponding to the exploratory task instruction set; and the supplementary instruction text is parsed to extract new action instruction feature words and new conditional trigger words to correct the original conditionally ambiguous task nodes. The task orchestration processing module concatenates the area attention features, object confirmation features, and semantic correction features in the order of collection time to form a feedback feature vector, and uses the feedback feature vector to correct the multi-dimensional parameter search interval and node probability.
[0049] During the correction process, if the regional focus feature falls within the task region search interval, it indicates a correspondence between the command and control end's focus range and the situational changes generated by the exploratory task instruction set. The task orchestration module then compresses the boundary of the task region search interval to the coverage area corresponding to the regional focus feature, while increasing the node probability of the corresponding node branch within that coverage area. If the object confirmation feature matches the execution object data, the target category search interval is compressed to the target attribute range corresponding to the execution object data, and the node probability of the corresponding node branch within that target attribute range is increased. If the semantic correction feature contains new conditional trigger words, the new conditional trigger words are matched with the trigger categories in the task template table. The matching results replace the original trigger categories in the conditionally ambiguous task node, and the trigger timing search interval and the handling intensity search interval are reconstructed accordingly. After correction, each node branch under the same conditionally ambiguous task node needs to undergo node probability normalization processing to ensure that the sum of node probabilities remains one. Subsequently, the task orchestration module calculates the distribution entropy value of the intent probability task graph according to the corrected node probabilities. The distribution entropy value is used to measure the degree of uncertainty in the current intention probability task graph. During calculation, the probability dispersion of node probabilities under each conditional fuzzy task node is statistically analyzed. The more concentrated the node probabilities are, the lower the distribution entropy value. The preset convergence threshold can be configured according to the task complexity, for example, set to 0.5; the preset number of steps can be set to three to limit the number of times shadow inference is repeated.
[0050] After correcting the multi-dimensional parameter search range and node probabilities, the task orchestration processing module needs to quantitatively evaluate the uncertainty of the entire intent probability task graph to determine whether the system's understanding of the command intent has reached a level sufficient to translate it into a final instruction. This evaluation process is mainly achieved by calculating the distribution entropy value of the intent probability task graph. Utilizing the mathematical properties of information entropy, it can objectively reflect the concentration or dispersion of probability distributions under various conditionally ambiguous task nodes. During calculation, the task orchestration processing module directly uses the following formula: ; In the calculation formula, the left side of the equation The distribution entropy value, representing the final calculated value of the intention probability task graph, It is not the entropy value of a single conditional fuzzy task node, but rather a global evaluation index obtained by averaging the node distribution entropy values of all conditional fuzzy task nodes. It is the upper limit parameter in the first-level summation calculation on the right-hand side of the equation. , representing the total number of conditionally ambiguous task nodes currently existing in the intent probability task graph, subscript This is used to iterate through these N nodes sequentially during the accumulation process. The upper bound parameter is nested within the second level of summation calculation. Representing the The total number of node branches extending outwards from a given conditionally fuzzy task node, index This indicates that the index numbers of the specific branches are traversed sequentially under the current node. This represents the result after the feedback feature vectors have been corrected in the previous steps and normalized again. The first conditionally fuzzy task node The formula calculates the node probability values corresponding to each branch. It uses a base-2 logarithm to obtain the self-information corresponding to each branch probability, multiplies it by the node probability itself, and accumulates these values sequentially to accurately capture the decision divergence within each fuzzy node. Finally, by taking the overall negative value and dividing by the total number of nodes N, it obtains the average uncertainty level of the entire intention probability task graph in the current state. Following the above calculation rules, when the probabilities under each conditional fuzzy task node gradually converge to a certain optimal or most definite branch, the calculated distribution entropy value... It will then show a continuous downward trend.
[0051] Specifically, after the exploratory task instruction set is incorporated into the main simulation training data stream, the task orchestration processing module records the field-of-view center coordinate change data, entity selection state change data, and supplementary instruction text generated by the command and control terminal within a preset feedback time window. The field-of-view center coordinate change data consists of continuously sampled field-of-view center coordinates, which the task orchestration processing module matches with the task region search interval corresponding to the exploratory task instruction set. If the field-of-view center coordinate change data enters the task region search interval within the preset feedback time window and stays for a preset time percentage, the task orchestration processing module generates a region attention feature; if the field-of-view center coordinate change data only briefly passes through the region, the intensity of the region attention feature is low. The entity selection state change data is used to match the identifier of the execution object data corresponding to the exploratory task instruction set. When the match is consistent, an object confirmation feature is generated; when the match is inconsistent, the reselected entity identifier is recorded, providing a basis for correcting the target category search interval. The supplementary instruction text still extracts new action instruction feature words and new condition trigger words according to the aforementioned parsing method, forming semantic correction features. The region focus features, object confirmation features, and semantic correction features are concatenated in the order of collection time to form a feedback feature vector, which serves as the direct input for correcting the multidimensional parameter search interval and node probability.
[0052] When using feedback feature vectors to correct the multidimensional parameter search interval and node probability, the task orchestration processing module first reads the regional interest features. When a regional interest feature falls within the task regional search interval, the boundary of the task regional search interval is compressed to the coverage area corresponding to the regional interest feature. The coverage area can be determined by the minimum bounded area formed by the field of view center coordinate change data within a preset feedback time window, or by superimposing the field of view ratio to obtain the actual visible range. The node probability of the node branch corresponding to this coverage area is increased, and the increase can be determined according to the duration and positional overlap of the regional interest feature. When the object confirmation feature is consistent with the execution object data, the target category search interval is compressed to the target attribute range corresponding to the execution object data, while simultaneously increasing the node probability of the node branch corresponding to this target attribute range. If the object confirmation feature points to a new entity identifier, the task orchestration processing module will read the target attribute of the new entity identifier in the current simulation situation data and correct the target category search interval with this target attribute range.
[0053] When the semantically corrected features contain new conditional trigger words, the task orchestration module matches these new trigger words with the trigger categories in the task template table. Upon successful matching, the original trigger categories in the conditionally ambiguous task node are replaced, and the trigger timing search interval and the disposal intensity search interval are reconstructed accordingly. For example, if the new conditional trigger word corresponds to a stricter timing constraint, the trigger timing search interval will be narrowed; if the new conditional trigger word corresponds to a lower disposal level, the disposal intensity search interval will exclude high-consumption disposal levels. The node probabilities of each corrected node branch are normalized, ensuring that the sum of node probabilities under the same conditionally ambiguous task node remains one. Normalization can be reassigned according to the proportion of the corrected probability value of each node branch to the total probability value, avoiding the loss of comparability of node probabilities due to the superposition of different correction sources. After completion, the corrected multidimensional parameter search interval and node probabilities are written back to the intent probability task graph, and the original intent probability task graph is updated to the new state.
[0054] When the distribution entropy value is lower than the preset convergence threshold, the task path with the highest node probability is extracted and converted into an equipment-level task operation sequence; when the distribution entropy value is not lower than the preset convergence threshold within a preset number of steps, the candidate task decomposition scheme is regenerated based on the corrected multidimensional parameter search interval until the equipment-level task operation sequence is obtained, and the equipment-level task operation sequence is output to complete the task arrangement.
[0055] Furthermore, when the distribution entropy value does not fall below a preset convergence threshold within a preset number of steps, a minimum clarification data generation step is also included: Read the node probability of each conditional fuzzy task node in the intent probability task graph, and calculate the distribution entropy value corresponding to each conditional fuzzy task node; The distribution entropy value corresponding to each conditional fuzzy task node can be calculated using the following formula: ; in, Indicates the first The uncertainty of a conditionally fuzzy task node itself. This indicates the number of branches corresponding to this node. Indicates the first The first conditionally fuzzy task node The node probability of each branch. This node distribution entropy value is used to compare the divergence of the probability distribution within nodes of fuzzy tasks under different conditions. The system can select... The task node with the most ambiguous conditions is designated as the task node to be clarified.
[0056] Select the conditionally ambiguous task node with the highest distribution entropy value as the task node to be clarified. From the multi-dimensional parameter search range corresponding to the task node to be clarified, extract the parameter dimension with the largest range as the dimension to be clarified; Based on the task type and dimension to be clarified of the task node to be clarified, select the corresponding clarification text fragment from the task template table to generate the minimum clarification data; The system receives the clarification result data returned for the minimum clarification data, converts the clarification result data into a limited range of the dimension to be clarified, and recalculates the node probability after correcting the multidimensional parameter search interval using the limited range.
[0057] Furthermore, the task path with the highest node probability is extracted and converted into an equipment-level task operation sequence, including: In the intention probability task graph, the node branch with the highest node probability is selected step by step from the starting task node until the ending task node is reached, thus forming the target task path. Read the execution object data, task parameter data, and node connection relationships corresponding to each task node in the target task path; The execution object data, task parameter data, and node connection relationships are mapped according to the operation fields that can be identified by the simulated equipment entity to generate an initial equipment-level task operation sequence. The initial equipment-level mission operation sequence is compared with the already executed trial mission instruction set, duplicate mission operations are deleted, and the latest entity state data generated by the trial mission instruction set is retained. Based on the latest entity status data, correct the location, timing, and resource fields in the remaining task operations to obtain and output the equipment-level task operation sequence.
[0058] In one specific embodiment, when the distribution entropy value is lower than a preset convergence threshold, the task orchestration processing module considers the intent probability task graph to have converged to an executable state. It then selects nodes step-by-step from the starting task node along the branch with the highest node probability until reaching the ending task node, forming the target task path. The task nodes, execution object data, task parameter data, and node connection relationships in the target task path are then converted into an equipment-level task operation sequence. If the distribution entropy value does not fall below the preset convergence threshold within a preset number of steps, the task orchestration processing module does not directly abandon orchestration. Instead, it regenerates candidate task decomposition schemes based on the corrected multi-dimensional parameter search interval and re-enters the loop of shadow simulation copy deduction, path comparison, exploratory task instruction set execution, and feedback feature vector acquisition. This loop continues until an equipment-level task operation sequence is obtained. If convergence is not achieved after reaching the preset number of steps, a minimum clarification data generation step can be entered to break the uncertainty with a minimal amount of supplementary data. The final equipment-level task operation sequence is output to the main simulation training data stream, which drives the corresponding simulated equipment entity to execute, thereby completing the task orchestration.
[0059] The preset convergence threshold is a flexible control benchmark dynamically calculated by the system based on the tactical attributes of the core mission and the current urgency of the confrontation. Specifically, the task orchestration module first retrieves the baseline fault tolerance rate of the type of task from the simulation training task library based on the initially parsed action instruction feature words and back-maps it to the basic convergence threshold. For example, an extremely low basic threshold is assigned to irreversible fire strike tasks to pursue absolute intent clarity, while a relatively high basic threshold is assigned to reversible area maneuver tasks to improve system response agility. Subsequently, the system extracts the situation evolution rate and enemy-friendly combat time window from the current simulation situation data across dimensions and quantifies the current confrontation urgency coefficient. Finally, based on the basic convergence threshold, the confrontation urgency coefficient is used to dynamically adjust it in a positive or negative way. In critical moments when the battlefield situation deteriorates rapidly and the decision-making time is extremely compressed, the convergence threshold is adaptively increased to tolerate slight non-core dimension intent ambiguity, forcing the algorithm to converge quickly and output equipment-level task operation sequences to seize the opportunity. In the period of calm situation and ample time for reconnaissance deployment, the convergence threshold is strictly reduced to drive the algorithm to make full use of feedback feature vectors and even cooperate with the minimum clarification data generation step for multiple rounds of deep focusing. Based on the adaptive threshold adjustment mechanism of the mission bottom line and real-time situational pressure, it perfectly solves the problems of insufficient accuracy in the steady period and deadlock in the high-pressure period caused by the reliance on fixed convergence conditions in traditional algorithms, and ensures the dynamic optimal balance between the timeliness of command and control decision-making and the fidelity of intent parsing in complex adversarial environments.
[0060] Specifically, if the distribution entropy value of the intent probability task graph does not fall below a preset convergence threshold within a preset number of steps, the task orchestration processing module proceeds to the minimum clarification data generation step. This step does not query all uncertain content; instead, it first reads the node probabilities of each conditionally ambiguous task node in the intent probability task graph and calculates the distribution entropy value corresponding to each conditionally ambiguous task node. The conditionally ambiguous task node with the highest distribution entropy value indicates that its node branches are still scattered, and the task orchestration processing module selects it as the task node to be clarified.
[0061] Subsequently, from the multi-dimensional parameter search intervals corresponding to the task nodes to be clarified, the interval spans of the task region dimension, target category dimension, triggering timing dimension, and handling intensity dimension are compared, and the parameter dimension with the largest interval span is extracted as the dimension to be clarified. The comparison of interval spans adopts a unified normalized scale, so that the spatial range, temporal range, and handling level can all be compared under the same standard.
[0062] Once the task nodes and dimensions to be clarified are determined, the task orchestration module selects the corresponding clarification text fragments from the task template table based on the task type and dimension of the task node to be clarified, generating minimal clarification data. Minimal clarification data focuses solely on the dimension to be clarified, minimizing the amount of additional information required from the accuser. After receiving the clarification result data returned for the minimal clarification data, the task orchestration module converts the clarification result data into a defined range for the dimension to be clarified. For example, if the clarification result data points to a task area, it is converted into a defined range for the task area dimension; if it points to a certain level of action, it is converted into a defined range for the level of action intensity dimension. This defined range is then used to correct the multi-dimensional parameter search interval and recalculate the node probability. Because the minimal clarification data is generated only after shadow deduction and feedback correction have failed to converge, the system can achieve intent convergence in most cases through implicit feedback, introducing only a small amount of supplementary data when necessary.
[0063] Regarding the specific method for selecting the corresponding clarification text fragment from the task template table, the task orchestration processing module first uses the task type and the dimension to be clarified of the task node as the joint search key, performs index matching in the task template table, directly locates and extracts the preset semantic query template that is strongly bound to the specific intention disagreement. For example, when the task type is "area air defense" and the dimension to be clarified is "disposal intensity", the placeholder sentence "Please confirm the firing authorization level: [ ] or [ ]?" is extracted. Subsequently, the system reads the discrete boundary values or high-probability candidate categories, such as warning fire and free engagement, within the current multi-dimensional parameter search interval for the dimension to be clarified. It then uses specific situational context parameters as dynamic variables to fill the empty slots in the extracted preset template. Through this combination of static semantic template mapping and dynamic interval parameter instantiation, abstract probability distribution doubts can be instantly transformed into highly focused, clearly defined, and concisely expressed natural language queries. This ensures that commanders only need to make a minimal selection confirmation to accurately eliminate the logical discrepancies with the highest entropy value in the system.
[0064] Once the intent probability task graph reaches convergence, the task orchestration processing module extracts the task path with the highest node probability and converts it into an equipment-level task operation sequence. In the intent probability task graph, the task orchestration processing module starts from the initial task node and selects the node branch with the highest node probability level by level until it reaches the end task node, forming the target task path. If there are node branches with the same node probability at a certain level, the node branch with lower resource consumption and a better match to the distance condition is selected based on the resource status data and task distance constraints in the current simulation situation data. After the target task path is formed, the task orchestration processing module reads the execution object data, task parameter data, and node connection relationships corresponding to each task node in the target task path. The task parameter data includes determined parameters such as task area, target category, trigger timing, and handling intensity. The node connection relationships are used to maintain the sequence and dependencies in the equipment-level task operation sequence.
[0065] During field mapping, the task orchestration module transforms the execution object data, task parameter data, and node connection relationships according to the operation fields identifiable by the simulated equipment entity, generating an initial equipment-level task operation sequence. The operation fields identifiable by the simulated equipment entity generally include operation type, execution object, spatial, timing, resource, and termination condition fields. Execution object data is written to the execution object field, the task area to the spatial field, the trigger timing to the timing field, the resource requirements corresponding to the handling intensity to the resource field, and node connection relationships are transformed into preconditions between each operation. Since the exploratory task instruction set has already been executed in the main simulation training data stream, the task orchestration module needs to compare the initial equipment-level task operation sequence with the already executed exploratory task instruction set, deleting duplicated task operations and retaining the latest entity state data generated by the exploratory task instruction set. Finally, the task orchestration module corrects the position, timing, and resource fields in the remaining task operations based on the latest entity state data, obtaining and outputting the equipment-level task operation sequence. The resulting equipment-level task operation sequence can be naturally connected with the results of the exploratory executions that have already occurred, without generating repetitive actions or state rollbacks in the main simulation training data stream.
[0066] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention, they should all fall within the protection scope of the present invention.
[0067] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0068] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. An equipment simulation training task scheduling method oriented to command and coordination cooperation, characterized in that, Includes the following steps: Simultaneously acquire unstructured command text, interactive behavior feature data, and current simulation situation data from the command and control terminal. The interactive behavior feature data includes the coordinates of the field of view center and the entity selection status. Unstructured instruction text is parsed into action instruction feature words, tactical entity nouns and condition trigger words. Combined with interactive behavior feature data and current simulation situation data, an intent probability task graph containing multiple task nodes and node probabilities is generated. A multi-dimensional parameter search interval is formed for the conditionally ambiguous task nodes in the intent probability task graph. Discrete sampling of the intent probability task map is performed according to the multidimensional parameter search interval to generate multiple candidate task decomposition schemes. Each candidate task decomposition scheme is then injected into the shadow simulation copy formed by copying the current simulation situation data to obtain multiple sets of situation evolution trajectory data. Path comparison is performed on multiple sets of situation evolution trajectory data to extract common task nodes that meet logical consistency and have risk values below the preset risk threshold, forming a set of tentative task instructions, and then the set of tentative task instructions is incorporated into the main simulation training data stream for execution. Collect feedback feature vectors after the execution of the exploratory task instruction set, correct the multidimensional parameter search interval and node probability based on the feedback feature vectors, and calculate the distribution entropy value of the intent probability task graph; When the distribution entropy value is lower than the preset convergence threshold, the task path with the highest node probability is extracted and converted into an equipment-level task operation sequence; when the distribution entropy value is not lower than the preset convergence threshold within a preset number of steps, the candidate task decomposition scheme is regenerated based on the corrected multidimensional parameter search interval until the equipment-level task operation sequence is obtained, and the equipment-level task operation sequence is output to complete the task arrangement.
2. The method for arranging equipment simulation training tasks for command and control coordination according to claim 1, characterized in that, The mapping generates an intent probability task graph containing multiple task nodes and node probabilities, including: Unstructured command text is segmented and syntactically ordered to extract action command feature words, tactical entity nouns, and conditional trigger words. Read the task template table from the simulation training task library, match the action instruction feature words with the standard task actions in the task template table, and generate candidate task nodes; By associating tactical entity names with entity identifiers in the current simulation situation data, the execution object data corresponding to the candidate task nodes can be obtained; Based on the order of action instruction feature words in the unstructured instruction text and the positional relationship of the execution object data in the current simulation situation data, the execution order between candidate task nodes is determined and node connection relationships are generated. Mark the candidate task nodes corresponding to the conditional trigger words as conditional fuzzy task nodes, and generate node branches based on the trigger categories corresponding to the conditional trigger words in the task template table; Construct an intent probability task graph based on node connection relationships, conditional fuzzy task nodes, and node branches, and assign initial node probabilities to node branches according to the number of trigger categories corresponding to each node branch.
3. The method for arranging equipment simulation training tasks for command and control coordination according to claim 2, characterized in that, For conditionally fuzzy task nodes in the intent probability task graph, a multi-dimensional parameter search interval is formed, including: Read the trigger category corresponding to the fuzzy task node and break down the trigger category into task area dimension, target category dimension, trigger timing dimension and handling intensity dimension; Based on the projection position of the field of view center coordinates in the current simulation situation data, the central region of the task region dimension is determined, and the task region search interval is formed by combining the distribution range of entities around the central region. Based on the entity identifier corresponding to the selected entity state, read the target attribute of the entity identifier in the current simulation situation data to form a target category search range; Based on the timing constraints corresponding to the conditional trigger words in the task template table, and combined with the enemy-friendly distance change data in the current simulation situation data, a trigger timing search interval is formed; Based on the response level corresponding to the conditional trigger words, read the firepower availability status and reconnaissance availability status in the current simulation situation data to form a response intensity search range; The task area search range, target category search range, trigger timing search range, and treatment intensity search range are merged into a multi-dimensional parameter search range and written back to the corresponding conditional fuzzy task node.
4. The method for arranging equipment simulation training tasks for command and control coordination according to claim 3, characterized in that, Discrete sampling of the intent probability task graph is performed according to the multidimensional parameter search interval to generate multiple candidate task decomposition schemes, including: Encode each dimension in the multidimensional parameter search interval with a predetermined number of bits to obtain the discrete value sequence corresponding to each dimension; The discrete value sequences are combined according to the node connection relationship in the intention probability task graph to generate multiple parameter combination data; Parameter combinations that do not meet the requirements of entity availability, task distance limits, and resource occupancy in the current simulation situation data are removed to obtain valid parameter combinations. Fill the conditional fuzzy task node in the intent probability task graph with the effective parameter combination data to generate a candidate task path with definite parameters. Based on the task nodes, execution object data, and node connection relationships in the candidate task path, multiple candidate task decomposition schemes are obtained.
5. The method for arranging equipment simulation training tasks for command and control coordination according to claim 4, characterized in that, Each candidate task decomposition scheme is injected into a shadow simulation copy formed by replicating the current simulation situation data, resulting in multiple sets of situation evolution trajectory data, including: Copy the entity status data, environmental status data, and resource status data from the current simulation situation data to generate a shadow simulation copy with the same number of candidate task decomposition schemes. Write the task nodes in each candidate task decomposition scheme into the corresponding shadow simulation copy in the execution order, and drive the shadow simulation copy to perform the simulation step by the same simulation step. At the end of each simulation step, record entity location data, detection result data, resource consumption data, and threat change data in the shadow simulation copy; According to the candidate task decomposition scheme number, the entity location data, detection result data, resource consumption data, and threat change data are sequentially spliced together to form situational evolution trajectory data corresponding to the candidate task decomposition scheme.
6. The method for arranging equipment simulation training tasks for command and control coordination according to claim 5, characterized in that, Path comparison is performed on multiple sets of situational evolution trajectory data to extract common task nodes that meet logical consistency and have risk values below a preset risk threshold, forming a set of exploratory task instructions, including: According to the execution order of the task nodes, prefix comparison is performed on the candidate task paths corresponding to multiple sets of situation evolution trajectory data to extract the common task nodes that appear in each candidate task path. Read the resource consumption data, threat change data, and detection result data of the common task nodes in the corresponding situational evolution trajectory data; The resource consumption data, threat change data, and detection result data are normalized and then weighted and summed according to the weights of the corresponding task types in the task template table to obtain the risk value of the common task node. Retain public task nodes with risk values below the preset risk threshold, and delete public task nodes that have conflicting execution objects; Based on the execution order of the retained common task nodes in the intent probability task graph, generate a set of exploratory task instructions.
7. The method for arranging equipment simulation training tasks for command and control coordination according to claim 6, characterized in that, Collect feedback feature vectors after the execution of the exploratory task instruction set, including: After the exploratory task instruction set is incorporated into the main simulation training data stream, the field of view center coordinate change data, entity selection state change data, and supplementary instruction text generated by the command and control terminal within the preset feedback time window are recorded. The coordinate change data of the field of view center is matched with the search interval of the task area corresponding to the exploratory task instruction set to generate regional interest features; The entity's selected state change data is matched with the execution object data corresponding to the exploratory task instruction set to generate object confirmation features; New action instruction feature words and new conditional trigger words are extracted from the supplementary instruction text to generate semantic correction features; The region-of-interest features, object-confirmation features, and semantic correction features are concatenated in chronological order of collection to form a feedback feature vector.
8. The method for arranging equipment simulation training tasks for command and control coordination according to claim 7, characterized in that, The multidimensional parameter search interval and node probability are corrected based on the feedback feature vector, including: When the region of interest feature falls within the task region search interval, the boundary of the task region search interval is compressed to the coverage area corresponding to the region of interest feature, and the node probability of the node branch corresponding to the coverage area is increased. When the object confirmation features are consistent with the execution object data, the target category search interval is compressed to the target attribute range corresponding to the execution object data, and the node probability of the node branch corresponding to the target attribute range is increased; When the semantic correction feature contains new conditional trigger words, the new conditional trigger words are matched with the trigger categories in the task template table, the original trigger categories in the conditionally ambiguous task nodes are replaced, and the trigger timing search range and the handling intensity search range are reconstructed accordingly. The node probabilities of each branch after correction are normalized so that the sum of the node probabilities under the same conditional fuzzy task node remains one. Write the corrected multidimensional parameter search range and node probabilities back into the intent probability task graph.
9. A method for arranging equipment simulation training tasks for command and control coordination according to claim 8, characterized in that, When the distribution entropy value does not fall below the preset convergence threshold within a preset number of steps, a minimum clarification data generation step is also included: Read the node probability of each conditional fuzzy task node in the intent probability task graph, and calculate the distribution entropy value corresponding to each conditional fuzzy task node; Select the conditionally ambiguous task node with the highest distribution entropy value as the task node to be clarified. From the multi-dimensional parameter search range corresponding to the task node to be clarified, extract the parameter dimension with the largest range as the dimension to be clarified; Based on the task type and dimension to be clarified of the task node to be clarified, select the corresponding clarification text fragment from the task template table to generate the minimum clarification data; The system receives the clarification result data returned for the minimum clarification data, converts the clarification result data into a limited range of the dimension to be clarified, and recalculates the node probability after correcting the multidimensional parameter search interval using the limited range.
10. A method for arranging equipment simulation training tasks for command and control coordination according to claim 9, characterized in that, Extract the task path with the highest node probability and convert it into an equipment-level task operation sequence, including: In the intention probability task graph, the node branch with the highest node probability is selected step by step from the starting task node until the ending task node is reached, thus forming the target task path. Read the execution object data, task parameter data, and node connection relationships corresponding to each task node in the target task path; The execution object data, task parameter data, and node connection relationships are mapped according to the operation fields that can be identified by the simulated equipment entity to generate an initial equipment-level task operation sequence. The initial equipment-level mission operation sequence is compared with the already executed trial mission instruction set, duplicate mission operations are deleted, and the latest entity state data generated by the trial mission instruction set is retained. Based on the latest entity status data, correct the location, timing, and resource fields in the remaining task operations to obtain and output the equipment-level task operation sequence.