Command decision agent construction system and method, electronic equipment and storage medium

By constructing a command and decision-making intelligent agent system, utilizing heterogeneous computing power collaborative scheduling and cognitive enhancement knowledge system, optimizing model architecture and autonomous planning, the system solves the problems of insufficient intelligence and model generalization ability in traditional command and decision-making, and achieves efficient and accurate command and decision support.

CN121457579APending Publication Date: 2026-02-03INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511414189.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Traditional command and decision-making relies on the subjective experience of commanders and lacks intelligence, resulting in inconsistent simulation results. Existing generalized large models lack the ability to generalize and understand the business in command and decision-making operations, making it difficult to support practical applications.

Method used

A command and decision-making intelligent agent system is constructed, including a computing power scheduling component, a cognitive enhancement component, a large model training and promotion component, and an intelligent agent autonomous planning component. Through heterogeneous computing power collaborative scheduling, environmental coupling modeling, and knowledge graphs, the model architecture is optimized, and the construction and optimization of the command and decision-making intelligent agent are realized by combining autonomous planning and knowledge transfer.

Benefits of technology

It improves the generalization ability and business understanding ability of the command and decision-making intelligent agent, realizes efficient, accurate and intelligent command and decision-making, supports intelligent management of the whole process, and solves the problems of insufficient model generalization ability and insufficient business understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121457579A_ABST
    Figure CN121457579A_ABST
Patent Text Reader

Abstract

The invention provides a command decision-making agent construction system and method, electronic equipment and a storage medium. The command decision-making agent construction system comprises a computing power scheduling component, a command decision-making agent construction component, a command decision-making agent construction component and a command decision-making agent construction component, the cognitive enhancement component is used for constructing a cognitive enhancement knowledge system; the large model training and pushing assembly is used for carrying out model training based on a cognitive enhanced knowledge system to obtain a command decision intelligent agent or carrying out command decision reasoning based on the command decision intelligent agent under the condition that the computational power constraint requirement of command decision is met; the agent autonomous planning component is used for optimizing the command decision agent based on reasoning experience data accumulated in the command decision reasoning process of the command decision agent; and the agent platform application component is used for constructing an agent platform with autonomous perception and decision execution capabilities, so that the command decision agent realizes command decision reasoning and command decision execution based on the agent platform. And good application of artificial intelligence in command decision business is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a command decision-making agent construction system and method, an electronic device, and a storage medium. BACKGROUND

[0002] Traditional command decision-making relies on the subjective experience of commanders, lacks intelligence, and has a certain sketchiness. In the deduction process, the deduction personnel need to conduct subjective analysis and judgment according to the specific situation, and design specific actions. The same set of plans may lead to different deduction results when conducted by different personnel, making it difficult to accurately reflect the problems existing in the command decision-making itself and supporting practical applications.

[0003] Artificial intelligence is accelerating its penetration into various industries, and determining command decisions based on artificial intelligence has become the core direction of future development. However, the fields that require the formulation of command decisions have strong business professional characteristics, and the existing general-purpose large models generally have insufficient model generalization ability and insufficient business understanding ability, making it difficult to be practically applied in command decision-making businesses. SUMMARY

[0004] The present application provides a command decision-making agent construction system and method, an electronic device, and a storage medium to solve the problem of poor practical application of artificial intelligence in command decision-making businesses.

[0005] The present application provides a command decision-making agent construction system, comprising: A computing power scheduling component for constructing a heterogeneous computing power collaborative scheduling system to provide computing power support for the construction of a command decision-making agent; A cognitive enhancement component for constructing a cognitive enhancement knowledge system, wherein the cognitive enhancement knowledge system is obtained based on environment coupling modeling and a knowledge graph; A large model training and pushing component for model training based on the cognitive enhancement knowledge system under the constraint requirement of command decision-making computing power, or command decision-making reasoning based on the command decision-making agent; An agent autonomous planning component for optimizing the command decision-making agent based on the reasoning experience data accumulated in the command decision-making reasoning process of the command decision-making agent; An agent platform application component for constructing an agent platform with autonomous perception and decision execution capabilities to enable the command decision-making agent to implement command decision-making reasoning and execute command decisions based on the agent platform.

[0006] According to the command decision-making agent construction system provided by the present application, the heterogeneous computing power collaborative scheduling system comprises: The cloud computing center is configured with a heterogeneous acceleration architecture and a distributed training framework, is used for aggregating multi-source data, performing data processing on the multi-source data, and training a model of a command decision-making agent based on the distributed training framework; The edge computing node is deployed with a lightweight inference framework, and is used for scheme deduction based on the lightweight inference framework. The terminal device layer is used for collecting multi-source data and performing offline decision-making.

[0007] According to the command decision-making agent construction system provided by the application, the cognitive enhancement knowledge system is coupled with the mathematical modeling of multi-dimensional physical effect variables, and environmental data is collected to build an environmental influence factor library based on the environmental data, so as to realize environmental coupling modeling.

[0008] According to the command decision-making agent construction system provided by the application, the cognitive enhancement knowledge system uses the structured association capability of the knowledge graph to extract entities, entity relationships and rules, and builds a knowledge graph based on the extracted entities, entity relationships and rules, so that the large model training component trains a model based on the knowledge graph and prior knowledge.

[0009] According to the command decision-making agent construction system provided by the application, the large model training component trains a model based on the cognitive enhancement knowledge system under the constraint requirement of the computing power of the command decision-making, to obtain a command decision-making agent, or to make a command decision-making inference based on the command decision-making agent.

[0010] According to the command decision-making agent construction system provided by the application, the optimization model architecture is a network structure designed by an automatic search algorithm; The compressed parameter scale is obtained by integrating prior knowledge into the model architecture or loss function of the command decision-making agent, and using hierarchical sensitivity analysis to identify and optimize redundant parameters in the command decision-making agent; The training and inference integrated optimization is to use mixed precision training and gradient scaling technology in the training process of the command decision-making agent, and to introduce sparse regularization in the training process, and to optimize the calculation graph acceleration matrix operation according to the input data characteristics of the command decision-making agent in the inference stage.

[0011] According to the command decision-making agent construction system provided by the application, the agent autonomous planning component uses an agent decision optimization system, a hierarchical multi-objective reinforcement learning framework and a domain knowledge updating engine to accumulate inference experience data accumulated in the command decision-making agent during the command decision-making inference process. The agent decision optimization system adopts a deep deterministic policy gradient algorithm, an experience replay buffer is designed, so that reasoning experience data is stored in the experience replay buffer, and during the command decision deduction process of the command decision agent, decision making is performed based on a Monte Carlo tree search algorithm. In the hierarchical multi-objective reinforcement learning framework, the decision task to be executed is decomposed into multiple levels, and a corresponding reward function is set for each level, and the command decision agent is optimized based on the reward function. The domain knowledge updating engine is used for knowledge transfer learning based on meta-learning technology.

[0012] According to the command decision agent construction system provided by the application, the domain knowledge updating engine uses meta-learning technology to establish a similarity evaluation system, determines the relevance of different command scenarios based on the similarity evaluation system, extracts the command decision under the target command scenario based on the relevance, and realizes knowledge transfer learning.

[0013] According to the command decision agent construction system provided by the application, the agent platform application component comprises: The intelligent command decision engine is used for command decision reasoning based on the command decision agent. The performance evaluation engine is used for performance evaluation and performance prediction of the command decision. The data flywheel engine is used for correcting the command decision that does not achieve the target effect, and optimizing the command decision agent based on the corrected command decision.

[0014] The application also provides a command decision agent construction method applied to the command decision agent construction system described in any one of the preceding aspects, and the method comprises: A cognitive enhancement knowledge system and a heterogeneous computing power collaborative scheduling system are constructed, wherein the cognitive enhancement knowledge system is obtained based on environment coupling modeling and a knowledge graph; Under the computing power constraint requirement of the command decision, model training is performed based on the cognitive enhancement knowledge system and the heterogeneous computing power collaborative scheduling system, and the command decision agent is obtained. An agent platform with autonomous perception and decision execution capability is constructed, so that the command decision agent can realize command decision reasoning and execute the command decision based on the agent platform.

[0015] The application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the command decision agent construction method described in any one of the preceding aspects when executing the computer program.

[0016] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the command decision-making agent construction method.

[0017] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the command decision-making agent construction method.

[0018] The command decision-making agent construction system, method, electronic device and storage medium provided by the application, the computing power scheduling component constructs a heterogeneous computing power collaborative scheduling system, and the heterogeneous computing power collaborative scheduling system provides computing power support for the construction of the command decision-making agent. The cognitive enhancement component constructs a cognitive enhancement knowledge system, the cognitive enhancement knowledge system is obtained based on environment coupling modeling and a knowledge graph, and the cognitive enhancement knowledge system provides interpretable, traceable and evolvable knowledge support for the command decision-making agent, so that the command decision-making agent has stronger generalization ability. The large model training and pushing component performs model training based on the cognitive enhancement knowledge system under the condition that the computing power constraint requirement of the command decision-making is met, and obtains the command decision-making agent. The model training based on the large model training and pushing component considers the precision and practicability of the model and meets the computing power constraint requirement of real-time command. The trained command decision-making agent can perform command decision-making reasoning to obtain a command decision. The agent autonomous planning component supports the command decision-making agent to accumulate reasoning experience data in the command decision-making reasoning process, so as to optimize the command decision-making agent based on the reasoning experience data and improve the accuracy of the command decision-making agent. The agent platform application component constructs an agent platform having autonomous perception and decision execution capability, and the command decision-making agent performs command decision-making reasoning and executes the command decision based on the agent platform. The cognitive enhancement knowledge system and the heterogeneous computing power collaborative scheduling system constructed by the application solve the problems of insufficient model generalization ability and insufficient business understanding ability of current general large models, effectively cope with the challenges of insufficient command decision-making efficiency and difficult improvement of intelligent level, and enable artificial intelligence to be well applied in command decision-making business. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0020] Figure 1 is one of the schematic diagrams of the command decision-making agent construction system provided by the application.

[0021] Figure 2Fig. 2 is a schematic diagram of a command decision-making intelligent agent construction system according to an example embodiment of the present application.

[0022] Figure 3 Fig. 3 is a schematic diagram of a command decision-making intelligent agent construction method according to an example embodiment of the present application.

[0023] Figure 4 Fig. 4 is a schematic diagram of an electronic device according to an example embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0025] Figure 1 Fig. 1 is a schematic diagram of a command decision-making intelligent agent construction system according to an example embodiment of the present application. As shown in Fig. 1, in an example embodiment, the command decision-making intelligent agent construction system comprises: Figure 1 a computing power scheduling component 110, configured to construct a heterogeneous computing power collaborative scheduling system to provide computing power support for construction of the command decision-making intelligent agent; a cognitive enhancement component 120, configured to construct a cognitive enhancement knowledge system, wherein the cognitive enhancement knowledge system is obtained based on environment coupling modeling and a knowledge graph; a large model training and pushing component 130, configured to perform model training based on the cognitive enhancement knowledge system under the computing power constraint requirement of command decision-making, to obtain the command decision-making intelligent agent, or to perform command decision-making reasoning based on the command decision-making intelligent agent; an intelligent agent autonomous planning component 140, configured to optimize the command decision-making intelligent agent based on reasoning experience data accumulated in the process of command decision-making reasoning of the command decision-making intelligent agent; an intelligent agent platform application component 150, configured to construct an intelligent agent platform with autonomous perception and decision execution capability, so that the command decision-making intelligent agent realizes command decision-making reasoning and execution of command decision-making based on the intelligent agent platform.

[0026] ​In the embodiment of the present application, the computing power scheduling component constructs a heterogeneous computing power collaborative scheduling system, which provides computing power support for the construction of the command and decision-making agent. The computing power scheduling component, by constructing a heterogeneous computing power collaborative scheduling system based on a "cloud-edge-end" architecture, establishes an elastic and expandable resource pool, an intelligent task distribution mechanism and a differentiated service quality guarantee strategy, realizes the full-link computing power optimization configuration from data acquisition to command and decision-making generation, and provides computing power resource basis for upper-layer large models and agents. Through intelligent scheduling of computing power, the collaborative scheduling optimization of resources and applications is realized, and the inference efficiency is improved, such as differentiated computing power adaptation for performance prediction, model training, real-time deduction and other scenes, meeting the multi-element demand of computing power accuracy and speed in different business scenarios.

[0027] The cognitive enhancement component constructs a cognitive enhancement knowledge system, which is obtained based on environment coupling modeling and knowledge graph. The cognitive enhancement component, by collecting, processing and fusing multi-source heterogeneous data in the field, provides high-quality data sources for knowledge base construction and large model training, constructs a command and decision-making special knowledge base based on processed field data, adopts a "physical-information-cognitive" ternary space coupling mechanism, constructs a cognitive enhancement knowledge system based on "environment coupling modeling + knowledge graph", and provides explainable, traceable and evolving knowledge support for the command and decision-making agent, aiming to improve the learning process of the command and decision-making agent, so that the command and decision-making agent can more efficiently, stably and deeply extract knowledge from data, learn rules, and obtain stronger generalization ability.

[0028] The large model training component, based on the cognitive enhancement knowledge system, trains the model to obtain the command and decision-making agent, or performs command and decision-making reasoning based on the command and decision-making agent under the computing power constraint requirement of command and decision-making. The large model training component, combined with the scene characteristics of command and decision-making, optimizes the model architecture, compresses the parameter size, and integrates training and prediction to realize model training based on the cognitive enhancement knowledge system, and obtain the command and decision-making agent. Based on the large model training component, the model is trained, considering the accuracy and practicality of the model, and meeting the computing power constraint requirement of real-time command. The trained command and decision-making agent can perform command and decision-making reasoning to obtain the command and decision-making. Command and decision-making refers to analyzing, evaluating and selecting from multiple possible action plans in an uncertain and dynamic environment to achieve a specific goal, and finally forming an action decision and instruction plan.

[0029] The agent autonomous planning component is used to support the command and decision-making agent to accumulate reasoning experience data in the command and decision-making reasoning process, so as to optimize the command and decision-making agent based on the reasoning experience data and improve the accuracy of the command and decision-making agent.

[0030] The agent platform application component constructs an agent platform with autonomous perception and decision execution capability, and a command and decision agent implements command and decision reasoning and executes command and decision based on the agent platform.

[0031] In the embodiment of the present application, the computing power network, large model, agent and other technologies are integrated with the command scene, the gap between the general large model and the large model in the professional field of command decision is filled, the problems of insufficient model generalization ability and insufficient business understanding ability of the current general large model are solved, the challenges of insufficient command decision efficiency and difficult improvement of intelligent level are effectively responded, and finally the construction of the command decision professional large model capability and the agent platform is realized, supporting the whole process of intelligentization of "pre-event scheme formulation and recommendation, in-event scheme execution and tracking, and post-event effect evaluation and optimization".

[0032] In an exemplary embodiment of the present application, the heterogeneous computing power collaborative scheduling system comprises: The cloud computing power hub is configured with a heterogeneous acceleration architecture and a distributed training framework, is used for aggregating multi-source data, performing data processing on the multi-source data, and training a command and decision agent based on the distributed training framework; The edge computing power node is deployed with a lightweight inference framework, and is used for scheme deduction based on the lightweight inference framework; The terminal device layer is used for collecting multi-source data and executing offline decision.

[0033] As shown in the embodiment of the present application, Figure 2 The heterogeneous computing power collaborative scheduling system specifically comprises a cloud computing power hub, an edge computing power node and a terminal device layer.

[0034] The cloud computing power hub is deployed in a strategic data center, adopts a heterogeneous acceleration architecture of CPU+GPU+NPU, is equipped with a distributed training framework (such as PaddlePaddle, TensorFlow), aggregates data in different regions through federated learning, fuses multi-source intelligence such as satellite and radar, is responsible for model training and global data processing, processes PB (PetaByte) level structured / unstructured data, supports multi-dimensional feature retrieval and similarity matching. Global data processing refers to data processing of data in the construction and reasoning process of the command and decision agent. Multi-dimensional feature retrieval refers to retrieval based on multi-dimensional features.

[0035] The edge computing power node is deployed with a lightweight inference framework (such as TensorRT, ONNX Runtime), ensures <100ms decision instruction generation delay, has real-time response capability, and undertakes performance evaluation and rapid scheme deduction.

[0036] Terminal device layer, including terminal devices such as command vehicles and unmanned aerial vehicle clusters on the terminal side, mainly supporting data collection and backhaul, offline decision-making, and the like, for multi-target tracking and path planning in an offline environment, real-time image processing and feature extraction, and running lightweight decision tree models.

[0037] In an example embodiment of the present application, the cognitive enhancement knowledge system performs mathematical modeling and coupled calculation on multi-dimensional physical effect variables, collects environmental data, and builds an environmental impact factor library based on the environmental data to realize environmental coupled modeling.

[0038] In the embodiment of the present application, the environmental coupled modeling is realized through a fusion modeling of “physical mechanism + data driving”, a dynamic environmental impact factor library is built, and full-link causal reasoning of “physical effect - environmental impact - effectiveness evaluation” is realized to meet the demand for high-confidence data for real combat command and decision-making in complex scenarios.

[0039] Specifically, the cognitive enhancement knowledge system considers multi-dimensional physical effect variables, performs mathematical modeling and coupled calculation on the multi-dimensional physical effect variables, and realizes unified representation of different dimensional parameters through tensor analysis method. The multi-dimensional physical effect variables refer to physical variables involved in command and decision-making reasoning of a command and decision-making agent, such as electromagnetic interference, meteorological parameters, target material, and the like, when the command and decision-making agent performs command and decision-making reasoning in a combat field.

[0040] Environmental data is collected, and an environmental impact factor library is built based on the collected environmental data. When collecting environmental data, a distributed updating mechanism based on federated learning is adopted to support real-time collection of environmental data in data sources such as satellite remote sensing, unmanned aerial vehicles, and ground sensors.

[0041] The cognitive enhancement knowledge system also includes effectiveness evaluation for command and decision-making, and a plurality of effectiveness evaluation modules are built to realize full-link causal reasoning of “physical effect - environmental impact - effectiveness evaluation”.

[0042] The performance evaluation module is constructed by a corresponding performance evaluation layer, and the performance evaluation layer includes a physical effect layer, an environment impact layer and a performance evaluation layer. The effect prediction module is used for predicting the effect that can be achieved after executing the command decision, the environment interference evaluator is used for evaluating the environmental impact factors involved in the execution of the command decision, and the performance evaluation function is used for evaluating the effect that can be achieved after executing the command decision. In an exemplary embodiment of the present application, the cognitive enhancement knowledge system uses the structured association capability of the knowledge graph to extract entities, entity relationships and rules, and constructs a knowledge graph based on the extracted entities, entity relationships and rules, so that the large model training component performs model training based on the knowledge graph and prior knowledge.

[0043] In the embodiment of the present application, the cognitive enhancement knowledge system uses the structured association capability of the knowledge graph to mine the implicit information in domain entities (such as equipment, personnel, terrain, etc.), relationships (such as tactical cooperation, command affiliation, etc.) and rules (such as production rules converted from regulations, etc.), and then constructs a knowledge graph. Through the knowledge graph, the problems of data scarcity, high noise and complex association are solved, the scale, quality and semantic richness of the data are improved, clear semantic information is provided for model training, and prior knowledge such as equipment mechanism model, tactical rules and historical cases is integrated into the model training process, solving the problem of insufficient model generalization ability under small sample conditions.

[0044] In an exemplary embodiment of the present application, the large model training component, by optimizing the model architecture, compressing the parameter scale, and training and optimizing in an integrated manner, performs model training based on the cognitive enhancement knowledge system under the requirement of satisfying the computing power constraint of command decision, to obtain a command decision intelligent agent, or performs command decision reasoning based on the command decision intelligent agent.

[0045] In the embodiment of the present application, the large model training component, combined with the scene characteristics of command decision, optimizes the model architecture, compresses the parameter scale, and trains and optimizes in an integrated manner, to realize the whole-process light-weight of the command decision intelligent agent from training to reasoning based on the cognitive enhancement knowledge system under the premise of ensuring decision accuracy, and to balance the accuracy and practicality of the model, to satisfy the computing power constraint requirement of real-time command decision.

[0046] In an exemplary embodiment of the present application, the optimized model architecture is a network structure designed by an automatic search algorithm. The compression parameter scale is achieved by incorporating prior knowledge into the model architecture or loss function of the command and decision-making agent, and by using hierarchical sensitivity analysis to identify and optimize redundant parameters in the command and decision-making agent. The integrated training and inference optimization employs hybrid precision training and gradient scaling techniques during the training process of the command and decision-making agent, introduces sparse regularization during training, and combines hardware characteristics to optimize the computational graph for the target hardware to accelerate matrix operations during the inference phase of the command and decision-making agent, and adjusts the inference path according to the characteristics of the input data of the command and decision-making agent.

[0047] In this embodiment of the invention, the optimized model architecture is an efficient and lightweight network structure designed for the command and decision-making agent through automated search algorithms (such as reinforcement learning and Bayesian optimization), which maintains feature representation ability while reducing computational load.

[0048] Compressing parameter size involves incorporating prior knowledge into the model architecture or loss function through physical rule encoding, reducing the model's learning of redundant features, and using hierarchical sensitivity analysis to identify redundant parameters, thereby improving parameter utilization efficiency.

[0049] The integrated training-inference optimization employs mixed-precision training to reduce floating-point computations during the training phase of the command and decision-making agent, combines gradient scaling techniques to avoid numerical overflow, and introduces sparsity regularization during training to automatically sparsify model weights. Furthermore, during the inference phase, it leverages hardware characteristics to accelerate matrix operations through optimized computational graphs tailored to the target hardware. Simultaneously, it dynamically adjusts the inference path based on the characteristics of the command and decision-making agent's input data, reducing redundant computations and thus saving computational resources and lowering latency.

[0050] Mixed-precision training intelligently combines FP16 (half-precision floating-point numbers) and FP32 (single-precision floating-point numbers) in different parts of model training to significantly reduce the memory required to train the model and greatly speed up the training process, while with almost no loss of the model's final accuracy.

[0051] Gradient scaling is a technique used during model training to indirectly scale gradients by scaling the loss, preventing gradients from underflowing in low-precision computations (such as FP16).

[0052] Optimize the computation graph for the target hardware to run as fast and resource-efficiently as possible on the target hardware (such as CPU, GPU, NPU) to accelerate matrix operations.

[0053] In an exemplary embodiment of the present invention, the agent autonomous planning component utilizes an agent decision optimization system, a hierarchical multi-objective reinforcement learning framework, and a domain knowledge update engine to enable the reasoning experience data accumulated during the command and decision-making process of the agent to make command and decision-making inference.

[0054] In this embodiment of the invention, the intelligent agent autonomous planning component, through the intelligent agent decision optimization system, hierarchical multi-objective reinforcement learning framework, and domain knowledge update engine, supports the command and decision intelligent agent to autonomously accumulate reasoning experience data in the process of actual command and decision reasoning. Based on the reasoning experience data, the evaluation rules and decision strategies are dynamically updated, forming a closed-loop improvement system of "training-actual combat-evolution" for the command and decision intelligent agent, and continuously iterating and optimizing.

[0055] In an exemplary embodiment of the present invention, the agent decision optimization system adopts a deep deterministic policy gradient algorithm, designs an experience replay buffer so that inference experience data is stored in the experience replay buffer, and makes decisions based on the Monte Carlo tree search algorithm during the command decision-making agent's command decision deduction process. In the hierarchical multi-objective reinforcement learning framework, the decision-making task to be executed is decomposed into multiple levels, and a corresponding reward function is set for each level. The command and decision-making agent is optimized based on the reward function. Domain knowledge update engine, used for knowledge transfer learning based on meta-learning technology.

[0056] In this embodiment of the invention, the agent decision optimization system adopts the Deep Deterministic Policy Gradient (DDPG) algorithm framework and designs a dynamic experience replay buffer to achieve temporal storage of inference experience data. Secondly, during the command decision-making agent's command decision-making simulation process, Monte Carlo Tree Search (MCTS) is introduced for multi-step look-ahead decision-making. Through offline policy learning, historical command decision data is efficiently utilized. Then, the evaluation rules adopt a dual network structure (Target Network + Online Network), and the decision policy update frequency is set to a dynamic adjustment mechanism of the first level (24 hours / time) and the second level (4 hours / time).

[0057] In the hierarchical multi-objective reinforcement learning framework, the decision-making task to be executed is decomposed into multiple levels of "planning-deployment-execution". Each level is set with a corresponding reward function. Through reinforcement learning and knowledge-driven autonomous learning, the command and decision-making agent can autonomously discover new command and decision-making patterns and optimize decision-making strategies. At the same time, during the autonomous evolution process, the human intervention mechanism is retained in the loop, which effectively solves the credibility problem of autonomous decision-making system.

[0058] Specifically, the decision-making tasks to be executed are first broken down into multiple levels of "planning-deployment-execution". The planning level uses Value Iteration Network (VIN) to process the high-dimensional state space and is responsible for handling macro-level objectives and resource allocation. The deployment level uses Hierarchical Abstract Markov Decision Process (HAM) to generate combat sequences and organize personnel. The execution level applies the Option Framework to achieve real-time decision-making and control the real-time action decisions of personnel and equipment.

[0059] Then, a corresponding reward function is set for each level. Through reinforcement learning and knowledge-driven autonomous learning, the command and decision-making agent can guide decision-making patterns and optimize decision-making strategies. Before calling up capabilities, the reward model is used to estimate the reward for the planned capability execution result. If the estimated reward score exceeds a pre-set threshold, the capability call-up process begins. If the estimated reward score is lower than the threshold, it indicates that the current planning may not produce the desired effect, and the process reverts to the capability planning stage for replanning or human intervention. At the same time, the results of this analysis and planning are saved to the memory module for future reference and learning.

[0060] Meanwhile, during the autonomous decision-making process, the human intervention mechanism is retained, allowing personnel to monitor and correct the command decisions made by the intelligent command agent in real time. By integrating the computational advantages of machine learning with the experience and judgment of personnel, the credibility problem of autonomous decision-making systems can be effectively solved.

[0061] In an exemplary embodiment of the present invention, the domain knowledge update engine utilizes meta-learning technology to establish a similarity evaluation system, determines the correlation between different command scenarios based on the similarity evaluation system, extracts command decisions under the target command scenario based on the correlation, and realizes knowledge transfer learning.

[0062] In this embodiment of the invention, the domain knowledge update engine integrates professional knowledge such as equipment performance parameters and combat experience into the learning process of the command and decision-making agent. When faced with a new command scenario, the command and decision-making agent uses meta-learning technology to quickly extract command decisions under the target scenario similar to the new command scenario, thereby realizing knowledge transfer to meet the ever-changing scenario training and decision-making needs.

[0063] Specifically, firstly, a model-agnostic meta-learning (MAML) algorithm is used to establish a transfer learning foundation model. Secondly, a scene feature encoder is designed to transform command scenarios into computable feature vectors. Then, considering factors such as task type, environmental complexity, terrain matching degree, and time constraints, a similarity evaluation system is established. This system quantifies the correlation between different command scenarios. Based on this correlation, historical similar target command scenarios are retrieved in real time, enabling rapid extraction of command decisions under similar scenarios and achieving knowledge transfer.

[0064] In an exemplary embodiment of the present invention, the intelligent agent platform application component includes: The intelligent command and decision engine is used to perform command and decision reasoning based on the command and decision intelligent agent; The performance evaluation engine is used to evaluate and predict the effectiveness of command decisions. The data flywheel engine is used to correct command decisions that fail to achieve the target effect, and to optimize the command decision-making agent based on the corrected command decisions.

[0065] In this embodiment of the invention, the intelligent agent platform application component, by constructing an intelligent agent platform with autonomous perception and decision-making execution capabilities, provides an end-to-end capability set including model training, knowledge acquisition and processing, intelligent agent development, operation monitoring, and operational analysis. This effectively improves the efficiency of constructing and continuously iterating command and decision-making intelligent agents, supports dynamic decision-making in complex scenarios such as decision simulation and command and dispatch, meets constantly changing decision-making needs, and realizes the construction of a smart management system covering the entire process from "pre-event to in-event to post-event." The intelligent agent platform application component includes an intelligent command and decision-making engine, a performance evaluation engine, and a data flywheel engine.

[0066] The effectiveness assessment engine, combining actual command performance data, models the command environment, constructing electromagnetic interference models, meteorological impact models, and combat effectiveness models that closely resemble the actual command environment. This enables high-confidence assessments of the effectiveness of command decisions, providing more accurate effectiveness data support for applications such as scheme formulation and simulation. Based on this, it performs rapid and high-confidence effectiveness assessments of the command decision schemes formulated by commanders, using equipment matching, and predicts the combat effectiveness of the current schemes, assisting commanders in identifying shortcomings and quickly generating more cost-effective solutions.

[0067] The intelligent command and decision-making engine performs real-time analysis, fusion, correlation, and comprehensive processing of massive amounts of discrete intelligence information in complex environments to generate comprehensive intelligence data. Then, through simulation exercises and situational awareness exercises, it collects and summarizes massive amounts of data from the simulation process to obtain information on troop and firepower deployment, target selection, strike effects, and battle losses, providing commanders with analysis and prediction of the effects of different command decisions. Simultaneously, based on optimization algorithms, it intelligently plans action options at different stages of command decision-making and proposes optimization and adjustment suggestions for command decisions, providing decision support for commanders to optimize and revise their command decisions, thereby improving the scientific nature and accuracy of command decisions.

[0068] In practical applications, the data flywheel engine records command decisions that have poor performance, i.e., those that have failed to achieve the target results. This includes errors identified by the intelligent system and errors actively marked by personnel. This data is then fed back to the knowledge base, either online or offline. Commanders will intervene with this knowledge and collaborate with the large model to correct these erroneous data, forming high-density knowledge. Finally, this knowledge is applied to the next version of the command decision-making agent or to the next round of iterative training, thereby continuously improving the capabilities of the command decision-making agent.

[0069] In this invention, computing power networks, large-scale models, and intelligent agents are integrated with command scenarios to construct a professional command and decision-making intelligent agent, realizing an intelligent management system covering the entire process from "pre-event to in-event to post-event." By bridging the gap between general-purpose large-scale models and large-scale models in the command and decision-making field, this invention addresses the common problems of insufficient model generalization ability and insufficient business understanding in current general-purpose large-scale models, effectively responding to challenges such as insufficient command and decision-making efficiency and difficulty in improving the level of intelligence. Furthermore, it effectively solves the challenges currently faced in the command and decision-making field, such as the difficulty in coordinating basic computing network resources and the challenges of adapting current industry large-scale models to domestic computing power and lightweight deployment. Specifically, it includes computing power scheduling components, cognitive enhancement components, large-scale model training and promotion components, intelligent agent autonomous planning components, and intelligent agent platform application components. Among them, the computing power scheduling component achieves coordinated scheduling optimization of resources and applications through intelligent scheduling methods, such as differentiated computing power adaptation for scenarios like performance simulation, model training, and real-time inference, to meet the diverse needs of different business scenarios for computing power accuracy and speed; the cognitive enhancement component constructs a cognitive enhancement knowledge system based on "environment-coupled modeling + knowledge graph" to provide knowledge support for command and decision-making intelligent agents; and the large model training and inference component, combined with the characteristics of command and decision-making scenarios, optimizes the model architecture, compresses parameter scale, and integrates training and inference optimization to achieve the transition from training to inference while ensuring decision-making accuracy. The entire process is lightweight, balancing model accuracy and practicality while meeting the computational constraints of real-time command. The agent autonomous planning component, through the agent decision optimization system, hierarchical multi-objective reinforcement learning framework, and domain knowledge update engine, supports the command decision agent to autonomously accumulate reasoning experience data in realistic simulations, forming a closed-loop improvement system of "training-real combat-evolution". The agent platform application component, by building an agent platform with autonomous perception and decision execution capabilities, supports dynamic decision-making in complex scenarios such as command decision simulations and command dispatch, to meet the ever-changing scenario training and decision-making needs.

[0070] The method for constructing a command and decision-making intelligent agent provided by this invention is described below. The method described below can be referred to in correspondence with the system described above. It should be noted that the method provided in the embodiments below and the system provided in the embodiments above belong to the same concept, and the specific execution method of each step has been described in detail in the system embodiments, and will not be repeated here.

[0071] In one exemplary embodiment of the present invention, please refer to Figure 3 , Figure 3 This is an exemplary embodiment illustrating a method for constructing a command and decision-making intelligent agent, applied to any of the preceding command and decision-making intelligent agent construction systems, the method comprising: Step 310: Construct a cognitive enhancement knowledge system and a heterogeneous computing power collaborative scheduling system; wherein, the cognitive enhancement knowledge system is obtained based on environment coupling modeling and knowledge graph.

[0072] In this embodiment of the invention, a cognitively enhanced knowledge system and a heterogeneous computing power collaborative scheduling system are constructed. The cognitively enhanced knowledge system, based on "environmental coupling modeling + knowledge graph," provides knowledge support for the command and decision-making intelligent agent. The heterogeneous computing power collaborative scheduling system, through intelligent scheduling methods, achieves collaborative scheduling optimization of resources and applications, meeting the diverse needs of different business scenarios for computing power accuracy and speed.

[0073] Step 320: Under the condition of meeting the computing power constraints of command and decision-making, the model is trained based on the cognitive enhancement knowledge system and the heterogeneous computing power collaborative scheduling system to obtain the command and decision-making intelligent agent.

[0074] In this embodiment of the invention, by optimizing the model architecture, compressing the parameter scale, and integrating training and promotion optimization, the model is trained in conjunction with a cognitive enhancement knowledge system to obtain a command and decision-making intelligent agent, taking into account the characteristics of command and decision-making scenarios.

[0075] Step 330: Construct an intelligent agent platform with autonomous perception and decision-making execution capabilities, so that the command and decision-making intelligent agent can realize command and decision-making reasoning and execute command and decision-making based on the intelligent agent platform.

[0076] In this embodiment of the invention, an intelligent agent platform with autonomous perception and decision-making execution capabilities is constructed. The command and decision-making intelligent agent realizes command and decision-making reasoning and executes command and decision-making based on the intelligent agent platform.

[0077] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a command and decision-making intelligent agent construction method, which includes: constructing a cognitive enhancement knowledge system and a heterogeneous computing power collaborative scheduling system; wherein the cognitive enhancement knowledge system is obtained based on environmental coupling modeling and knowledge graphs; Under the constraint of computing power requirements for command and decision-making, a command and decision-making intelligent agent is obtained by training the model based on a cognitive augmentation knowledge system and a heterogeneous computing power collaborative scheduling system. Construct an intelligent agent platform with autonomous perception and decision-making execution capabilities, so that command and decision-making intelligent agents can realize command and decision-making reasoning and execute command decisions based on the intelligent agent platform.

[0078] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0079] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the command and decision-making intelligent agent construction method provided by the above methods. The method includes: constructing a cognitive enhancement knowledge system and a heterogeneous computing power collaborative scheduling system; wherein the cognitive enhancement knowledge system is obtained based on environmental coupling modeling and knowledge graph. Under the constraint of computing power requirements for command and decision-making, a command and decision-making intelligent agent is obtained by training the model based on a cognitive augmentation knowledge system and a heterogeneous computing power collaborative scheduling system. Construct an intelligent agent platform with autonomous perception and decision-making execution capabilities, so that command and decision-making intelligent agents can realize command and decision-making reasoning and execute command decisions based on the intelligent agent platform.

[0080] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the command and decision-making intelligent agent construction method provided by the above methods, the method comprising: constructing a cognitively enhanced knowledge system and a heterogeneous computing power collaborative scheduling system; wherein, the cognitively enhanced knowledge system is obtained based on environmental coupling modeling and knowledge graph; Under the constraint of computing power requirements for command and decision-making, a command and decision-making intelligent agent is obtained by training the model based on a cognitive augmentation knowledge system and a heterogeneous computing power collaborative scheduling system. Construct an intelligent agent platform with autonomous perception and decision-making execution capabilities, so that command and decision-making intelligent agents can realize command and decision-making reasoning and execute command decisions based on the intelligent agent platform.

[0081] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A command and decision-making intelligent agent construction system, characterized in that, include: The computing power scheduling component is used to build a heterogeneous computing power collaborative scheduling system to provide computing power support for the construction of command and decision-making intelligent agents; A cognitive enhancement component is used to construct a cognitive enhancement knowledge system; wherein, the cognitive enhancement knowledge system is obtained based on environment coupling modeling and knowledge graph; The large model training and inference component is used to train models based on a cognitive enhancement knowledge system to obtain a command and decision intelligent agent, or to perform command and decision reasoning based on the command and decision intelligent agent, under the condition of meeting the computing power constraints of command and decision. The intelligent agent autonomous planning component is used to optimize the command and decision-making intelligent agent based on the reasoning experience data accumulated during the command and decision-making reasoning process; The intelligent agent platform application component is used to build an intelligent agent platform with autonomous perception and decision-making execution capabilities, so that command and decision-making intelligent agents can realize command and decision-making reasoning and execute command and decision-making based on the intelligent agent platform.

2. The command and decision-making intelligent agent construction system according to claim 1, characterized in that, The heterogeneous computing power collaborative scheduling system includes: The cloud computing power hub is equipped with a heterogeneous acceleration architecture and a distributed training framework, which is used to aggregate multi-source data, process multi-source data, and train models for command and decision-making intelligent agents based on the distributed training framework. Edge computing nodes are equipped with a lightweight inference framework for performing solution simulations based on the lightweight inference framework. The terminal device layer is used to collect multi-source data and perform offline decisions.

3. The command and decision-making intelligent agent construction system according to claim 1, characterized in that, The cognitive enhancement knowledge system achieves environmental coupling modeling by mathematically modeling and coupling calculations of multi-dimensional physical effect variables, collecting environmental data, and constructing an environmental impact factor library based on the environmental data.

4. The command and decision-making intelligent agent construction system according to claim 1, characterized in that, The large model training and inference component optimizes the model architecture, compresses the parameter scale, and integrates training and inference. Under the condition of meeting the computing power constraints of command and decision-making, it trains the model based on the cognitive enhancement knowledge system to obtain the command and decision-making intelligent agent, or performs command and decision-making reasoning based on the command and decision-making intelligent agent.

5. The command and decision-making intelligent agent construction system according to claim 4, characterized in that, The optimized model architecture is a network structure designed using an automated search algorithm; The compression parameter scale is achieved by incorporating prior knowledge into the model architecture or loss function of the command and decision-making agent, and by using hierarchical sensitivity analysis to identify and optimize redundant parameters in the command and decision-making agent. The integrated training and inference optimization employs hybrid precision training and gradient scaling techniques during the training process of the command and decision-making agent, introduces sparse regularization during training, and combines hardware characteristics to optimize the computational graph for the target hardware to accelerate matrix operations during the inference phase of the command and decision-making agent, and adjusts the inference path according to the characteristics of the input data of the command and decision-making agent.

6. The command and decision-making intelligent agent construction system according to claim 1, characterized in that, The intelligent agent autonomous planning component utilizes an intelligent agent decision optimization system, a hierarchical multi-objective reinforcement learning framework, and a domain knowledge update engine to leverage the reasoning experience data accumulated during the command and decision-making process of the intelligent agent. Among them, the agent decision optimization system adopts a deep deterministic policy gradient algorithm, designs an experience replay buffer so that the reasoning experience data is stored in the experience replay buffer, and makes decisions based on the Monte Carlo tree search algorithm during the command decision intelligent agent’s command decision deduction process. In the hierarchical multi-objective reinforcement learning framework, the decision-making task to be executed is decomposed into multiple levels, and a corresponding reward function is set for each level. The command and decision-making agent is optimized based on the reward function. Domain knowledge update engine, used for knowledge transfer learning based on meta-learning technology.

7. The command and decision-making intelligent agent construction system according to any one of claims 1 to 6, characterized in that, The intelligent agent platform application components include: The intelligent command and decision engine is used to perform command and decision reasoning based on the command and decision intelligent agent; The performance evaluation engine is used to evaluate and predict the effectiveness of command decisions. The data flywheel engine is used to correct command decisions that fail to achieve the target effect, and to optimize the command decision-making agent based on the corrected command decisions.

8. A method for constructing a command and decision-making intelligent agent, characterized in that, The method, applied to the command and decision-making intelligent agent construction system according to any one of claims 1 to 7, comprises: A cognitive enhancement knowledge system and a heterogeneous computing power collaborative scheduling system are constructed; wherein, the cognitive enhancement knowledge system is obtained based on environment coupling modeling and knowledge graph; Under the constraint of computing power requirements for command and decision-making, a command and decision-making intelligent agent is obtained by training the model based on a cognitive augmentation knowledge system and a heterogeneous computing power collaborative scheduling system. Construct an intelligent agent platform with autonomous perception and decision-making execution capabilities, so that command and decision-making intelligent agents can realize command and decision-making reasoning and execute command decisions based on the intelligent agent platform.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the command and decision-making intelligent agent construction method as described in claim 8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the command and decision-making intelligent agent construction method as described in claim 8.