Human-machine hybrid decision system based on multiple agents

By using a human-machine hybrid decision-making system based on multi-agent systems, the problems of information overload and lack of human intelligence in autonomous agents under traditional models are solved. This system enables accurate, flexible and reliable decision support in complex scenarios, improving the collaborative efficiency and scientific nature of decision-making.

CN121541478BActive Publication Date: 2026-03-27NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In decision support for complex scenarios, existing technologies often fall into information overload due to traditional experience-driven models, making it difficult to respond quickly. Meanwhile, fully autonomous intelligent agent models lack the global perspective and creative judgment of humans, resulting in a lack of flexibility in the scheduling and control of intelligent agents and a lack of unified standards for interaction between system modules, thus failing to fully leverage the advantages of the technology.

Method used

A human-machine hybrid decision-making system based on multi-agent technology is adopted, including a simulation and inference platform, an agent framework, a standardized interaction interface, and a dynamic scheduling unit. Through the collaborative operation of situation analysis agents and action decision-making agents, combined with the human-machine collaboration unit, data transmission and joint decision-making are realized, which solves the module barriers and control conflicts, and improves the collaborative efficiency of the system and the scientificity and reliability of the decision-making.

Benefits of technology

It achieves precise decision support in complex scenarios, taking into account both forward-looking and dynamic adaptability, and improves the rationality and reliability of decision-making. By leveraging the advantages of machine processing of massive amounts of information and the deep integration of human wisdom, it enhances the flexibility and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541478B_ABST
    Figure CN121541478B_ABST
Patent Text Reader

Abstract

The application relates to a multi-agent-based man-machine hybrid decision system. The system comprises: a simulation deduction platform, which outputs full-quantity situation data and updated situation data containing multi-domain entity states; an agent framework, which deploys at least one group of confrontation agents, each group of agents comprising a situation analysis agent and an action decision agent; a standardized interaction interface, which realizes data transmission between each group of agents, the simulation deduction platform and a man-machine cooperation unit; a dynamic scheduling unit, which independently controls each group of agents to manage the whole life cycle and resolves the control right conflict of multiple groups of agents on the same entity through a last instruction priority rule; and a man-machine cooperation unit, which receives user parameter control instructions for each group of agents, transmits the instructions to the agent framework through the standardized interaction interface, and outputs structured analysis results and decision schemes generated by the two types of agents to the user. The method can improve the rationality and reliability of decision making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a man-machine hybrid decision system based on multiple agents. BACKGROUND

[0002] With the rapid development of the field of reasoning and intelligent collaborative support, the continuous upgrading of computer technology and hardware computing environment promotes the deep integration of modeling simulation, big data and artificial intelligence technology, giving birth to a new generation of intelligent reasoning tools. The collaborative application of these technologies has freed reasoning from the limitations of traditional sand table games, with the core characteristics of data-driven and intelligent assistance, and can simulate complex scenarios of multi-domain collaboration to provide quantitative support and scenario verification for decision making. It has become an important means of decision support in modern complex environments.

[0003] Current reasoning mainly relies on two core modes: one is the traditional experience-oriented mode relying on human global judgment, scene intuition and practical experience to make decisions; the other is the completely autonomous agent decision-making mode, which completes the situation analysis and action planning independently through algorithm models. However, these two modes and existing related solutions still have obvious limitations: the traditional experience-oriented mode is prone to information overload in complex scenarios with multi-domain interweaving and information explosion, making it difficult to achieve rapid response and unable to fully utilize the advantages of efficient computing brought by technology; the completely autonomous agent mode lacks the unique global perspective, creative judgment and complex situation coping ability of humans, making it difficult to adapt to non-quantitative variables and dynamic confrontation needs. At the same time, both modes fail to achieve the deep integration of human wisdom and machine intelligence, and the scheduling and control of agents lack flexibility, and the interaction between modules lacks a unified standard, resulting in the inability to complement each other's advantages, ultimately affecting the scientificity, flexibility and actual scene fitting degree of reasoning and decision making. SUMMARY

[0004] Therefore, it is necessary to provide a man-machine hybrid decision system based on multiple agents to solve the above technical problems.

[0005] A man-machine hybrid decision system based on multiple agents, the system comprising:

[0006] a simulation reasoning platform for building a reasoning environment, outputting full-quantity situation data and updated situation data containing multi-domain entity states, and executing decision instructions issued by agents;

[0007] An agent framework is deployed with at least one set of adversarial agents, each set of agents including a situation analysis agent and an action decision agent, the action decision agent including a pre-planning agent and a real-time planning agent; the situation analysis agent receives full-scope situation data and updated situation data and processes to generate structured analysis results, the action decision agent generates decision schemes and decision instructions according to the structured analysis results, and transmits the decision instructions to a simulation deduction platform through a standardized interaction interface;

[0008] A standardized interaction interface is used to realize data transmission between each set of agents, the simulation deduction platform and the man-machine collaborative unit;

[0009] A dynamic scheduling unit independently controls each set of agents to manage the whole life cycle, and resolves the conflict of control right of multiple sets of agents on the same entity through a pre-set last instruction priority rule;

[0010] A man-machine collaborative unit receives user parameter control instructions for each set of agents and transmits them to the agent framework through the standardized interaction interface, and outputs the structured analysis results and decision schemes generated by the two types of agents to the user, assisting in realizing man-machine joint decision and decision confirmation.

[0011] The above-mentioned multi-agent-based man-machine hybrid decision system can provide accurate and comprehensive data support for decision-making through the simulation deduction platform outputting full-scope situation data and updated situation data containing multi-domain entity states, and the structured processing of the situation analysis agent in the agent framework, fully utilizing the advantages of machine in efficiently processing massive information. Through the cooperative operation of the pre-planning agent and the real-time planning agent in the agent framework, the forward-looking and dynamic adaptability of decision-making can be considered, and the decision-making needs of complex scenarios can be adapted. Through the standardized interaction interface, efficient data transmission between each set of agents, the simulation deduction platform and the man-machine collaborative unit can be realized, which can break down the barriers between system modules and improve overall collaborative efficiency. Through the dynamic scheduling unit for independent control of the whole life cycle of the agent and the control right conflict resolution mechanism, the stable operation of the system can be guaranteed, and the flexibility of agent resource scheduling can be enhanced. Through the man-machine collaborative unit, user parameter control and decision-making can be realized, which can deeply integrate human wisdom and machine intelligence, making the decision-making both scientifically based and practical, significantly improving the rationality and reliability of decision-making. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 It is a structural diagram of the deduction adversarial framework in one embodiment;

[0013] Figure 2 It is a task setting structure diagram in one embodiment;

[0014] Figure 3Fig. 1 is a schematic diagram of a command hierarchy under the control of a unit order in one embodiment;

[0015] Figure 4 Fig. 2 is a schematic diagram of an adversarial scheduling framework in one embodiment;

[0016] Figure 5 Fig. 3 is a schematic diagram of a human-machine hybrid collaborative decision timing in one embodiment;

[0017] Figure 6 Fig. 4 is a schematic diagram of an agent decision framework in one embodiment;

[0018] Figure 7 Fig. 5 is a schematic diagram of an agent application mode flow in one embodiment;

[0019] Figure 8 Fig. 6 is a schematic diagram of an agent running state flow in one embodiment;

[0020] Figure 9 Fig. 7 is a schematic diagram of a single agent application timing in one embodiment;

[0021] Figure 10 Fig. 8 is a schematic diagram of an agent configuration parameter flow in one embodiment;

[0022] Figure 11 Fig. 9 is a timing diagram when a client application applies multiple agents in one embodiment. DETAILED DESCRIPTION

[0023] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0024] In one embodiment, a multi-agent based human-machine hybrid decision system is provided, comprising:

[0025] a simulation deduction platform, configured to build a deduction environment, output full-quantity situation data and update situation data containing multi-domain entity states, and execute decision instructions issued by the agent;

[0026] an agent framework, configured to deploy at least one group of adversarial agents, each group of agents including a situation analysis agent and an action decision agent, the action decision agent including a pre-planning agent and a real-time planning agent; the situation analysis agent receives the full-quantity situation data and the update situation data and processes and generates structured analysis results, the action decision agent generates decision schemes and decision instructions according to the structured analysis results, and transmits the decision instructions to the simulation deduction platform through a standardized interaction interface;

[0027] A standardized interaction interface is used to realize data transmission between each group of intelligent agents, the simulation deduction platform and the man-machine collaborative unit.

[0028] A dynamic scheduling unit is used to independently manage and control each group of intelligent agents to manage the whole life cycle, and to resolve the control right conflict of multiple groups of intelligent agents on the same entity through a pre-set last instruction priority rule.

[0029] A man-machine collaborative unit is used to receive parameter regulation instructions of each group of intelligent agents and transmit them to the intelligent agent framework through the standardized interaction interface, and output the structured analysis results and decision schemes generated by the two types of intelligent agents to the user, thereby assisting in realizing man-machine joint decision and decision confirmation.

[0030] In the above-mentioned multi-agent-based man-machine hybrid decision system, the simulation deduction platform outputs whole quantity situation data and updated situation data containing the state of multi-domain entities, which, in combination with the structured processing of the situation analysis intelligent agent in the intelligent agent framework, can provide accurate and comprehensive data support for decision-making, and fully exert the advantage of machine in efficiently processing massive information. Through the collaborative operation of the pre-planned intelligent agent and the real-time planned intelligent agent in the intelligent agent framework, the forward-looking and dynamic adaptability of decision-making can be taken into account, and the decision-making requirements of complex scenarios can be adapted. Through the standardized interaction interface, efficient data transmission between each group of intelligent agents, the simulation deduction platform and the man-machine collaborative unit can be realized, which can break through the barriers between system modules and improve the overall collaborative efficiency. Through the dynamic scheduling unit, the whole life cycle of the intelligent agent is independently managed and controlled, and the control right conflict resolution mechanism is realized, which can ensure the stable operation of the system and enhance the flexibility of intelligent agent resource scheduling. Through the man-machine collaborative unit, user parameter regulation and decision-making joint decision can be realized, which can deeply integrate human wisdom and machine intelligence, so that the decision-making has both scientific basis and actual demand, and the rationality and reliability of the decision-making are significantly improved.

[0031] In one embodiment, the standardized interaction interface includes: a platform operation interface used to obtain whole quantity situation data and updated situation data, and to realize the construction and deletion of tasks, grouping, auxiliary reference points and special areas; a task setting interface used to configure patrol tasks, control tasks and support tasks, including a public interface of task enabling, starting and stopping, entity allocation and deletion; an entity control interface used to realize transmission of fine control instructions of multi-domain entities; the fine control instructions include maneuvering, target selection, equipment control and path planning instructions; a rule modification interface used to modify control rules according to content classification, the content classification including overall setting, equipment control setting, equipment use rule and deployment adjustment rule.

[0032] In this embodiment, as shown in FIG. 2, the standardized interaction interface includes a platform operation interface, a task setting interface, an entity control interface and a rule modification interface. Figure 1The illustrated deduction confrontation framework diagram. This framework breaks the limitations of traditional deduction system module wall barriers and low interaction efficiency, through standardized design and fine division, not only guarantees the independent operation ability of each module, but also realizes the efficient flow of data and control instructions, lays a solid foundation for the deep cooperation of subsequent man-machine mixed decision. The basic deduction platform is a war game deduction engine, which can be controlled through a visual operation interface or a program language, and can run independently, support human-computer confrontation and human-human confrontation; confrontation scheduling is the core of the whole framework, including deduction scheduling and multi-process control modules, as Figure 4 The agent interface interaction can encapsulate each operation according to the function of the simulation deduction engine and the constraints of human-human confrontation, and encapsulate the interface into four parts: platform operation, task setting, entity control and order modification.

[0033] The platform operation interface is the core channel for the agent to obtain environmental information and control the basic configuration of the platform, including two core functions: situation acquisition and platform-level operation. As the source of information for agent decision-making, situation acquisition provides structured multi-domain entity information, covering both sides of the situation data: one side of the situation contains the state information of entities such as air, water surface and mobile devices, and the other side of the situation is divided into entity information and statistical information, entity information covers various types of operating units and equipment, and statistical information includes operating effect, resource consumption and other data, providing quantitative reference for agent decision-making; Platform-level operation supports agents or users to realize task creation and deletion, operating unit grouping management, auxiliary reference point calibration, special area division and other basic configuration operations through the interface, meeting the individual adjustment needs of the deduction scene.

[0034] The task setting interface focuses on the standardized configuration of various operating tasks, such as Figure 2 The task setting structure diagram as shown, the interface development of patrol task, control task and support task has been completed, there are a total of 6 interfaces, which are common operation interfaces under the task setting interface, and their function is to provide full life cycle operation functions for patrol tasks, control tasks and support tasks. Different types of tasks have built-in exclusive operating logic, which can drive entities to run autonomously without additional control: patrol task supports automatic interception configuration for different types of targets, can select the corresponding patrol mode according to the operating requirements, and realizes the target monitoring and interception in the specified area; Control task supports autonomous operation on objects in the target list, allows configuration of accompanying support units to improve the safety and effectiveness of task execution; Support task supports auxiliary operations such as electronic reconnaissance, energy supply and signal interference, and the task entity will carry out cruise support according to the preset path. In addition, the interface also provides common operation functions for various tasks, including task activation, start and end time setting, operating entity allocation and deletion, etc., realizing the standardized control of the whole life cycle of the task.

[0035] The entity control class interface is used to implement fine-grained control of multi-domain operation units, and adapts to more than 1,000 different types of entities (including air operation units, surface carriers, underwater devices, ground facilities, satellites, etc.). The control instructions supported by this interface cover core operations such as operation unit maneuver path planning, target selection and locking, device switching and parameter adjustment, and navigation route optimization. The agent can output specific control instructions through this interface to achieve precise regulation and control of entity operation behavior, meeting the strategy execution requirements in complex scenarios.

[0036] The order modification interface is used to implement multi-level and multi-dimensional operation rule configuration and adjustment, such as the command level hierarchy diagram under order control as shown in Figure 3 The rules at each level follow the principles of default inheritance and independent modification: low-level rules automatically follow the rules of the upper level when not modified, and are independently effective when modified, allowing differences with the upper level rules. According to the content, it can be divided into overall setting, device control setting, device usage rule and deployment adjustment rule, which correspond to global operation specification, electronic device operation control, device operation parameter configuration, operation unit retreat and redeployment condition setting, etc. To simplify the management of hierarchical rules, the interface also provides a rule reset function that can quickly restore the rules of a specified level to the default state and re-subordinate to the upper level rules, improving the flexibility and efficiency of rule configuration.

[0037] As shown in Figure 4 The upper multi-process control module relies on the resource pool, agent pool, and scenario pool of the reasoning platform to achieve flexible scheduling of resources and instances. It can manage multiple reasoning processes in parallel, allocate corresponding agents, reasoning scenarios, and platform resources to each process, and support multi-scenario parallel reasoning. The lower reasoning scheduling module is the core of the process control. Starting from "reasoning start", it successively completes agent initialization (corresponding to the _init_0 operation in the figure), agent reset (corresponding to the reset(situation) operation), and then drives the reasoning to proceed. During the process, the execution agents of different parties complete the situation time step processing (corresponding to the step(situation, time) operation) and action instruction generation through the rule module and learning module, respectively. At the same time, it is judged whether this round is over through is_done(). If this round is over and the reasoning termination condition is not met, the agent reset and reasoning promotion are executed in a loop until the reasoning is completely finished, realizing the automatic closed-loop control of the reasoning process.

[0038] The architecture eliminates compatibility problems between modules through standardized interface encapsulation, reduces the difficulty of agent development and integration, and improves the efficiency of agent training and verification through the collaborative design of multi-process control and deduction scheduling. The multi-mode control and independent running ability of the basic deduction platform adapt to the use needs of different users (ordinary users and developers). The hierarchical rule configuration and fine-grained entity control improve the authenticity and flexibility of the deduction scene, and provide stable and efficient bottom support for the subsequent agent collaboration and human-machine interaction in human-machine hybrid decision-making.

[0039] In one embodiment, the system further comprises: a human-machine collaboration unit receiving a rule modification instruction issued by a user, the instruction containing a control level and rule content; the control level includes a deduction side level, a task level, a grouping level, or an individual level; the agent framework analyzes the rule modification instruction, judges the subordinate relationship of the target control level and the upper level rule; if it is a low-level rule modification, the rule configuration of the corresponding agent is updated according to the instruction content, and it no longer obeys the upper level rule; when receiving a rule reset instruction, the target level rule is restored to the default state and re-submits to the upper level rule control.

[0040] After building an efficient and stable interaction between the agent and the deduction engine, the core proposition further focuses on how to build an organic collaborative system architecture - deeply integrate the intelligent agents with different functions and scattered deployment and human decision-making wisdom to form an efficient collaborative decision-making capability. The human-machine hybrid decision-making system architecture of the present invention is the result of careful design following the internal rules of the OODA cycle, which takes the scientific classification and precise positioning of intelligent agents as the core cornerstone, realizes resource planning through a clear and clear overall framework, adapts to complex and variable deduction environments relying on real-time and efficient dynamic scheduling mechanism, ensures seamless collaboration between users and machine agents, and realizes the organic unity of forward-looking planning and response to unexpected scenarios.

[0041] The efficiency of the human-machine hybrid decision-making framework fundamentally depends on the functional design and division of labor of the internal agent ecosystem. Human-machine hybrid decision-making cannot adopt a single agent design, but according to the internal logic of modern command and decision-making, the agent is scientifically divided into two categories: situational analysis agents and action decision agents, the latter is further divided into pre-planning and real-time planning modes. This classification method aims to accurately correspond to the core needs of commanders in the judgment and decision-making dimensions, so that various types of agents can perform their respective functions and deeply collaborate to jointly extend and enhance the cognitive ability of commanders.

[0042] Among them, the situation analysis type agent plays the role of intelligent perception and cognitive center. Its core mission is to transform the massive and raw bottom layer data provided by the engine into structured knowledge and deep insight, greatly expanding the commander's situational understanding (Situational Understanding) ability. This type of agent is not a simple data filter, but a specialist analyst using advanced artificial intelligence algorithms such as deep learning and time series prediction models. In terms of function, they can achieve intent recognition of non-self resource deployment patterns, infer their possible operational purposes and action plans, quickly identify the model / ship type of ambiguous information obtained by reconnaissance units, reduce the uncertainty of identification between the two opposing sides, and quantitatively analyze the operational capabilities of the two opposing sides, calculate the threat level, operational window and defense resilience, and provide data support for decision-making. In terms of method, the output is not a dry data report, but through high-lighting, labeling, trend lines, probability pop-up windows, and other highly visualized ways, seamlessly integrated into the main interface of the deduction, directly empowering the commander's judgment process, enabling them to form a faster and more accurate understanding of the confrontation environment.

[0043] The action decision type agent plays the role of intelligent staff and executive officer. They directly serve the decision-making and action process of the commander, responsible for generating feasible action plans and implementing them. According to the difference between the deduction stage and the decision-making timeliness, this type of agent is divided into two complementary modes: pre-planning agent and real-time planning agent.

[0044] In one embodiment, the pre-planning agent is used to generate decision-making plans before deduction or during the intermittent period. In this embodiment, the pre-planning agent acts before the deduction starts or during the intermittent period of task stage conversion. It is mainly used to solve complex optimization problems such as resource deployment, task planning, and resource allocation. It usually uses operations research algorithms such as linear programming and genetic algorithms or game search models based on a large number of self-play training, can search for multiple high-quality action plans in a wide range of scheme space, and perform deduction evaluation through war game simulation, finally present several optimal schemes and their expected gains and losses to the commander for final decision, greatly improving the scientificity and efficiency of planning.

[0045] In one embodiment, the real-time planning agent is used to generate decision-making solutions in the deduction process under sudden scenarios. In this embodiment, the real-time planning agent acts in the deduction process. It is mainly used to deal with sudden situations, instantaneous operation opportunities or unplanned events, requiring the generation of response instructions within seconds or even milliseconds. Its technical implementation is often more lightweight, often using rule-based expert systems, fast heuristic search or reactive reinforcement learning models. The commander or scheduling framework can dynamically allocate execution units (such as a standby operation unit) and target units (such as a suddenly appearing non-self target) to it, which will then generate specific and executable instructions such as interception, avoidance, etc. After the commander confirms it with one key, it is quickly sent to the simulation engine, thereby realizing the real-time AI-assisted decision-making in intense confrontation.

[0046] Through the fine division of the functions and timing of the agent, a solid foundation is laid for the realization of truly human-machine hybrid intelligence. The effectiveness of human-machine hybrid decision-making not only depends on agents with different functions, but also depends on a system architecture that can integrate these elements organically and regulate their coordination methods. This framework adopts a layered decoupling, service-based system architecture, whose core design philosophy is to separate human judgment, machine calculation, and system scheduling. Through the definition of clear interfaces and protocols, flexible, reliable, and extensible human-machine collaboration is achieved.

[0047] As shown in the human-machine hybrid collaborative decision-making timing diagram in Figure 5 , the human-machine hybrid collaborative decision-making timing describes the interaction and working order of humans, machines, and systems in a complete OODA cycle, which defines the workflow of hybrid decision-making.

[0048] In the human-machine hybrid decision-making mode, humans can import developed agents through the deduction client, edit the parameters required by the agent such as units, targets, or location points, and run the agent. At this time, the deduction is created by the deduction client, the scenario is loaded, the situation is received from the deduction server, and the situation is sent to the agent application framework. The agent application framework goes through the same situation processing process as the development framework, sends the situation to the agent, makes the same decision as the development mode, and sends the decision instruction to the client, which is sent to the deduction server by the client.

[0049] (1) Observation & Orientation stage: The war game engine generates new deduction scenario state data. The scheduling framework automatically or according to human instructions triggers the situation analysis type agent. This type of agent processes the original data and pushes the analysis results (such as intent, threat, identification conclusion) to the human-machine interaction interface in a visual form to assist the commander in forming a cognitive advantage.

[0050] (2) Decision stage: the commander makes a decision strategy based on the current situation awareness. This decision can be: macro instruction, triggering pre-planned agent to generate scheme and simulation evaluation; micro instruction, triggering real-time planning agent.

[0051] (3) Action and feedback stage: the scheme or instruction generated by the agent (such as specific resource mobilization list, interception route) is issued to the war game engine through the interface, thereby driving the evolution of the situation change and starting the next cycle.

[0052] This process clearly establishes the hybrid decision-making paradigm of machine perception, human-machine joint judgment, human decision-making, machine generation, and human approval, ensuring that humans are always at the core of the decision-making loop.

[0053] As Figure 6 The agent decision framework schematic diagram is shown, the agent running state is from the life cycle perspective of a single agent, which describes the internal state transition process from being called to returning the result. This process is the cornerstone of the dynamic management of the scheduling framework. The operation mechanism of the agent in the human-machine hybrid decision-making is analyzed from three aspects of the decision framework, application mode, and running state of the agent. The situation received by the agent development framework can be divided into full-amount situation (a full-amount frame will be sent once at the beginning of the evolution) and updated situation. The development framework creates an agent object according to the agent class, and assigns values to the unit and task attributes of the agent object according to the evolution side where the agent is located. After receiving the updated situation, the agent attribute values are updated again. Ensure that the agent object can easily obtain the task, unit, and other situation attributes.

[0054] In one embodiment, the system adopts a server-client distributed architecture, and the server and the client establish a long connection through a standardized interaction interface. The simulation evolution platform and the dynamic scheduling unit are deployed on the server side. The man-machine cooperation unit and the agent framework are deployed on the client side. The agent framework of the client side realizes the bidirectional transmission of instructions and data with the simulation evolution platform and the dynamic scheduling unit of the server side through the standardized interaction interface.

[0055] In this embodiment, the agent application mode process schematic diagram is shown in Figure 7 The war game client imports the use of the agent, dynamically adjusts the configuration parameters of the agent, and can reflect the human-machine cooperation. Humans can import the developed agent, add entity unit parameters and target parameters.

[0056] The intelligent agent makes decisions according to expert rules or operational optimization algorithms, or machine learning intelligence, and so on, and the human and machine cooperate to complete the control target or other operation tasks. The function implemented by the intelligent agent can be to use the configured entity unit to immediately prepare to operate the target, use what kind of operation equipment, how far the distance is, and use how many operation equipment to control the target; the human makes decisions through the client about which unit to control, when to control which target, and the specific control process is implemented by which intelligent agent. In combination with the intelligent agent countermeasures scheduling framework, the intelligent agent running state flow is as shown in Figure 8 This intelligent agent running state flow is a prerequisite for realizing the hot plug and dynamic scheduling of the intelligent agent. It abstracts each intelligent agent as an independent, stateless service unit, and its life cycle is completely managed by the scheduling framework. This makes any intelligent agent can be started, stopped, updated or replaced without affecting the overall operation of the system.

[0057] In an embodiment, receiving user parameter control instructions for each group of intelligent agents and transmitting to the intelligent agent framework through a standardized interaction interface includes: the human-machine cooperation unit of the client receives the user input parameter control instructions through a visual interaction interface; the parameter control instructions include a control entity list, target object parameters and operation preference parameters; the human-machine cooperation unit performs format specification checking and permission range verification on the parameter control instructions, transmits the parameter control instructions that pass the verification to the intelligent agent framework on the server side through a standardized interaction interface, and the interface real-time feedback transmission state; after the intelligent agent framework receives the instructions, it updates the configuration parameters of the corresponding intelligent agent and synchronously adjusts the decision logic and operation priority, and feeds back the parameter update result to the human-machine cooperation unit of the client through the standardized interaction interface.

[0058] Specifically, the human-machine cooperation unit realizes the deep interaction between the user and the intelligent agent through the intelligent agent scheduling function, and the core includes two major functional modules of intelligent agent editing and intelligent agent message output, which are implemented as follows:

[0059] The agent scheduling mainly includes two parts of agent editing and agent message output functions. The agent editing can realize the management function of the agent, mainly including agent import, agent running, agent element addition (entity unit, target, reference point), agent element display, agent configuration saving, etc. On the left side of the agent editing interface is the function interface of the agent, which can realize adding, starting and deleting the agent. There are two ways to add the agent: importing a single agent and importing the agent configuration. The single agent is imported by clicking + at the bottom left of the agent editing interface, and the agent file to be imported can be selected after clicking +. The agent editing function supports saving the current agent scheme (saving the current agent and the current agent element), and the saved configuration can be imported by clicking the import agent configuration function. The agent message output function can realize five kinds of message display (pure text, key information, circular area, arrow, and polygon area). In the agent scheduling, we provide the agent message output interface for the user. After the user develops the agent, the user can realize the display of the five types of messages through the agent message output interface. When the agent calls the message output interface, the agent message output text message will be displayed on the client interface.

[0060] In one embodiment, the dynamic scheduling unit independently controls each group of agents to manage the whole life cycle, including: when the agent is loaded, the dynamic scheduling unit assigns a unique identifier to the agent, registers the corresponding capability element information to the service directory, sets the agent state to idle, synchronizes the operable multi-domain entity list of the agent to the simulation and deduction platform, and completes the permission initialization; when the preset trigger condition is met, the dynamic scheduling unit sends a task request to the agent, switches the agent state to running, and notifies the simulation and deduction platform to push the full-amount situation data and update situation data to the agent; when receiving the pause instruction, the dynamic scheduling unit marks the agent state as paused and stops distributing new tasks, and sends a pause notification to the simulation and deduction platform to pause the situation data pushing; when receiving the uninstall instruction, the dynamic scheduling unit unregisters the agent identifier, releases the occupied computing resources, switches the agent state to terminated, and sends a control right backflow notification to the simulation and deduction platform to update the entity control right record; when the dynamic scheduling unit controls the agent to pause or uninstall, the multi-domain entity controlled by the agent continues to execute the last valid instruction until the dynamic scheduling unit allocates a new control right or issues a new decision instruction, ensuring the continuity of the deduction process.

[0061] In this embodiment, the human-machine hybrid decision-making framework must have the ability to dynamically manage agents during runtime to adapt to complex and variable deduction environments. The dynamic scheduling and hot plug mechanism of the present invention is the technical core of realizing this flexibility. This mechanism ensures that a single agent can be updated, debugged, loaded, or unloaded independently without interrupting the entire deduction system, while effectively coordinating access conflicts of multiple agents to the same deduction resource, thereby ensuring the continuity of deduction and the reliability of the system.

[0062] The life cycle of a single agent is managed by the scheduling framework, and the smooth switching of its run-pause-end states is the basis of hot plug. The application timing diagram when a client imports an agent is as shown in Figure 9 .

[0063] When the user decides to pause the agent's operation based on the current situation, the entire agent operation can be paused; or some units or targets need to be removed from the agent (some units no longer have suitable control conditions, etc., or the target has escaped, etc.). Before the execution of each step of the decision-making function of the agent, the updated configuration parameters are obtained after requesting the parameters configured by the client, and the decision-making process is updated. The specific timing is as follows:

[0064] Run timing: the user selects an agent program (such as a Python script or model file) through the client interface, and the client uploads it to the scheduling framework. The framework assigns it a unique ID and registers its capability meta-information (such as controllable entity types, supported tasks) in the service directory, and sets the state to idle (Idle). When the trigger condition is met, the scheduling framework sends a task request to the agent. Before calculation, the agent will first call an updated real-time configuration parameter through the client interface (such as the list of controlled entities for this decision, specified target objects, strategy preferences, etc., as shown in Figure 10 ). This ensures that the decision is based on the latest user intentions and deduction scene situation, and realizes real-time fine-tuning of human behavior to agents. The agent combines real-time parameters and global situation for calculation, generates decision-making instructions (such as entity A moving towards target B), and returns the result to the framework and issues it to the engine for execution.

[0065] Pause timing: The user issues a pause instruction through the client, and the scheduling framework marks the agent state as paused and immediately stops dispatching new decision tasks to it. In this phase, the situation processing module remains active, continues to receive and process broadcast global situation information to maintain the timeliness of its internal situation image, and is ready to resume operation at any time. The decision control module is suspended and no longer generates new decision instructions. All entities controlled by the agent will strictly follow the last valid instruction issued before it was paused to continue their actions until the instruction is executed or a new instruction is received. For example, if the last instruction is to fly to point X, the entity will continue to execute the patrol command, showing an automatic driving-like behavior, until the agent is awakened or taken over by another agent.

[0066] Stop / Kill timing: The user issues a stop instruction, and the scheduling framework sends a termination signal to the agent, which is then unregistered from the service directory. The agent's situation processing and decision modules are completely suspended, releasing all computing resources. Similar to the pause state, the entities it controls will continue to execute their last instruction to ensure the continuity and physical rationality of the reasoning operation behavior, avoiding the entity from being stiff due to the agent's exit. The control is automatically released back to the scheduling framework, allowing the user to reassign.

[0067] In one embodiment, the conflict of control right of multiple groups of agents on the same entity is resolved by a pre-set last instruction priority rule, which includes: when an agent generates a decision instruction for a target entity, adding its own identifier and sending timestamp to the decision instruction, and sending it to the standardized interaction interface; the dynamic scheduling unit synchronously acquires the decision instruction through the standardized interaction interface, and checks the permission matching degree of the sending agent on the target entity according to the pre-constructed entity control right table; the entity control right table includes a multi-domain entity unique identifier, a current control right attribution agent identifier, and an instruction update timestamp field; when the check is passed, the entity control right table is queried, if the current control right of the target entity belongs to other agents, the sending timestamp of the current decision instruction is compared with the instruction update timestamp, if the current sending timestamp is later, the entity control right table is updated, the control right of the target entity is switched to the current sending agent, and the instruction update timestamp is replaced synchronously; the standardized interaction interface issues the entity control right to the updated decision instruction to the simulation reasoning platform, and the simulation reasoning platform sends the updated decision instruction to the target entity, realizing the conflict resolution of the control right of multiple agents on the same entity.

[0068] Compared with single-agent scheduling, when multiple agents are enabled at the same time, and their control ranges have intersections, the conflict of control right on the same entity is inevitable. The application timing of multiple agents is as shown in Figure 11 .

[0069] The framework adopts the following mechanisms for resolution:

[0070] Last instruction priority principle: The scheduling framework maintains a physical control table, which records in real time the current control of each entity by an agent. When an agent tries to send an instruction to an entity, no matter who controls the entity previously, the framework will accept and execute the latest received instruction. The agent sending the instruction will automatically obtain the control of the entity and update the control table.

[0071] Timing logic and conflict resolution: This mechanism means that the control of the entity is dynamic and competitive. The later activated agent can immediately override the previous agent's instruction to the same entity. This not only gives high-level agents the ability to override low-level agent instructions, realizing flexible cross-level command, but also allows users to immediately start another agent for intervention and correction when they find that an agent has made a mistake, thus ensuring the ultimate control of humans over agents in mechanism.

[0072] Decision consistency guarantee: To avoid the shaking of entity behavior caused by frequent control switching, the agent should have certain coordination awareness in decision-making. Usually, the agent needs to read the current entity control state before calculation. If its target entity has been controlled by other agents, it can choose to communicate and negotiate or give up decision-making, thus avoiding the generation of invalid instructions at the algorithm level, and together with the hardware rules at the scheduling level, forming a complete conflict resolution scheme.

[0073] The dynamic hot plug and scheduling mechanism of the present application defines a clear state timing and last instruction priority conflict resolution principle, making the entire system a high-resilience and high-flexibility organism, which not only supports flexible decision-making at the strategy level, but also realizes real-time human-in-the-loop control at the system level.

[0074] In one embodiment, the agent framework deploys four types of functionally differentiated agents, corresponding to two modes of pre-planning and real-time planning, respectively, to achieve decision support in complex scenarios through algorithm optimization and process closed loop, including resource target allocation agent (pre-planning agent), air operation multi-target agent (real-time planning agent), cruise device trajectory planning agent (real-time planning agent), and early warning device cooperative planning agent (pre-planning agent). The resource target allocation agent first acquires state data and target information of multi-domain entities through the situation collection module, filters the core parameters and performs structured processing; then calls the operations research algorithm or game search model, constructs a multi-objective optimization function, and solves the optimal allocation scheme under the constraints of resource quantity, target value, etc.; after the scheme is generated, the control and management instructions are issued to the simulation and deduction platform through the standardized interaction interface; finally, the effect evaluation module is used to calculate the target control success rate and resource utilization efficiency, forming a decision-making closed loop. The air operation multi-target agent calculates the target threat matrix and resource action success rate matrix through the situation analysis module, weights the two matrices to obtain a comprehensive evaluation matrix, and sets a threshold to select high-priority targets; when the target evaluation value exceeds the threshold, the control instruction generation is automatically triggered; at the same time, the built-in operation unit guarantee rules are used, when the operation unit resource is below the preset proportion, the resource needs to be supplemented or there is no effective action capability, the return mechanism is automatically triggered to ensure task continuity. The cruise device trajectory planning agent is suitable for trajectory optimization and obstacle avoidance in dynamic environments. Users specify operation units, reference points, and other parameters through the client. The agent first determines the current task state of the operation unit, and if there is a conflict task, it is automatically canceled; then the reference point is set as the waypoint to guide the operation unit to the specified location; when the operation unit reaches the reference point and meets the launch conditions, the cruise device is triggered to launch; after successful launch, the trajectory is dynamically updated by the preset step size to avoid environmental obstacles, and at the same time, the path extension algorithm is used to ensure that multiple cruise devices arrive at the target area simultaneously, improving the task success rate. The early warning device cooperative planning agent is suitable for multi-unit cooperative detection scenarios. When the simulation is started, the cooperative strategy of the early warning device and the accompanying operation unit is started simultaneously, the accompanying operation unit constructs a protection range around the early warning device, and the patrol area is dynamically updated according to the real-time position of the early warning device; the early warning device uses the route optimization algorithm to maximize the detection range while avoiding environmental risks, drives along the preset heading at the initial stage and turns on the detection function, adjusts the driving direction according to the situation change, and uses the hovering maneuver mode when entering the key area to achieve continuous coverage detection in the key area.The functions and performances of the four types of agents are verified through a simulation environment, the resource target allocation agent accurately matches resources and targets through a multi-objective optimization algorithm and a closed-loop decision-making process, and improves the rationality of resource utilization and the pertinence of target control; the air operation multi-objective agent quickly locks high-priority targets relying on matrix weighted evaluation, reduces invalid consumption, and ensures task continuity; the cruise device trajectory planning agent improves the quality of complex task completion through dynamic obstacle avoidance and collaborative synchronization strategy; and the early warning device collaborative planning agent realizes the optimal balance between detection efficiency and operation safety, and provides stable information support for decision-making.

[0075] The four types of agents work in coordination, which verifies the core value of the present application, that is, for the first time in this type of deduction scene, the transition from static pre-programming to dynamic real-time scheduling, from a single agent to multi-functional cooperation, and from manual operation to bidirectional empowerment of man and machine is realized. This systematic solution from theory to practice not only provides a specific technical approach to improve the intelligent level of deduction and the actual application efficiency, but also aims to provide users with an intelligent staff who is online at all times, has comprehensive capabilities, and can be dispatched at any time. Ultimately, the intelligent staff empowers decision-making and moves towards a new paradigm of decision-making with human-machine cooperation.

[0076] The technical features of the above embodiments can be combined arbitrarily, and to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0077] The above-described embodiments only express several implementation manners of the present application, the description is more specific and detailed, but it should not be understood as a limitation on the scope of the application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A human-machine hybrid decision-making system based on multi-agent systems, characterized in that, The system includes: The simulation and inference platform is used to build the inference environment, output full situational data and updated situational data containing the states of entities in multiple domains, and execute decision-making instructions issued by the intelligent agent. An intelligent agent framework is provided, which deploys at least one group of adversarial intelligent agents. Each group of intelligent agents includes a situation analysis intelligent agent and an action decision-making intelligent agent. The action decision-making intelligent agent includes a pre-planning intelligent agent and a real-time planning intelligent agent. The situation analysis intelligent agent receives full situation data and updated situation data and processes them to generate structured analysis results. The action decision-making intelligent agent generates decision schemes and decision instructions based on the structured analysis results and transmits the decision instructions to the simulation and inference platform through a standardized interactive interface. A standardized interactive interface is used to enable data transmission between various intelligent agents, the simulation platform, and the human-machine collaborative unit; The dynamic scheduling unit independently manages each group of intelligent agents to manage their entire lifecycle, and resolves conflicts of control over the same entity by multiple groups of intelligent agents through a pre-set last instruction priority rule; The human-machine collaboration unit receives the user's parameter adjustment instructions for each group of intelligent agents and transmits them to the intelligent agent framework through a standardized interaction interface. At the same time, it outputs the structured analysis results and decision schemes generated collaboratively by the two types of intelligent agents to the user, assisting in realizing human-machine co-judgment and decision confirmation. The method of resolving control conflicts among multiple agents over the same entity through a pre-set last instruction priority rule includes: When an intelligent agent generates a decision instruction for a target entity, it adds its own identifier and a sending timestamp to the decision instruction before sending it to the standardized interaction interface. The dynamic scheduling unit synchronously obtains the decision instructions through the standardized interaction interface and verifies the matching degree of the sending agent's permissions to the target entity according to the pre-built entity control rights table; the entity control rights table includes a unique identifier for a multi-domain entity, the identifier of the agent currently in control, and an instruction update timestamp field. When the verification passes, the entity control table is queried. If the target entity's current control belongs to another intelligent agent, the sending timestamp of the current decision instruction is compared with the instruction update timestamp. If the current sending timestamp is later, the entity control table is updated, the control of the target entity is switched to the current sending intelligent agent, and the instruction update timestamp is replaced synchronously. The standardized interaction interface sends the updated decision instructions regarding entity control to the simulation platform, which then sends the updated decision instructions to the target entity, thereby resolving conflicts in control of the same entity among multiple agents.

2. The system according to claim 1, characterized in that, The dynamic scheduling unit independently manages each group of intelligent agents to manage the entire lifecycle, including: When the agent is loaded, the dynamic scheduling unit assigns a unique identifier to the agent, registers the corresponding capability information to the service directory, sets the agent's state to idle, and synchronizes the list of operable multi-domain entities of the agent with the simulation platform to complete the permission initialization. When the preset triggering conditions are met, the dynamic scheduling unit sends a task request to the agent, switches the agent's state to running, and notifies the simulation platform to push full situational data and updated situational data to the agent. Upon receiving a pause command, the dynamic scheduling unit marks the agent's state as paused and stops distributing new tasks, and sends a pause notification to the simulation platform to suspend the push of situational data. Upon receiving the unload command, the dynamic scheduling unit cancels the agent's identifier, releases the occupied computing resources, switches the agent's state to terminate, and sends a control return notification to the simulation platform to update the entity control record.

3. The system according to claim 1, characterized in that, When the dynamic scheduling unit manages the agent to pause or unload, the multi-domain entities controlled by the agent continue to execute the last valid instruction until the dynamic scheduling unit allocates new control or issues a new decision instruction, ensuring the continuity of the simulation process.

4. The system according to claim 1, characterized in that, The system adopts a server-client distributed architecture, with the server and client establishing a long connection through a standardized interaction interface. The simulation platform and dynamic scheduling unit are deployed on the server side; The human-machine collaboration unit and intelligent agent framework are deployed on the client side; The client-side intelligent agent framework achieves bidirectional transmission of instructions and data with the server-side simulation and deduction platform and dynamic scheduling unit through a standardized interactive interface.

5. The system according to claim 1, characterized in that, The pre-planned agent is used to generate decision-making schemes before simulations or during inter-stage intervals.

6. The system according to claim 1, characterized in that, The real-time planning agent is used to generate decision-making solutions for unexpected scenarios during the simulation process.

7. The system according to claim 1, characterized in that, The system also includes: The human-machine collaboration unit receives rule modification instructions issued by the user. The instructions include control levels and rule content. The control levels include inference level, task level, group level, or individual level. The intelligent agent framework parses rule modification instructions and determines the subordinate relationship between the target control level and the superior rule; If it is a low-level rule modification, the rule configuration of the corresponding agent will be updated according to the instruction content, and the rule of the higher level will no longer be followed; When a rule reset command is received, the target level rule is restored to its default state and reverts to being controlled by the next higher level rule.

8. The system according to claim 1, characterized in that, The standardized interaction interface includes: The platform operation interface is used to acquire and update full situational data, as well as to construct and delete tasks, groups, auxiliary reference points, and special areas. The task settings interface is used to configure patrol tasks, control tasks, and support tasks, including public interfaces for task activation, start and termination, entity assignment, and deletion. An entity control interface is used to transmit fine-grained control commands for multi-domain entities; the fine-grained control commands include maneuvering, target selection, equipment control, and path planning commands. The rule modification interface is used to modify control rules according to content categories, including overall settings, device management settings, device usage rules, and deployment adjustment rules.

9. The system according to claim 1, characterized in that, The step of receiving user parameter adjustment instructions for each group of intelligent agents and transmitting them to the intelligent agent framework through a standardized interaction interface includes: The human-machine collaboration unit on the client side receives parameter adjustment instructions input by the user through a visual interactive interface; the parameter adjustment instructions include a list of control entities, target object parameters, and job preference parameters; The human-machine collaboration unit performs format standardization and permission scope verification on the parameter control instructions, and transmits the verified parameter control instructions to the intelligent agent framework on the server through a standardized interaction interface. The interface provides real-time feedback on the transmission status. After receiving the instruction, the intelligent agent framework updates the configuration parameters of the corresponding intelligent agent and adjusts the decision-making logic and task priority in sync. It then feeds back the parameter update results to the human-machine collaboration unit on the client through a standardized interaction interface.

Citation Information

Patent Citations

  • Knowledge-driven war game deduction intelligent decision-making method

    CN113435598A

  • Man-machine intelligent game system

    CN113633994A