Unmanned equipment agent control method and device based on AI-OODA ring

By adopting an AI-OODA ring-based agent control method in unmanned equipment, combined with OR-Tools and DQN algorithms, the problem of low efficiency in identifying scene changes and path planning in group actions of unmanned equipment is solved, and more efficient and low energy consumption is achieved.

CN120194702APending Publication Date: 2025-06-24AEROSPACE TIMES FEIHONG TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328298.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the existing unmanned equipment group operations, it is difficult for artificial intelligence systems to identify scene changes in time and accurately, resulting in low efficiency in flight control and mission execution of the agent, high capabilities of airborne computers and high energy consumption.

Method used

The unmanned equipment agent control method based on the AI-OODA ring is adopted. By obtaining the task information to be executed, a data model is established, path planning is used using OR-Tools, and training is combined with DQN reinforcement algorithm to realize the optimal path planning and real-time decision-making of the agent.

Benefits of technology

It improves the accuracy and efficiency of unmanned equipment path planning, reduces the requirements for onboard computer capabilities, reduces energy consumption, and realizes real-time obstacle avoidance and target recognition of the agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120194702A_ABST
    Figure CN120194702A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart parks, in particular to an unmanned equipment agent control method and device based on an AI-OODA ring, computer equipment and a storage medium. The method comprises the steps of obtaining to-be-executed task information, and establishing a data model according to the to-be-executed task information; creating a task point index and a path planning model through OR-Tools according to the data model; and performing task path planning through the path planning model according to the task point index to obtain an optimal path for executing the task. According to the method, an OODA ring is used as a theoretical basis of capability generation, role-driven intelligent agent behaviors are used as a generation target, and an implementation mode is provided for the behavior capability of the intelligent agent from four links of perception, cognition, decision making and execution through a command and control intelligent agent capability generation technology. The information interaction technology serves as an interface of capability generation and is responsible for correctly calling information in knowledge data and driving realization and evolution of the capability of the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of unmanned aerial vehicle equipment, and in particular to an intelligent agent control method, device, computer device and storage medium for unmanned equipment based on the AI-OODA loop. Background Technique

[0002] The ant colony tactic refers to relying on information to execute mission command, forming a large combat unit through multiple small combat carriers, so as to perform precise, efficient, and all-round three-dimensional defense and attack mission forms. In the future, when hundreds of unmanned intelligent devices perform tasks in a swarm or ant colony, the coordinated execution of tasks by hundreds of unmanned intelligent devices will pose higher requirements on the decision-making control system. Each unmanned intelligent device has to complete a series of tasks such as collecting data, summarizing data, aerial formation, and executing attack instructions through cluster algorithms.

[0003] Disadvantages and deficiencies of the prior art: The group actions of unmanned equipment require the machine to have a high degree of autonomy and intelligence. When the command and control scenario suddenly changes, the current behavior must be terminated in a timely manner, and the artificial intelligence system may not be able to identify the changes in a timely and accurate manner. Although high-level autonomous intelligence is the development goal of unmanned clusters, at this stage and for a long time to come, the activities of unmanned devices still heavily rely on the remote assistance of control stations, which pose relatively high requirements on the flight control, ground control, formation control, inter-station communication, avoiding enemy detection, and command and control capabilities of control station personnel of intelligent agents. Information must be exchanged in real time between intelligent agents to ensure orderly operation, and then tasks such as task allocation and target selection are involved. The realization of these functions all requires the operation of a powerful artificial intelligence system. The higher the level of artificial intelligence, the greater the requirements for the capabilities of on-board computers and the higher the energy consumption.

[0004] How to effectively evaluate the capabilities of intelligent agents has become an urgent problem to be solved at this stage. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide an intelligent agent control method, device, computer device and storage medium for unmanned equipment based on the AI-OODA loop that can improve the accuracy of path planning.

[0006] The first aspect provides an intelligent agent control method for unmanned equipment based on the AI-OODA loop, including:

[0007] Obtain information on tasks to be executed, and establish a data model according to the information on tasks to be executed. The information on tasks to be executed includes: task target coordinates, demand quantity, intelligent agent task carrying capacity, and distance matrix;

[0008] Create a task point index and a path planning model through OR-Tools according to the data model, where the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem;

[0009] Perform task path planning through the path planning model according to the task point index to obtain the optimal path for task execution.

[0010] In one embodiment, the obtaining the information of the task to be executed and establishing a data model according to the information of the task to be executed, the information of the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix, including:

[0011] Generate task target coordinates, and there are multiple task target coordinates;

[0012] Generate the demand quantity of the task target, and generate the agent task carrying capacity according to the demand quantity;

[0013] Calculate the distances between the task target coordinates according to the task target coordinates to obtain the distance matrix.

[0014] In one embodiment, the generating the task target coordinates, and there are multiple task target coordinates, includes:

[0015] Randomly generate a set of task target positions in a two-dimensional plane, and each task target position has a fixed demand quantity.

[0016] In one embodiment, the creating a task point index and a path planning model through OR-Tools according to the data model, where the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem, includes:

[0017] Summarize the generated task target positions, the demand quantity, the agent task carrying capacity, and the distance matrix into a data model.

[0018] Manage the task point index through "RoutingIndexManager" of OR-Tools, define the constraints and objectives of the path planning problem, and pass the number of unmanned equipment, the number of tasks, and the task target positions to the path planning model.

[0019] In one embodiment, the performing task path planning through the path planning model according to the task point index to obtain the optimal path for task execution, includes:

[0020] Define through a Markov decision process:

[0021] State space: Each state consists of two parts. The first part is the equipment state, and the second part is the node state. The state information includes: the position, speed, computing load, and power of the agent.

[0022] Action space: An action is defined as selecting an equipment and a node for access. The action information includes: the motion posture, computing power allocation, and node scheduling strategy of the agent.

[0023] Transition rule: The transition rule τ transfers the previous state to the next state according to the executed action, that is;

[0024] Reward function: The reward is defined as the negative value of the maximum value. The reward is calculated by separately accumulating the travel times of multiple trips of each piece of equipment.

[0025] In one embodiment, the reward and punishment settings of the reward function include:

[0026] The agent will receive a reward when approaching the target point and a punishment when moving away from the target point;

[0027] The closer the agent is to the obstacle, the greater the punishment;

[0028] According to the angle between the coordinates of the agent in the previous five steps to the coordinates before executing the action and the coordinates before executing the action to the coordinates after executing the action, it is calculated whether the trajectory is smooth;

[0029] According to whether the agent has passed through the coordinates of the previous ten steps, it is determined whether it has taken a detour. If it has passed through the coordinates of the previous ten steps, a certain punishment will be given;

[0030] Reaching the end point gives a reward.

[0031] In one embodiment, the method for controlling an unmanned equipment agent based on the AI-OODA loop further includes:

[0032] After a large number of rounds of training through the DQN reinforcement algorithm, in the path planning route in the virtual-real fusion scenario, the unmanned equipment travels according to the planned route;

[0033] On-the-spot decision-making occurs when the unmanned equipment automatically avoids obstacles in real time and discovers a target, automatically generates the next decision task, and issues it to the corresponding equipment to execute the task.

[0034] In a second aspect, there is provided a control device for an unmanned equipment agent based on the AI-OODA loop, including:

[0035] An acquisition module, configured to acquire information about a task to be executed, and establish a data model according to the information about the task to be executed. The information about the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix;

[0036] A task point index and path planning model creation module, which is used to create a task point index and a path planning model through OR-Tools according to the data model, wherein the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem;

[0037] A task path planning module, which is used to perform task path planning through the path planning model according to the task point index to obtain an optimal path for task execution.

[0038] A third aspect provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any of the above embodiments are implemented.

[0039] A fourth aspect provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any of the above embodiments are implemented.

[0040] The above-mentioned AI-OODA loop-based unmanned equipment agent control method, device, computer device and storage medium obtain information on tasks to be executed, establish a data model according to the information on tasks to be executed; create a task point index and a path planning model through OR-Tools according to the data model; perform task path planning through the path planning model according to the task point index to obtain an optimal path for task execution. This application uses the OODA loop as the theoretical basis for ability generation, takes the agent behavior driven by roles as the generation goal, and through the command and control agent ability generation technology, provides an implementation method for the behavior ability of the agent from four links: perception, cognition, decision-making, and execution. The information interaction technology, as the interface for ability generation, is responsible for correctly retrieving information from knowledge data and driving the realization and evolution of the agent's ability. The effectiveness evaluation method of the agent uses deep learning evaluation indicators, deeply docks with the ability generation technology in the OODA link, and effectively evaluates the agent's ability by comprehensively using traditional effectiveness evaluation methods and intelligent effectiveness evaluation methods. Brief Description of the Drawings

[0041] Figure 1 It is an application scenario diagram of the AI-OODA loop-based unmanned equipment agent control method in an embodiment;

[0042] Figure 2 It is a flowchart of the AI-OODA loop-based unmanned equipment agent control method in an embodiment;

[0043] Figure 3 It is a schematic diagram of the update of the initial rescue path in an embodiment;

[0044] Figure 4 It is a structural block diagram of an intelligent agent control device for unmanned equipment based on the AI-OODA loop in an embodiment;

[0045] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0046] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] The AI-OODA loop-based intelligent agent control method provided by the present application can be applied to an application environment as Figure 1 shown. The AI-OODA loop-based intelligent agent control method models around the core capabilities of the control agent, and uses the data construction method oriented to the generation of agent capabilities as the base to support the generation of the capabilities of the command and control agent. Taking the OODA loop as the theoretical basis for capability generation and the behavior of the agent driven by the role as the generation goal, through the command and control agent capability generation technology, implementation methods are provided for the behavioral capabilities of the agent from four links: perception, cognition, decision-making, and execution. The information interaction technology, as the interface for capability generation, is responsible for correctly retrieving the information in the knowledge data and driving the realization and evolution of the agent capabilities. The effectiveness evaluation method of the agent uses deep learning evaluation indicators, deeply connects with the capability generation technology in the OODA link, and effectively evaluates the agent capabilities by comprehensively using traditional effectiveness evaluation methods and intelligent effectiveness evaluation methods.

[0048] In an embodiment, an AI-OODA loop-based intelligent agent control method for unmanned equipment is provided, including the following steps: obtaining information on a task to be executed, establishing a data model according to the information on the task to be executed, where the information on the task to be executed includes: task target coordinates, demand quantity, intelligent agent task carrying capacity, and distance matrix; creating a task point index and a path planning model through OR-Tools according to the data model, where the task point index includes: starting point and ending point, and the path planning model is used to define the constraints and objectives of the path planning problem; performing task path planning through the path planning model according to the task point index to obtain an optimal path for executing the task.

[0049] Among them, the OODA loop, that is, the Observe-Orient-Decide-Act loop, is a model used to describe the decision-making process and is widely used in military and complex system control. In the control of unmanned equipment intelligent agents, the OODA loop can help the intelligent agent quickly respond to environmental changes and optimize the task execution process. For example, in cluster electronic warfare unmanned aerial vehicles, the OODA loop is used to analyze and model the combat process, significantly improving the combat effectiveness.

[0050] Specifically, the intelligent agent control device for unmanned equipment based on the AI-OODA loop uses an AI algorithm based on reinforcement learning and combines the 3D environment modeling technologies of AirSim and UE4 to construct a virtual environment (including real machines, real vehicles, and real scenes) one-to-one for the strategy research and evolutionary training of various capabilities, obtaining the algorithm model in the intelligent agent, applying it to the real scene, acquiring the data after real-scene application (including equipment task process information), and inputting this data into the training environment to continue iterative evolution, forming a continuously evolving closed-loop process. To verify the stability of the evolution strategy of the command and control intelligent agent and the effectiveness of the evolution method, a relationship evaluation model of the effectiveness of the intelligent agent evolution strategy is used to verify the stability of the evolution strategy of the command and control intelligent agent and the effectiveness of the evolution method. Each research content complements and supports each other, forming an efficient and dynamically adaptable evolution framework of the command and control intelligent agent as shown in the figure.

[0051] In this embodiment, the present application takes the OODA loop as the theoretical basis for ability generation, takes the intelligent agent behavior driven by the role as the generation goal, and provides an implementation method for the behavior ability of the intelligent agent from four links: perception, cognition, decision-making, and execution through the command and control intelligent agent ability generation technology. The information interaction technology, as the interface for ability generation, is responsible for correctly retrieving the information in the knowledge data and driving the realization and evolution of the intelligent agent ability. The effectiveness evaluation method of the intelligent agent uses deep learning evaluation indicators, deeply connects with the ability generation technology in the OODA link, and effectively evaluates the intelligent agent ability by comprehensively using traditional effectiveness evaluation methods and intelligent effectiveness evaluation methods.

[0052] In one of the embodiments, the information of the task to be executed is obtained, and a data model is established according to the information of the task to be executed. The information of the task to be executed includes: task target coordinates, demand quantity, intelligent agent task carrying capacity, and distance matrix, including: generating task target coordinates, where there are multiple task target coordinates; generating the demand quantity of the task target, and generating the intelligent agent task carrying capacity according to the demand quantity; calculating the distances between each task target coordinate according to the task target coordinates to obtain the distance matrix.

[0053] Among them, the task target coordinates: the geographical location information of the task point. The demand quantity of the task target: the workload or resource requirement to be completed at each task point. The resource demand quantity: for example, the quantity of goods to be transported, the amount of electricity to be allocated, etc. The workload: for example, the task time or task difficulty to be completed. The intelligent agent task carrying capacity: the maximum task quantity or resource quantity that the intelligent agent can undertake. The distance matrix: the distance information between task points, used for path planning.

[0054] Specifically, the intelligent agent control device for unmanned equipment based on the AI-OODA loop generates task target coordinates, and at the same time generates the demand quantity and the task carrying capacity of the intelligent agent. The demand quantity for each task target and the maximum task carrying capacity of each intelligent agent are randomly generated to ensure the diversity and realism in the simulation environment. Calculate the Euclidean distance between each task target to generate a distance matrix for subsequent path optimization calculation.

[0055] In one embodiment, task target coordinates are generated. There are multiple task target coordinates, including: randomly generating a set of task target positions in a two-dimensional plane, and each task target position has a fixed demand quantity.

[0056] Specifically, the intelligent agent control device for unmanned equipment based on the AI-OODA loop randomly generates task target positions in a two-dimensional plane, assigns randomly generated demand quantities to each task target, and at the same time randomly generates the maximum task carrying capacity of each intelligent agent, which can ensure the diversity and realism of the simulation environment.

[0057] In this embodiment, a random number generator is used to generate the x and y coordinates of the task target within a specified range. The demand quantity of the task target represents the resources or workload required to complete each task. The demand quantity can be generated within a specified range by a random number generator to ensure diversity. The maximum task carrying capacity of the intelligent agent represents the maximum amount of tasks or resources that the intelligent agent can undertake. The carrying capacity can also be generated within a specified range by a random number generator to ensure the diversity of the simulation environment.

[0058] In one embodiment, a task point index and a path planning model are created through OR-Tools according to the data model. Among them, the task point index includes: the starting point and the ending point. The path planning model is used to define the constraints and objectives of the path planning problem, including: summarizing the generated task target positions, demand quantities, intelligent agent task carrying capacities, and distance matrix into a data model. Manage the task point index through the "RoutingIndexManager" of OR-Tools, define the constraints and objectives of the path planning problem, and pass the number of unmanned equipment, the number of tasks, and the task target positions to the path planning model.

[0059] Specifically, RoutingIndexManager is the core component in OR-Tools for managing task point indexes, and it plays an important role in the path planning problem. The intelligent agent control device for unmanned equipment based on the AI-OODA loop first summarizes the generated task target positions, demand quantities, intelligent agent task carrying capacities, and distance matrix into a data model for OR-Tools to use.

[0060] Furthermore, the intelligent agent control device for unmanned equipment based on the AI-OODA loop creates a management site index and a path planning model. The RoutingIndexManager provided by OR-Tools is used to manage the task point index. The constructor of the RoutingIndexManager requires the following parameters: the number of task points including all task points and the starting / ending points. The number of intelligent agents, that is, the number of vehicles or drones involved in the path planning. The starting point (depot) refers to the starting and ending points of the path, usually the first task point. By reasonably using the RoutingIndexManager method, the correctness and flexibility of the path planning problem can be ensured.

[0061] In one of the embodiments, the intelligent agent control device for unmanned equipment based on the AI-OODA loop performs task path planning through the path planning model according to the task point index, and obtains the optimal path for executing the task, including: being defined through a Markov decision process: a state space module, which is used to represent each state S t =(V t , X t ) ∈ S is composed of two parts. The first part is the equipment state V t , and the second part is the node state S t . The state information includes: the position, speed, computing load, and power of the intelligent agent; an action space module, which is used to define the action as selecting an equipment and a node for access. The action information includes: the motion posture, computing power allocation, and node scheduling strategy of the intelligent agent; a transfer rule module, which is used to define the transfer rule τ to transfer the previous state s to the next state s t , that is, s t+1 =(V t+1 , X t+1 )=(V t+1 , X t ) = τ(V t ); a reward function module, which is used to define the reward as the negative value of this maximum value. The reward is calculated by separately accumulating the multiple travel times of each piece of equipment.

[0062] As Figure 2 shown, specifically, in order to minimize the maximum task execution time of all equipment, the intelligent agent control device for unmanned equipment based on the AI-OODA loop defines the reward as the negative value of this maximum value. The reward is calculated by separately accumulating the multiple travel times of each piece of equipment. The state perceived by the intelligent agent at time t is s t . Under the mechanism of exploration and exploitation, the intelligent agent has a certain probability of randomly selecting the next action, or selecting an action by referring to past learning experiences, and takes the action a t, after that, the environment will give the agent the corresponding reward and punishment value rt, and then the agent will repeat the above process and continue to perceive the state s at time t+1 t+1 , in order to obtain new rewards, thereby continuously improving its own decision-making strategy.

[0063] In one embodiment, the intelligent agent control device for unmanned equipment based on the AI-OODA loop considers various factors, including the movement direction of the agent, the distance from obstacles, trajectory smoothness, and whether it goes back the same way, etc.

[0064] Specifically, there will be a reward when the agent approaches the target point, and a punishment when the agent moves away from the target point. By comparing the current distance and the next-step distance, if it approaches the target point, a positive reward will be given, and if it moves away, a negative reward will be given. Generally speaking, the punishment value is greater than the reward value to prevent the drone from hovering back and forth and stagnating. The collision probability is calculated based on the distance to the obstacle closest to the agent. The closer the distance, the greater the punishment. The closer the agent is to the obstacle, the greater the punishment. Whether the trajectory is smooth is calculated based on the angle between the coordinates of the agent in the previous five steps to the coordinates before the execution of the action and the coordinates before the execution of the action to the coordinates after the execution of the action. If the angle is too large, the trajectory is not smooth and a punishment is given. Whether it goes back the same way is determined according to whether the agent has passed through the coordinates of the previous ten steps. If it has passed through the coordinates of the previous ten steps, a certain punishment will be given. If the distance between the agent and the target point is less than the threshold, a large reward will be given.

[0065] In one embodiment, as Figure 3 and Figure 4 shown, the task path planning is carried out through the path planning model according to the task point index to obtain the optimal path for executing the task. It also includes: after a large number of rounds of training through the DQN reinforcement algorithm, the path planning route in the virtual-real fusion scenario, and the unmanned equipment travels according to the planned route. The on-the-spot decision-making occurs after the unmanned equipment automatically avoids obstacles in real time and discovers the target, automatically generates the next decision-making task, and issues it to the corresponding equipment to execute the task.

[0066] Specifically, the intelligent agent control device for unmanned equipment based on the AI-OODA loop adopts the DQN algorithm. The core of the algorithm is the Current Q network and the Target Q network. The Current Q network is responsible for interacting with the environment and completing the perception and decision-making functions of the DQN algorithm. The experience samples collected in each interaction will be stored in the experience pool for the training of the Current Q network. The Target Q network is responsible for generating the sample labels required for the training of the Current Q network.

[0067] It should be noted that in this embodiment, the calculation method of the loss value of the intelligent agent control device for unmanned equipment based on the AI-OODA loop using the DQN algorithm is:

[0068]

[0069] The mean squared error between the Q-estimate and the Q-target is calculated before each update of the deep reinforcement learning network. The Current Q value is used to evaluate the expected utility of the agent taking a specific action in a specific state. The Q-estimate is the Q value output by the Current Q network, which represents the expected return of the agent taking a certain flight action a in a given state S (representing a certain coordinate position in the point cloud map). This value evaluates the quality of the agent's actions and is calculated by the reward function. The Q-target is the Q value output by the Target Q network. The Target Q network is a copy of the Current Q network, and its parameters are updated at a slower rate during training to increase the stability of the algorithm.

[0070] In one embodiment, an intelligent agent control device for unmanned equipment based on the AI-OODA loop is provided, including:

[0071] An acquisition module, configured to acquire information about the task to be executed, and establish a data model according to the information about the task to be executed. The information about the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix;

[0072] A task point index and path planning model creation module, configured to create a task point index and a path planning model through OR-Tools according to the data model. Among them, the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem;

[0073] A task path planning module, configured to perform task path planning according to the task point index through the path planning model to obtain the optimal path for executing the task.

[0074] In one of the embodiments, acquiring information about the task to be executed and establishing a data model according to the information about the task to be executed, where the information about the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix, includes:

[0075] A first generation module, configured to generate task target coordinates, and there are multiple task target coordinates;

[0076] A second generation module, configured to generate the demand quantity of the task target, and generate the agent task carrying capacity according to the demand quantity;

[0077] An analysis module, configured to calculate the distances between the task target coordinates according to the task target coordinates to obtain a distance matrix.

[0078] In one of the embodiments, generating task target coordinates, and there are multiple task target coordinates, includes:

[0079] A task target location generation module, which is used to randomly generate a set of task target locations in a two-dimensional plane, and each task target location has a fixed demand.

[0080] In one embodiment, a task point index and a path planning model are created through OR-Tools according to a data model. Among them, the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem, including:

[0081] A summarization module, which is used to summarize the generated task target locations, demands, the task-carrying capacity of agents, and the distance matrix into a data model.

[0082] A transfer module, which is used to manage the task point index through the "RoutingIndexManager" of OR-Tools, define the constraints and objectives of the path planning problem, and transfer the number of unmanned equipment, the number of tasks, and the task target locations to the path planning model.

[0083] In one embodiment, task path planning is performed through the path planning model according to the task point index to obtain the optimal path for task execution, including:

[0084] Defined through a Markov decision process:

[0085] A state space module, which is used to represent each state S t =(V t , X t ) ∈ S consists of two parts. The first part is the equipment state V t , and the second part is the node state S t . The state information includes: the position, speed, computing load, and power of the agent;

[0086] An action space module, which defines the action as selecting an equipment and a node for access, and the action information includes: the motion posture, computing power allocation, and node scheduling strategy of the agent;

[0087] A transition rule module, which is used to define the transition rule τ to transfer the previous state s to the next state s t to s t+1 , that is, s t+1 =(V t+1 , X t+1 ) = τ(V t , X t );

[0088] A reward function module, which defines the reward as the negative value of this maximum value, and the reward is calculated by cumulatively summing the travel times of each piece of equipment for multiple trips.

[0089] In one embodiment, the reward and punishment settings of the reward function include:

[0090] A proximity-to-goal reward module, which gives a reward when the agent approaches the goal point and a punishment when the agent moves away from the goal point;

[0091] An obstacle reward module, where the closer the agent is to the obstacle, the greater the punishment;

[0092] A smoothness reward module, which calculates whether the trajectory is smooth based on the angle between the coordinates of the agent in the previous five steps and the coordinates before the execution of the action, and the coordinates before the execution of the action and the coordinates after the execution of the action;

[0093] A backtracking reward module, which determines whether the agent has taken a backtracking route based on whether it has passed through the coordinates of the previous ten steps. If it has passed through the coordinates of the previous ten steps, a certain punishment will be given;

[0094] An end-point reward module, which gives a reward when the end point is reached.

[0095] In one embodiment, the intelligent agent control method for unmanned equipment based on the AI-OODA loop further includes:

[0096] A training module, which, after a large number of rounds of training through the DQN reinforcement algorithm, plans the path in the virtual-real fusion scenario, and the unmanned equipment travels according to the planned route;

[0097] An on-the-spot decision-making module, which makes on-the-spot decisions when the unmanned equipment automatically avoids obstacles and discovers a target in real time, automatically generates the next decision task, and sends it to the corresponding equipment to execute the task.

[0098] Those skilled in the art can understand that Figure 5 the structure shown in

[0099] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer equipment to which the solution of the present application is applied. The specific computer equipment may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0100] Obtain the information of the task to be executed, establish a data model according to the information of the task to be executed, and the information of the task to be executed includes: task target coordinates, demand quantity, intelligent agent task carrying capacity, and distance matrix;

[0101] Create a task point index and a path planning model through OR-Tools according to the data model, where the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem;

[0102] Perform task path planning through the path planning model according to the task point index to obtain the optimal path for task execution.

[0103] In one embodiment, when the processor executes the computer program, it realizes obtaining the information of the task to be executed, establishing a data model according to the information of the task to be executed, and the information of the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix, including:

[0104] Generate task target coordinates, and there are multiple task target coordinates;

[0105] Generate the demand quantity of the task target, and generate the agent task carrying capacity according to the demand quantity;

[0106] According to the task target coordinates, calculate the distances between the task target coordinates to obtain the distance matrix.

[0107] In one embodiment, when the processor executes the computer program, it realizes generating task target coordinates, and there are multiple task target coordinates, including:

[0108] Randomly generate a set of task target positions in a two-dimensional plane, and each task target position has a fixed demand quantity.

[0109] In one embodiment, when the processor executes the computer program, it realizes creating a task point index and a path planning model through OR-Tools according to the data model, where the task point index includes: a starting point and an ending point, and the path planning model is used to define the constraints and objectives of the path planning problem, including:

[0110] Summarize the generated task target positions, demand quantities, agent task carrying capacities, and distance matrices into a data model;

[0111] Manage the task point index through the "RoutingIndexManager" of OR-Tools, define the constraints and objectives of the path planning problem, and pass the number of unmanned equipment, the number of tasks, and the task target positions to the path planning model.

[0112] In one embodiment, when the processor executes the computer program, it realizes performing task path planning through the path planning model according to the task point index to obtain the optimal path for task execution, including:

[0113] Define through the Markov decision process:

[0114] State space: Each state St =(V t , X t ) ∈ S consists of two parts. The first part is the equipment status V t , and the second part is the node status S t . The status information includes: the position, speed, computing load, and power of the agent;

[0115] Action space: The action is defined as selecting an equipment and a node for access. The action information includes: the motion posture, computing power allocation, and node scheduling strategy of the agent;

[0116] Transition rule: The transition rule τ transfers the previous state s to the next state s t according to the executed action, that is, s t+1 =(V t+1 , X t+1 )=(V t+1 , X t ) = τ(V t );

[0117] Reward function: The reward is defined as the negative value of this maximum value, and the reward is calculated by separately accumulating the travel times of multiple trips of each piece of equipment.

[0118] In one embodiment, when the processor executes the computer program, the reward and punishment settings of the reward function are implemented, including:

[0119] There will be a reward when the agent approaches the target point, and a punishment when the agent moves away from the target point;

[0120] The closer the agent is to the obstacle, the greater the punishment;

[0121] Whether the trajectory is smooth is calculated according to the angle between the coordinates of the agent in the first five steps to the coordinates before executing the action and the coordinates before executing the action to the coordinates after executing the action;

[0122] Whether the agent takes a detour is determined according to whether it has passed through the coordinates of the previous ten steps. If it has passed through the coordinates of the previous ten steps, a certain punishment will be given;

[0123] Reaching the end point gives a reward.

[0124] In one embodiment, the intelligent agent control method for unmanned equipment based on the AI-OODA loop further includes:

[0125] After a large number of rounds of training through the DQN reinforcement algorithm, the path planning route in the virtual-real fusion scenario, and the unmanned equipment travels according to the planned route;

[0126] Ad-hoc decision-making occurs after the unmanned equipment automatically avoids obstacles in real time and discovers a target, automatically generates the next decision-making task, and sends it to the corresponding equipment to execute the task.

[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0128] Obtain the information of the task to be executed, and establish a data model according to the information of the task to be executed. The information of the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix;

[0129] Create a task point index and a path planning model through OR-Tools according to the data model. Among them, the task point index includes: starting point and ending point, and the path planning model is used to define the constraints and objectives of the path planning problem;

[0130] Perform task path planning through the path planning model according to the task point index to obtain the optimal path for executing the task.

[0131] In one of the embodiments, when the computer program is executed by a processor, it realizes obtaining the information of the task to be executed, and establishing a data model according to the information of the task to be executed. The information of the task to be executed includes: task target coordinates, demand quantity, agent task carrying capacity, and distance matrix, including:

[0132] Generate task target coordinates, and there are multiple task target coordinates;

[0133] Generate the demand quantity of the task target, and generate the agent task carrying capacity according to the demand quantity;

[0134] Calculate the distances between the task target coordinates according to the task target coordinates to obtain the distance matrix.

[0135] In one of the embodiments, when the computer program is executed by a processor, it realizes generating task target coordinates, and there are multiple task target coordinates, including:

[0136] Randomly generate a set of task target positions in a two-dimensional plane, and each task target position has a fixed demand quantity.

[0137] In one of the embodiments, when the computer program is executed by a processor, it realizes creating a task point index and a path planning model through OR-Tools according to the data model. Among them, the task point index includes: starting point and ending point, and the path planning model is used to define the constraints and objectives of the path planning problem, including:

[0138] Summarize the generated task target positions, demand quantity, agent task carrying capacity, and distance matrix into a data model;

[0139] Manage the task point index through the "RoutingIndexManager" of OR-Tools, define the constraints and objectives of the path planning problem, and pass the number of unmanned equipment, the number of tasks, and the task target locations to the path planning model.

[0140] In one of the embodiments, when the computer program is executed by a processor, it realizes task path planning through a path planning model according to the task point index, and obtains the optimal path for executing the task, including:

[0141] Define it through a Markov decision process:

[0142] State space: Each state S t =(V t , X t ) ∈ S consists of two parts. The first part is the equipment state V t , and the second part is the node state S t . The state information includes: the position, speed, computing load, and power of the agent;

[0143] Action space: The action is defined as selecting an equipment and a node for access. The action information includes: the motion posture, computing power allocation, and node scheduling strategy of the agent;

[0144] Transition rule: The transition rule τ transfers the previous state s to the next state s t according to the executed action, that is, s t+1 =(V t+1 , X t+1 )=(V t+1 , X t )=τ(V t );

[0145] Reward function: The reward is defined as the negative value of this maximum value, and the reward is calculated by separately accumulating the travel times of multiple trips of each piece of equipment.

[0146] In one of the embodiments, when the computer program is executed by a processor, it realizes the reward and punishment settings of the reward function, including:

[0147] There will be a reward when the agent approaches the target point, and there will be a punishment when the agent moves away from the target point;

[0148] The closer the agent is to the obstacle, the greater the punishment;

[0149] Calculate whether the trajectory is smooth according to the angle between the coordinates of the agent in the previous five steps to the coordinates before executing the action and the coordinates before executing the action to the coordinates after executing the action;

[0150] Based on whether the agent has passed through the coordinates of the first ten steps, it is determined whether it takes a detour. If it has passed through the coordinates of the first ten steps, a certain penalty will be imposed;

[0151] Reaching the end point gives a reward.

[0152] In one of the embodiments, when the computer program is executed by a processor, it implements an intelligent agent control method for unmanned equipment based on the AI-OODA loop, and further includes:

[0153] After a large number of rounds of training through the DQN reinforcement algorithm, the path planning route in the virtual-real fusion scenario is obtained, and the unmanned equipment travels according to the planned route;

[0154] The on-the-spot decision-making occurs after the unmanned equipment automatically avoids obstacles in real time and discovers a target. It automatically generates the next decision-making task and sends it to the corresponding equipment to execute the task.

[0155] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0156] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0157] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An unmanned equipment intelligent agent control method based on AI-OODA loop, characterized in that: The method comprises: Acquire information about tasks to be performed, and establish a data model based on the information about tasks to be performed, wherein the information about tasks to be performed includes: task target coordinates, demand, agent task carrying capacity, and distance matrix; Creating a task point index and a path planning model through OR-Tools according to the data model, wherein the task point index includes: a starting point and an end point, and the path planning model is used to define constraints and goals of the path planning problem; The task path is planned according to the task point index through the path planning model to obtain the optimal path for executing the task.

2. The unmanned equipment intelligent agent control method based on AI-OODA loop according to claim 1 is characterized in that: The obtaining of information on tasks to be performed, and establishing a data model according to the information on tasks to be performed, wherein the information on tasks to be performed includes: task target coordinates, demand, agent task carrying capacity, and distance matrix, including: Generate task target coordinates, wherein the task target coordinates are multiple; Generate a demand for a task target, and generate an agent task carrying capacity based on the demand; According to the task target coordinates, the distances between the task target coordinates are calculated to obtain a distance matrix.

3. The unmanned equipment intelligent agent control method based on AI-OODA loop according to claim 2 is characterized in that: The generated task target coordinates are multiple task target coordinates, including: A set of task target positions are randomly generated in a two-dimensional plane, each of which has a fixed demand.

4. The unmanned equipment intelligent agent control method based on AI-OODA loop according to claim 2 is characterized in that: The task point index and the path planning model are created through OR-Tools according to the data model, wherein the task point index includes: a starting point and an end point, and the path planning model is used to define the constraints and objectives of the path planning problem, including: Aggregating the generated task target location, the demand, the agent task carrying capacity and the distance matrix into a data model; The task point index is managed by OR-Tools' "RoutingIndexManager", the constraints and objectives of the path planning problem are defined, and the number of unmanned equipment, the number of tasks and the task target positions are passed to the path planning model.

5. The unmanned equipment intelligent agent control method based on AI-OODA loop according to claim 1 is characterized in that: The step of performing task path planning through the path planning model according to the task point index to obtain an optimal path for executing the task includes: Defined by a Markov decision process: State space: Each state S t =(V t , X t )∈S consists of two parts. The first part is the equipment status V t , the second part is the node status S t ,The state information includes: the agent’s position, speed, computation load, and power; Action space: An action is defined as selecting a device and a node to access. The action information includes: the motion posture of the agent, computing power allocation, and node scheduling strategy; Transition rule: The transition rule τ depends on the action executed. The previous state s t Transition to next state s t+1 , that is, s t+1 =(V t+1 , X t+1 )=τ(V t , X t ); Reward function: The reward is defined as the negative of this maximum value, and the reward is calculated by accumulating multiple travel times for each device separately.

6. The unmanned equipment intelligent agent control method based on AI-OODA loop according to claim 5 is characterized in that: The reward and punishment settings of the reward function include: When the agent approaches the target point, there will be a reward, and when the agent moves away from the target point, there will be a penalty; The closer the agent is to the obstacle, the greater the penalty; Whether the trajectory is smooth is calculated based on the angle between the coordinates of the agent at the first five steps to the coordinates before the action is performed and the angle between the coordinates before the action is performed and the coordinates after the action is performed; Whether the agent goes back is determined based on whether it has passed through the previous ten coordinates. If it has passed through the previous ten coordinates, a certain penalty will be given. Rewards will be given upon reaching the finish line.

7. The unmanned equipment intelligent agent control method based on AI-OODA loop according to claim 1 is characterized in that: Also includes: After a large number of rounds of training using the DQN reinforcement algorithm, the path in the virtual-reality fusion scene is planned, and the unmanned equipment moves along the planned route; On-the-spot decision-making occurs after the unmanned equipment automatically avoids obstacles and discovers targets in real time, automatically generates the next decision task, and sends it to the corresponding equipment to execute the task.

8. An unmanned equipment intelligent agent control device based on AI-OODA loop, characterized in that: The device comprises: An acquisition module is used to acquire information of tasks to be performed, and establish a data model according to the information of tasks to be performed, wherein the information of tasks to be performed includes: task target coordinates, demand, agent task carrying capacity and distance matrix; A task point index and path planning model creation module, used to create a task point index and a path planning model through OR-Tools according to the data model, wherein the task point index includes: a starting point and an end point, and the path planning model is used to define the constraints and objectives of the path planning problem; The task path planning module is used to perform task path planning through the path planning model according to the task point index to obtain the optimal path for executing the task.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.