A second-order cluster increment enveloping control method, device and equipment

CN115879537BActive Publication Date: 2026-05-29NORTHWESTERN POLYTECHNICAL UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2022-12-14
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, the number of agents in a cluster system is fixed, resulting in a limited coverage area and low efficiency and success rate in task execution.

Method used

By establishing the communication topology of the original cluster system, initial state information is obtained, and the second-order communication topology and agent state information of the target cluster system are determined based on the agent's incremental instructions. Control quantities are calculated, and agent states are updated to achieve incremental encirclement control.

Benefits of technology

It expands the task scope of the cluster system, improves the efficiency and success rate of task execution, reduces resource waste, and extends the lifespan of hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115879537B_ABST
    Figure CN115879537B_ABST
Patent Text Reader

Abstract

The application discloses a second-order cluster increment encirclement control method and device and equipment, and relates to the field of artificial intelligence, and is used for solving the problems that the number of intelligent agents in a cluster system is invariable, and the efficiency and success rate of task execution are low in the prior art. The method comprises the following steps: a communication topology of an original cluster system is established, and initial state information of the original cluster system is obtained; a second-order communication topology corresponding to a target cluster system obtained by adding an increment intelligent agent set and state information of the intelligent agents are determined based on an intelligent agent increment instruction; control amounts of each intelligent agent in the target cluster system under a current encirclement target are calculated; and the state information of each intelligent agent in the target cluster system under the current encirclement target is updated according to the control amounts, so that the target increment encirclement control is realized. The method overcomes the difficulty that the encirclement range of the cluster system is limited in the encirclement task, expands the task range of the cluster system, and improves the efficiency and success rate of task execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a second-order cluster incremental encirclement control method, apparatus, and device. Background Technology

[0002] With the development and maturation of artificial intelligence theory and related research technologies, the capabilities of intelligent agents are constantly improving, and their application areas and scope are continuously expanding.

[0003] However, when performing tasks in complex and ever-changing environments, the capabilities of a single agent are insufficient to meet current demands. Therefore, researchers are considering using multiple individual agents to form a swarm system, enabling collaboration among these agents to accomplish more complex tasks. Research on swarm systems is an inevitable trend in the development of artificial intelligence technology and will play a crucial role in its advancement. Swarm control is the foundation and a key focus of research in swarm systems.

[0004] Encirclement control is a crucial application of swarm systems. In traditional encirclement control, the number of agents in a swarm system is fixed, limiting the encirclement range that the system can achieve. In this case, if the encirclement range needs to be expanded, the distance between agents increases, weakening the encirclement effect, and the target may escape through the gaps between agents. Furthermore, the lifespan of hardware and agent resources is finite during task execution, placing higher demands on the encirclement control of swarm systems.

[0005] Therefore, there is an urgent need to provide a more reliable second-order cluster incremental encirclement control scheme. Summary of the Invention

[0006] The purpose of this invention is to provide a second-order cluster incremental encirclement control method, apparatus, and device to solve the problems of the fixed number of agents in the cluster system, the limited encirclement range of the cluster system in the encirclement task, and the low efficiency and success rate of the task execution in the prior art.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] In a first aspect, the present invention provides a second-order cluster incremental encirclement control method, comprising:

[0009] Establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system;

[0010] Obtain incremental instructions from the agents; the incremental instructions from the agents must include at least the set of incremental agents.

[0011] Based on the incremental instructions of the agent, the second-order communication topology and the state information of the agent corresponding to the target cluster system are determined; the target cluster system is the cluster system obtained by adding an incremental set of agents to the original cluster system.

[0012] Calculate the control quantity of each agent in the target cluster system under the current surrounding target;

[0013] The state information of each agent in the target cluster system is updated based on the control variables under the current surrounding target to achieve incremental target encirclement control.

[0014] Compared with existing technologies, the second-order cluster incremental encirclement control method provided by this invention establishes the communication topology of the original cluster system and obtains the initial state information of the original cluster system; based on the incremental instructions of the agents, it determines the second-order communication topology and the state information of the agents corresponding to the target cluster system obtained by adding the incremental agent set; it calculates the control quantity of each agent in the target cluster system under the current encirclement target; and it updates the state information of each agent in the target cluster system according to the control quantity under the current encirclement target, thereby realizing target incremental encirclement control. This overcomes the dilemma of the limited encirclement range of the cluster system in the encirclement task, expands the task range of the cluster system, and improves the efficiency and success rate of task execution.

[0015] In a second aspect, the present invention provides a second-order cluster incremental encirclement control device, the device comprising:

[0016] The original communication topology establishment module is used to establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system.

[0017] The agent incremental instruction acquisition module is used to acquire agent incremental instructions; the agent incremental instructions include at least an incremental agent set;

[0018] The target cluster system information determination module is used to determine the second-order communication topology and the state information of the agent corresponding to the target cluster system based on the agent incremental instructions; the target cluster system is a cluster system obtained by adding an incremental set of agents to the original cluster system.

[0019] The control quantity calculation module is used to calculate the control quantity of each agent in the target cluster system under the current surrounding target;

[0020] The update module is used to update the state information of each agent in the target cluster system according to the control quantity under the current surrounding target, so as to realize incremental target encirclement control.

[0021] The technical effects achieved by the device-type solutions in the second aspect are the same as those achieved by the method-type solutions, and will not be elaborated here.

[0022] Thirdly, the present invention provides a second-order cluster incremental encirclement control device, the device comprising:

[0023] A communication unit / communication interface is used to establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system.

[0024] Obtain incremental instructions from the agents; the incremental instructions from the agents must include at least the set of incremental agents.

[0025] A processing unit / processor is used to determine the second-order communication topology and the state information of the agent corresponding to the target cluster system based on the incremental instructions of the agent; the target cluster system is a cluster system obtained by adding an incremental set of agents to the original cluster system.

[0026] Calculate the control quantity of each agent in the target cluster system under the current surrounding target;

[0027] The state information of each agent in the target cluster system is updated based on the control variables under the current surrounding target to achieve incremental target encirclement control.

[0028] Fourthly, the present invention provides a computer storage medium storing instructions that, when executed, implement the above-described second-order cluster incremental encirclement control method.

[0029] The technical effects achieved by the solutions provided in the third and fourth aspects are the same as those achieved by the method-based solutions, and will not be elaborated here. Attached Figure Description

[0030] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0031] Figure 1 A schematic diagram of a second-order cluster incremental encirclement control method provided by the present invention;

[0032] Figure 2 This is a schematic diagram comparing traditional encirclement control with incremental encirclement control in this invention.

[0033] Figure 3 This invention provides a simulation diagram of a second-order cluster incremental encirclement control example.

[0034] Figure 4 This invention provides a schematic diagram of the position evolution in a second-order cluster incremental encirclement control process.

[0035] Figure 5This invention provides a flowchart of incremental encirclement calculation in a second-order cluster incremental encirclement control method;

[0036] Figure 6 A schematic diagram of a second-order cluster incremental encirclement control device provided by the present invention;

[0037] Figure 7 This is a schematic diagram of a second-order cluster incremental encirclement control device provided by the present invention. Detailed Implementation

[0038] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.

[0039] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0040] In this invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0041] In existing technologies, for optimal control of swarm systems, traditional optimal control theory is designed based on the gradient of the optimal cost function, which is obtained by solving the Hamilton-Jacobi-Bellman equations. However, these equations are difficult to solve analytically due to their inherent nonlinearity. Over the past few decades, function approximation strategies based on reinforcement learning have been successfully applied to adaptive optimal control, becoming a popular method for solving complex control problems. Many of these methods have achieved good performance in various systems, but research on the optimal combination of backstepping and reinforcement learning for systems with nonlinear dynamics and strict feedback is scarce. Therefore, in summary, only by proposing more practical swarm encirclement control methods can we overcome the dilemma of the fixed number of agents in swarm systems, expand the task scope of swarm systems, and improve the efficiency and success rate of task execution.

[0042] In response, this invention provides a second-order cluster incremental encirclement control scheme.

[0043] Next, the solutions provided in the embodiments of this specification will be described in conjunction with the accompanying drawings:

[0044] Figure 1 This invention provides a schematic flowchart of a second-order cluster incremental encirclement control method, as shown below. Figure 1 As shown, the process may include the following steps:

[0045] Step 110: Establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system.

[0046] The communication topology of the original cluster system can include a first-order communication topology, and the number and form of agents in the topology can be adjusted according to the actual situation.

[0047] Step 120: Obtain agent incremental instructions; the agent incremental instructions must include at least the incremental agent set.

[0048] In practical applications, when it is necessary to increase the number of agents, an agent increment instruction can be obtained. This instruction can include at least the set of agents to be added, i.e., the incremental agent set. Of course, it can also include information such as the number of agents to be added and agent identifiers.

[0049] Step 130: Based on the incremental instructions of the agent, determine the second-order communication topology and the state information of the agent corresponding to the target cluster system; the target cluster system is the cluster system obtained by adding an incremental set of agents to the original cluster system.

[0050] After adding an incremental set of agents to the original cluster system, the new second-order communication topology and the state information of the agents in the newly formed target cluster system can be determined.

[0051] Step 140: Calculate the control quantity of each agent in the target cluster system under the current surrounding target.

[0052] Step 150: Update the state information of each agent in the target cluster system according to the control quantity under the current surrounding target to realize incremental target surrounding control.

[0053] Control variables can represent the control inputs of an incremental agent swarm system. Based on these control inputs, the agent's state is updated. The agent's state can include its operational status, such as changes in action, distance traveled, and orientation. State-based updates can control the incremental agent to move to the target location, achieving incremental target encirclement control.

[0054] Figure 1 The proposed method establishes the communication topology of the original cluster system and obtains its initial state information. Based on the incremental instructions of the agents, it determines the second-order communication topology and the state information of the agents in the target cluster system obtained by adding the incremental agents. It calculates the control quantities of each agent in the target cluster system under the current surrounding target. Based on the control quantities under the current surrounding target, it updates the state information of each agent in the target cluster system, thereby achieving incremental surrounding control of the target. This overcomes the predicament of the limited surrounding range of the cluster system in the surrounding task, expands the task range of the cluster system, and improves the efficiency and success rate of task execution.

[0055] Building upon traditional cluster system encirclement control methods, this scheme proposes incremental-encirclement control, combining optimal control, backstepping, and reinforcement learning to achieve optimal incremental encirclement control for the cluster system. Figure 2 As shown, in traditional encirclement control, if there is only one target, single-layer encirclement control can be used, and the target cannot escape. However, when there are multiple targets, such as... Figure 2 When two targets exist, agents must increase the distance between them to avoid collisions and thus encircle both targets. However, this increased distance creates feasible escape routes for the targets (black dashed lines), reducing the success rate of encirclement. Incremental encirclement control, on the other hand, cannot meet the demands of a swarm system due to its fixed number of agents and limited encirclement range. More agents are needed to expand the encirclement. Therefore, incremental encirclement control allows for a variable number of agents in the swarm system, increasing its flexibility and making it practically valuable.

[0056] First, incremental encirclement control is achieved by updating the communication topology and agent state using incremental information. Then, optimal cost function is introduced with error and input as optimization objectives to achieve optimal incremental encirclement control. Finally, based on this, the controller is designed using reinforcement learning based on backstepping.

[0057] For systems with nonlinear dynamic equations, this paper proposes an optimal incremental encirclement control method. Leveraging the nonlinear fitting capability of classical reinforcement learning networks (AC neural networks), a controller based on backstepping is designed, and its stability is proven using the designed Lyapunov function. Furthermore, combined with optimal control, an optimal incremental encirclement control method is proposed, effectively reducing resource waste and extending hardware lifespan. The specific implementation process will be explained in detail below.

[0058] based on Figure 1 In addition to the method described herein, this specification also provides some specific implementation methods of the method, which will be described below.

[0059] Optionally, based on the incremental instructions from the agent, the second-order communication topology and the state information of the agent corresponding to the target cluster system are determined; the target cluster system is a cluster system obtained by adding an incremental set of agents to the original cluster system, and may specifically include:

[0060] Based on the agent incremental instructions, an incremental agent set is added to the original cluster system to obtain an updated second-order communication topology; the updated second-order communication topology is composed of the adjacency matrix of the original cluster system, the adjacency matrix of the incremental agents, and the incremental connections between the incremental agents and the original agents.

[0061] The state information of the agent is updated based on the dynamic model of the original cluster system and the updated second-order communication topology.

[0062] This solution will be illustrated with specific examples, such as... Figure 3 As shown, the number and form of agents in the topology can be adjusted according to the actual situation. An example is used to demonstrate the effectiveness of the controller. The communication topology of the cluster system is established in MATLAB software, and the initial state information of the cluster system model is obtained. First, the number of agents in the cluster system is set to N. Then, a directed graph G of the communication topology is established, and its adjacency matrix... Store edge information, matrix Storage degree information, matrix Storing identity information, thus obtaining the Laplace matrix is And satisfy The initial state information includes the current position coordinates (x0, y0) and velocity v0 of agent 0; the state information of the surrounded target agent i is: position coordinates (x0, y0) and velocity v0. i y i ), velocity v i Incremental information is used to update the communication topology and the state of the agents.

[0063] For example: Figure 4As shown, the original cluster system has 6 agents, labeled as agents 1-6, with agent 7 representing the encircling target. Then, based on the communication topology relationships of agents 1-7, a directed communication topology graph G is generated, with its adjacency matrix... Store edge information, matrix Storage degree information, matrix Storing identity information, thus obtaining the Laplace matrix is And satisfy The initial state information includes the current position coordinates (x, y) of the six agents. i y i ), velocity v i , where i represents the agent number; the state information of the surrounded target agent 7 is: position coordinates (x7, y7) and velocity v7.

[0064] When determining the second-order communication topology and agent state information corresponding to the target cluster system based on the agent's incremental instructions, the following steps can be adopted:

[0065] At time t, an incremental agent N2 is requested to join, where M = N + N2. The updated communication topology is composed of the adjacency matrix of the original cluster system, the adjacency matrix of the incremental agents, and the incremental connections between the incremental agents and the original agents. For example, the incremental agents requested to join at time t could be agents numbered 8-12.

[0066] Based on the dynamic model of the cluster system and the updated communication topology Update the state X of the agent and store the position and velocity information of all agents.

[0067] Furthermore, calculating the control quantity of each agent in the target cluster system under the current surrounding target can specifically include:

[0068] Calculate the position error based on the state information of each agent;

[0069] Reinforcement learning based on backstepping to compute a virtual controller;

[0070] Calculate higher-order position error, velocity error, and higher-order velocity error;

[0071] The control quantity is calculated based on the position error, virtual controller, higher-order position error, velocity error, and higher-order velocity error.

[0072] Furthermore, the specific calculation process for the control quantity is as follows:

[0073] Input the state information of each agent and calculate the position error based on formula (1):

[0074]

[0075] Where, x1=(x 1,1 x 2,1 , ..., x M,1 ) represents the current position information of the agents in the communication topology, M = N + N², where M represents the agent numbers 1-6 and 8-12, x 0,1 ρ represents the current position information of target agent 7 in the communication topology, ρ is the bias defined according to the formation, i.e., the distance between the two agents, and ε1 is the position error.

[0076] The virtual controller is calculated using formula (2):

[0077]

[0078] in It is a virtual controller, h is the Gaussian function, h = [h1, h2, ..., h q ], where q is the number of neurons in the neural network, and β1 is a constant. By choosing their values, good tracking performance can be obtained. The optimal estimated weights for the AC neural network are:

[0079]

[0080]

[0081]

[0082] The higher-order position error is calculated using formula (3):

[0083]

[0084] in, f1 represents the unknown nonlinear dynamics existing within the system.

[0085] The speed error is calculated using formula (4):

[0086]

[0087] x2=(x 1,2 x 2,2 , ..., x M,2 ) represents the current speed information of the agent in the communication topology.

[0088] The higher-order velocity error is calculated using formula (5):

[0089]

[0090] f2 represents the unknown nonlinear dynamics existing within the system.

[0091] The control quantity is calculated using formula (6):

[0092]

[0093] Wherein, the initial control input is u o = (cost, sint), where the AC neural network weight update formulas are Equations (7) and (8):

[0094]

[0095]

[0096] Based on the control input, the agent's state information is updated using formula (9):

[0097] x=x+Δx·dt (9)

[0098] Where, f1 = sin(x1), f2 = sin(x2),

[0099] Additionally, after obtaining incremental instructions from the agent, such as Figure 5 As shown, it can be determined first whether agent increment is needed in the original cluster system; if agent increment is needed in the original cluster system, the state information of each agent in the incremental agent set is input, and the communication topology of the original cluster system and the state information of the agents are updated; if agent increment is not needed in the original cluster system, the communication topology of the original cluster system is kept unchanged.

[0100] Specifically, the communication topology is updated based on formula (10):

[0101]

[0102] The agent's state is updated using formula (11):

[0103]

[0104] The above scheme addresses the problem of an immutable number of agents and a limited encirclement range in a cluster system by employing an incremental encirclement control method. When the cluster system cannot meet the demand and more agents are needed to expand the encirclement range, the number of agents in the cluster system becomes variable, increasing the flexibility of the cluster system and demonstrating practical value. Furthermore, combined with optimal control, an optimal incremental encirclement control method is proposed, effectively reducing resource waste and extending the lifespan of the hardware. Based on the optimal incremental encirclement control, a controller is designed using reinforcement learning based on backstepping, leveraging the nonlinear fitting capability of classical reinforcement learning networks (AC neural networks). The stability of the controller is proven through the designed Lyapunov function.

[0105] Based on the same idea, the present invention also provides a second-order cluster incremental encirclement control device, such as... Figure 6 As shown, the device may include:

[0106] The original communication topology establishment module 610 is used to establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system.

[0107] The agent incremental instruction acquisition module 620 is used to acquire agent incremental instructions; the agent incremental instructions include at least an incremental agent set.

[0108] The target cluster system information determination module 630 is used to determine the second-order communication topology and the state information of the agent corresponding to the target cluster system based on the agent incremental instructions; the target cluster system is a cluster system obtained by adding an incremental set of agents to the original cluster system.

[0109] The control quantity calculation module 640 is used to calculate the control quantity of each intelligent agent in the target cluster system under the current surrounding target.

[0110] The update module 650 is used to update the state information of each agent in the target cluster system according to the control quantity under the current surrounding target, so as to realize the incremental surrounding control of the target.

[0111] based on Figure 6 The device may also include specific implementation units:

[0112] Optionally, the target cluster system information determination module 630 may specifically include:

[0113] The second-order communication topology update unit is used to add an incremental set of intelligent agents to the original cluster system based on the incremental instructions of the intelligent agents, so as to obtain the updated second-order communication topology; the updated second-order communication topology is composed of the adjacency matrix of the original cluster system, the adjacency matrix of the incremental intelligent agents, and the incremental connection between the incremental intelligent agents and the original intelligent agents.

[0114] The agent state information update unit is used to update the agent state information according to the dynamic model of the original cluster system and the updated second-order communication topology.

[0115] Optionally, the control quantity calculation module 640 may specifically include:

[0116] The position error calculation unit is used to calculate the position error based on the state information of each agent.

[0117] The virtual controller computation unit is used for reinforcement learning based on backstepping to compute the virtual controller.

[0118] The error calculation unit is used to calculate higher-order position errors, velocity errors, and higher-order velocity errors.

[0119] The control quantity calculation unit is used to calculate the control quantity based on the position error, virtual controller, higher-order position error, speed error, and higher-order speed error.

[0120] Optionally, the original communication topology establishment module 610 may specifically include:

[0121] The agent number setting unit is used to set the number of agents in the original cluster system;

[0122] The communication topology directed graph establishment unit is used to establish the communication topology directed graph of the original cluster system based on the number of agents in the original cluster system.

[0123] An initial state information acquisition unit is used to acquire the initial state information of the original cluster system; the initial state information includes the current position coordinates and velocity of the agents in the original cluster system.

[0124] Optionally, the target cluster system information determination module 630 may specifically include:

[0125] The target cluster system determination unit is used to add the incremental intelligent agent set to the original cluster system based on the intelligent agent incremental instructions to obtain the target cluster system.

[0126] The target second-order communication topology establishment unit is used to establish the target second-order communication topology corresponding to the target cluster system and to obtain the state information of each agent in the target cluster system.

[0127] Optional, Figure 6 The device may also include:

[0128] The incremental judgment module is used to determine whether agent increment is needed in the original cluster system;

[0129] The information update module is used to update the communication topology of the original cluster system and the status information of the agents by inputting the status information of each agent in the incremental agent set if agent increment is required in the original cluster system; otherwise, the communication topology of the original cluster system remains unchanged.

[0130] Based on the same approach, this specification also provides a second-order cluster incremental encirclement control device. For example... Figure 7 As shown. It can include:

[0131] A communication unit / communication interface is used to establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system.

[0132] Obtain incremental instructions from the agents; the incremental instructions from the agents must include at least the set of incremental agents.

[0133] A processing unit / processor is used to determine the second-order communication topology and the state information of the agent corresponding to the target cluster system based on the incremental instructions of the agent; the target cluster system is a cluster system obtained by adding an incremental set of agents to the original cluster system.

[0134] Calculate the control quantity of each agent in the target cluster system under the current surrounding target;

[0135] The state information of each agent in the target cluster system is updated based on the control variables under the current surrounding target to achieve incremental target encirclement control.

[0136] like Figure 7 As shown, the terminal device described above may also include a communication line. The communication line may include a path for transmitting information between the components described above.

[0137] Optional, such as Figure 7 As shown, the terminal device may further include a memory. The memory stores computer execution instructions for implementing the present invention, and the execution is controlled by a processor. The processor executes the computer execution instructions stored in the memory, thereby implementing the method provided in the embodiments of the present invention.

[0138] like Figure 7 As shown, the memory can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited to these. The memory can exist independently and be connected to the processor via communication lines. The memory can also be integrated with the processor.

[0139] Optionally, the computer execution instructions in the embodiments of the present invention may also be referred to as application code, and the embodiments of the present invention do not specifically limit this.

[0140] In a specific implementation, as one example, such as Figure 7 As shown, a processor may include one or more CPUs, such as Figure 7 CPU0 and CPU1 in the CPU.

[0141] In a specific implementation, as one example, such as Figure 7 As shown, the terminal device may include multiple processors, such as Figure 7 The processors in the system. Each of these processors can be a single-core processor or a multi-core processor.

[0142] The foregoing mainly describes the solutions provided by the embodiments of the present invention from the perspective of the interaction between various modules. It is understood that each module, in order to achieve the above functions, includes corresponding hardware structures and / or software units for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the present invention can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0143] The embodiments of the present invention can divide functional modules according to the above method examples. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in the embodiments of the present invention is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0144] The processor described in this specification may also function as a memory. The memory stores computer execution instructions for carrying out the present invention, and its execution is controlled by the processor. The processor executes the computer execution instructions stored in the memory, thereby implementing the method provided in the embodiments of the present invention.

[0145] The memory can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited to these. The memory can exist independently and be connected to the processor via communication lines. The memory can also be integrated with the processor.

[0146] Optionally, the computer execution instructions in the embodiments of the present invention may also be referred to as application code, and the embodiments of the present invention do not specifically limit this.

[0147] The methods disclosed in the above embodiments of the present invention can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0148] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).

[0149] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0150] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely exemplary descriptions of the invention as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if such modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include such modifications and modifications.

Claims

1. A second-order cluster incremental encirclement control method, characterized in that, include: Establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system; Obtain incremental instructions from the agent; The incremental instructions for the intelligent agent include at least a set of incremental intelligent agents; Based on the incremental instructions of the agent, the second-order communication topology and the state information of the agent corresponding to the target cluster system are determined; the target cluster system is the cluster system obtained by adding an incremental set of agents to the original cluster system. Calculate the control quantity of each agent in the target cluster system under the current surrounding target; Update the state information of each agent in the target cluster system according to the control quantity under the current surrounding target to achieve incremental target encirclement control; The process of determining the second-order communication topology and agent state information corresponding to the target cluster system based on the agent's incremental instructions specifically includes: Based on the agent incremental instructions, an incremental agent set is added to the original cluster system to obtain an updated second-order communication topology; the updated second-order communication topology is composed of the adjacency matrix of the original cluster system, the adjacency matrix of the incremental agents, and the incremental connections between the incremental agents and the original agents. Based on the dynamic model of the original cluster system and the updated second-order communication topology, update the state information of the agent; Calculating the control quantity of each agent in the target cluster system under the current surrounding target, specifically includes: Calculate the position error based on the state information of each agent; Reinforcement learning based on backstepping to compute a virtual controller; Calculate higher-order position error, velocity error, and higher-order velocity error; The control quantity is calculated based on the position error, virtual controller, higher-order position error, velocity error, and higher-order velocity error; the state information of the intelligent agent includes the changing action, the distance moved, and the changing orientation.

2. The method according to claim 1, characterized in that, Establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system, specifically including: Set the number of agents in the original cluster system; Based on the number of agents in the original cluster system, a directed graph of the communication topology of the original cluster system is established. Obtain the initial state information of the original cluster system; the initial state information includes the current position coordinates and velocity of the agents in the original cluster system.

3. The method according to claim 1, characterized in that, Based on the incremental instructions from the agent, the second-order communication topology corresponding to the target cluster system and the state information of the agent are determined, specifically including: Based on the incremental instructions of the intelligent agent, the incremental intelligent agent set is added to the original cluster system to obtain the target cluster system; Establish the target second-order communication topology corresponding to the target cluster system, and obtain the state information of each agent in the target cluster system.

4. The method according to claim 1, characterized in that, After obtaining the agent's incremental instructions, the process also includes: Determine whether agent increment is needed in the original cluster system; If it is necessary to increment the agent in the original cluster system, then input the state information of each agent in the incremental agent set, and update the communication topology of the original cluster system and the state information of the agents. If agent increment is not required in the original cluster system, the communication topology of the original cluster system remains unchanged.

5. A second-order cluster incremental encirclement control device, characterized in that, The device includes: The original communication topology establishment module is used to establish the communication topology of the original cluster system and obtain the initial state information of the original cluster system. An agent incremental instruction acquisition module is used to acquire agent incremental instructions; the agent incremental instructions include at least an incremental agent set; The target cluster system information determination module is used to determine the second-order communication topology and the state information of the agent corresponding to the target cluster system based on the agent incremental instructions. The target cluster system is a cluster system obtained by adding an incremental set of intelligent agents to the original cluster system; The control quantity calculation module is used to calculate the control quantity of each agent in the target cluster system under the current surrounding target; The update module is used to update the state information of each agent in the target cluster system according to the control quantity under the current surrounding target, so as to realize incremental target surrounding control. The target cluster system information determination module specifically includes: The second-order communication topology update unit is used to add an incremental set of intelligent agents to the original cluster system based on the incremental instructions of the intelligent agents, so as to obtain the updated second-order communication topology; the updated second-order communication topology is composed of the adjacency matrix of the original cluster system, the adjacency matrix of the incremental intelligent agents, and the incremental connection between the incremental intelligent agents and the original intelligent agents. The agent state information update unit is used to update the agent state information according to the dynamic model of the original cluster system and the updated second-order communication topology. The control quantity calculation module specifically includes: The position error calculation unit is used to calculate the position error based on the state information of each agent. The virtual controller computation unit is used for reinforcement learning based on backstepping to compute the virtual controller. The error calculation unit is used to calculate higher-order position errors, velocity errors, and higher-order velocity errors. The control quantity calculation unit is used to calculate the control quantity based on the position error, virtual controller, higher-order position error, velocity error, and higher-order velocity error; the state information of the intelligent agent includes the changing action, the distance moved, and the changing orientation.

6. A second-order cluster incremental encirclement control device, characterized in that the device... include: Memory, processor, and communication interface coupled to the processor; The memory stores computer programs that can be executed by the processor; When the processor runs the computer program, it executes the second-order cluster incremental encirclement control method as described in any one of claims 1 to 4.

7. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed, implement the second-order cluster incremental encirclement control method according to any one of claims 1 to 4.