Unmanned swarm dynamic allocation decision method and system based on formal methods
By applying partially ordered set product algorithm and simulated annealing algorithm in the unmanned cluster system, the problem of high complexity in the task planning of unmanned clusters in the existing technology is solved, and efficient dynamic task allocation and complex action reasoning are achieved.
Patent Information
- Application Number
- CN202410996049.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-24
AI Technical Summary
Existing formal methods are difficult to achieve efficient task planning for large-scale clusters under the constraints of the length of formal language formulas for unmanned cluster tasks, and the complexity of the conversion automaton leads to inefficient computing efficiency.
A complex model cluster dynamic allocation decision-making method based on formal methods is designed. Through partially ordered set product algorithm and simulated annealing algorithm, dynamic allocation of online tasks is realized, reducing algorithm complexity and improving inference ability.
Reducing the complexity of the algorithm from double exponent to polynomial, the ability to reason for complex actions is achieved, and efficient task planning for large-scale clusters is realized.
Smart Images

Figure CN118863438B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a dynamic allocation technology for online collaborative tasks of complex unmanned cluster systems, and specifically to a complex model cluster dynamic allocation method and system based on formal methods. The present invention is based on the unmanned cluster tasks described in formal languages. When the intelligent agents in the intelligent cluster have complex action models, the method and system consider the dynamic allocation decision of unmanned cluster tasks under multi-task and online task situations. Background Art
[0002] Unmanned swarm systems can demonstrate extremely high efficiency when working in parallel and collaboratively. Linear Temporal Logic (LTL) has been widely used in task planning of unmanned swarms due to its formal semantics and rich expressiveness, such as in transportation, maintenance, and service scenarios.
[0003] A standard centralized framework is based on the model checking algorithm: First, the LTL formula representing the unmanned swarm task is converted into a deterministic Robin automaton or a non-deterministic Büchi automaton through tools such as SPIN and LTL2BA. Second, a model describing the agent is used, such as a finite weighted transition system, a Markov decision process, or a Petri net. A product automaton is created between the automaton of the formula and the agent model. Finally, a graph search method is used to find an effective and acceptable path in the product automaton as the agent's plan. Since the complexity of building the product automaton will explode exponentially with the increase in the number of agents, researchers have made improvements in two directions. One direction is to modify the generation method of the product automaton, such as using sampling methods to generate only part of the automaton, or generating a smaller automaton based on the decomposable set method. The other direction is to abandon the generation of the product automaton, but first split the sequence in the automaton into subtasks, and use independence analysis to judge the logic between tasks, and then use algorithms that process combinatorial optimization to allocate tasks, such as branch and bound methods, auction algorithms, etc.
[0004] However, there is a basic step in the above methods, that is, the unmanned cluster task formula needs to be converted into the corresponding automaton, and this conversion may result in a double exponential complexity related to the length of the formula. In fact, for most LTL formulas with a length greater than 25, it takes more than 2 hours and 13GB of memory to calculate the corresponding non-deterministic Büchi automaton through the tool software LTL2BA. In addition, the complexity of the graph search-based method will increase exponentially with the number of agents because it needs to build a product automaton, but the algorithm can plan potential actions that need to be performed based on the system model. The algorithm based on combinatorial optimization has low computational complexity, but can only plan the tasks already mentioned in the formula and lacks the ability to reason about potential actions. Therefore, the existing formal method decision planning technology has not yet been able to break through the constraints of the length of the formal language formula of the unmanned cluster task, and it is difficult to achieve efficient task planning for large-scale clusters while retaining the reasoning ability. Summary of the invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a complex model cluster dynamic allocation decision method and system based on formal methods, which is aimed at unmanned clusters with complex action models under the constraints of formal languages. The dynamic decision method for online tasks is used to solve the dynamic allocation decision problem of online collaborative tasks of unmanned cluster systems. The algorithm can be reduced from double exponential complexity to polynomial complexity and has the ability to reason about complex actions.
[0006] Based on the formal method, the present invention first designs a poset product algorithm, which incrementally calculates the product of the poset of the dynamic LTL formula published online and the existing poset to obtain a poset that can fully satisfy all LTL formulas; then, the complete action trajectory chain is calculated by combining the complex action model of the agent and the unmanned cluster subtask. Finally, based on the agent function and timing constraints, a local search algorithm based on simulated annealing is designed to allocate actions and perform online updates for dynamic tasks.
[0007] The technical solution provided by the present invention is as follows:
[0008] A method and system for online task dynamic decision-making of clusters with complex action models based on formal methods. Considering a heterogeneous intelligent agent unmanned cluster system, a series of linear temporal logic (LTL) tasks are continuously published online and there are some resources that can interact with actions in a dynamic environment. The goal of the solution is to generate a task plan for the system online to efficiently meet these tasks. It includes the following steps:
[0009] 1) Formula to characterize a given unmanned cluster LTL task Transformed into a quasi-pose; including:
[0010] 11) Use the standard algorithm (software tool LTL2BA) to convert the task formula The transformation is expressed as a nondeterministic Büchi automaton (NBA): B = (Q, q0, ∑, δ, F), where Q is the state in the automaton, indicating the degree of task completion; is the initial state of the automaton, indicating that the task has not yet started to be executed; ∑ is the set of atomic propositions in the automaton, indicating all the actions involved in the task; δ is the function of the automaton state transition, indicating the action requirements corresponding to a part of the task; is an acceptable automaton state, indicating that the task has been completed.
[0011] 12) Using the partial order relationship analysis method, the non-deterministic Büchi automaton in step 11) is pruned to obtain the unmanned cluster subtask set Ω, and the partial order relationship between the subtasks in the task sequence is analyzed to continuously generate partial order sets P, which can be expressed as Where Ω represents the subtask set, ≠ indicates a partial order relation; Represents subtask ω i Should be in subtask ω j Before execution, (ω i ,ω j )∈≠ represents the subtask ω i and subtask ω j Cannot be executed simultaneously. Partial order relation The more elements in ≠, the more timing constraints there are in the subtask execution process, which means the lower the possibility of task parallelism, which will lead to low efficiency in task allocation execution in unmanned clusters. Among all the partially ordered sets generated, the partially ordered set that contains the least timing relations, that is, the partially ordered set that only contains necessary timing relations is called a quasi-partially ordered set, denoted as P R .
[0012] 12A) Prune the automaton,
[0013] By removing decomposable transition edges, the complexity of the automaton is reduced. A decomposable transition edge is defined as a transition from state q in automaton B. i to q j The transition edge of is decomposable if and only if there is another state q k Satisfy:q k ∈δ(q i ,σ ik ∪σ kj )and q j ∈δ(q k , σ kj ), where σik Indicates that from state q i Transfer to q k The set of actions to be performed, σ kj Indicates that from state q k Transfer to q j The set of actions to be performed, δ is the state transition function of the automaton, δ(q i ,σ ik ∪σ kj ) means that in satisfying the action requirement σ ik ∪σ kj The set of states in the automaton that can be reached after pruning. After removing the decomposable transfer edges of the automaton, the automaton will eventually only retain the indecomposable edges, that is, the smallest task unit that cannot be decomposed and can promote the development of the overall task. The automaton obtained after pruning is denoted as B - .
[0014] 12B) Search from the pruned automaton to the final state q0 n ∈F, we get the trajectory ρ=q0q1…q n , n is the total number of states in the trajectory except the initial state. Count the action requirements of the corresponding edge of each transition of this trajectory ρ, and convert the trajectory ρ into a set of subtasks in is the subtask number, is the action of the subtask, where is the trajectory ρ status; is the trajectory ρ status; Indicated in State execution action After that, the state of the automaton will transfer to
[0015] 12C) Using independence analysis to statistically calculate the temporal relationship of the subtask set Ω ≠, where Represents subtask ω i Should be in subtask ω j Before execution, (ω i ,ω j )∈≠ represents the subtask ω i and subtask ω j Cannot be executed simultaneously. Calculate the timing relationship by repeatedly swapping the order of adjacent subtasks. Collect the required timing relationships into a quasi-partial order set
[0016] 2) Design an arbitrary-time poset multiplication algorithm to incrementally compute the product of the new quasi-poset and the current quasi-poset.
[0017] Two partially ordered sets and The product of is expressed as in, is a partially ordered set The partially ordered set Satisfies two conditions: Given a running word w0 (a word is a subtask sequence) ① If but Where w is a new subtask sequence, represents all subtask sequences that satisfy the partial order set P1, represents all subtask sequences that satisfy the partial order set P2, where Denote all those satisfying P′ i ② If but
[0018] Therefore, when executing linear timing tasks When a new linear time task formula When publishing online, to get the formula The partially ordered set no longer needs to be generated Nondeterministic Büchi automaton (NBA). represents the new linear temporal logic formula obtained by integrating the current task and the newly released task. We only need to generate and NBA to find The partially ordered set P1 and Then, a set of partially ordered sets satisfying the two conditions in the definition of the product of partially ordered sets is generated through P1 and P2. The present invention designs an arbitrary time poset product algorithm for calculating the product of the given time budget t b Generate at least one valid poset product element within the time frame, and generate more poset product elements that meet the conditions if time permits; including:
[0019] 21) Subtask fusion: Generate all possible subtask sets Ω that can simultaneously satisfy the subtask sets Ω1 and Ω2 of the two partial order sets of the input;
[0020] 211) First use the collection Record the product of the finally searched partial order set, and initialize the search subtask combination sequence que=[(Ω1,M Ω)], where Ω1 is the set of subtasks in the first poset P1, M Ω It is the task mapping function that maps the subtasks in Ω2 to a new subtask set, which is initially set to empty.
[0021] 212) When the search sequence is not empty, that is, |que|>0 and the search time has not been used up, that is, t <t b , t is the current search time; start calculating a set of subtask combinations:
[0022] 213) From the search sequence, select a set of subtask combinations, including the subtask set Ω′ and the task mapping function M′ Ω ;
[0023] 214) Try to select two subtasks that have not been selected before from the first partial order set P1 and the second partial order set P2 If they satisfy or in Represents a subtask The action set contains The action set, On the contrary, record This represents the subtask in Ω2 The new subtask If they are not satisfied or Record This represents the subtask in Ω2 The last element of the new subtask collection will be express.
[0024] 22) Partial order inheritance: Calculate the partial order relationship between each subtask in the new subtask set Ω and construct a partial order set as an element in the partial order product set.
[0025] 221) Find the combination sequence que of search subtasks that satisfies |D(M Ω )|=|Ω2| subtask combination (Ω,M Ω ), where D(M Ω ) represents the task mapping function M Ω The domain of |D(M Ω )|=|Ω2| means that the current subtask set Ω contains all subtask information in the subtask set Ω2;
[0026] 222) Remove (Ω,M from que Ω ), first calculate the "≤" constraint, the calculation formula is: Where ≤1 represents the ≤ constraint set in the first poset P1, and ≤2 represents the ≤ constraint set in the second poset P2. Indicates that the ≤ constraint set in P2 is passed through the task mapping function M Ω Map it to the new subtask set Ω, then take the union with the ≤ constraint set in P1 to form a new "≤" constraint, and update the calculation Ω according to the new partial order relation (≤). Then calculate the "≠" constraint, the calculation formula is: Where ≠1 represents the ≠ constraint set in the first poset P1, and ≠2 represents the ≠ constraint set in the second poset P2. Indicates that the ≠ constraint set in P2 is passed through the task mapping function M Ω Map it to the new subtask set Ω, and then take the union with the ≠ constraint set in P1 to form a new "≠" constraint. Finally, add the newly generated partial order set P = (Ω, ≤, ≠) to middle.
[0027] 223) Determine whether the search time exceeds the time budget. If t>t b , then stop the calculation, otherwise go back to step 221).
[0028] 224) Return
[0029] The poset product computation process is recursively applied to each pair of posets to compute the product for the set of formulas Φ b =φ1,φ2,…,φ n The complete set of partially ordered sets of . That is, there is no need to calculate the long formula φ b =φ1∧φ2∧…∧φ n , but incrementally calculates the product between the partial ordered sets of sub-formulas, so that the partial ordered set that satisfies the entire set of formulas can be calculated in polynomial time.
[0030] 3) Calculate the action-path chain of each subtask in the obtained quasi-partial order set to obtain the satisfying subtask ω provided by agent n i The action-path chain;
[0031] After obtaining the updated quasi-posset, the present invention calculates the complete action-path chain of each subtask. The agent model is the action transfer system of the agent. i =(i,σ i )∈Ω f , where i is the subtask number, σ i is the set of actions required by the subtask, specifically represented as a set of atomic propositions {a i,m}, where the atomic proposition a i,mIndicates that the agent needs to be in area W m Execute a i action. But for a But cannot perform the corresponding action a i The intelligent agent, that is It cannot directly perform action a i It must first perform other actions to move to a feasible state Similarly, when executing action a i Afterwards, if the current state Then additional actions are required to transfer back to the initial state However, due to resource constraints, these other actions may need to be performed in other areas. n and environment model T information, and can satisfy the subtask ω i The action-path sequence of n is called the action-path sequence that agent n can provide to satisfy the subtask ω i The action-path chain:
[0032] Input: Subtask ω i =(i,σ i ), Agent Model E n
[0033] 31) Initialize the search sequence in is the agent model E n The current state of π is the corresponding action sequence, σ is the proposition sequence, and the action chain set is initialized Action-Path Chain Set
[0034] 32) When the search sequence is not empty, start the following search
[0035] 321) Select a cell from the search sequence q and delete it from q s,π,σ=q.pop();
[0036] Among them, s represents a state in the agent model En, π represents the corresponding action sequence, and σ represents the atomic proposition sequence.
[0037] Starting from the current state s, search for all feasible actions a, that is, actions a satisfying (s, a, s′)∈→ n Starting from the current state, try to execute this action, record the chain of actions executed successively π′=[π,a], and the result of executing the action σ′=σ∪λ n (a), λ n (a) is the agent model E nThe corresponding result of executing action a. If the current execution result meets the requirements of the task And the state requirement s′=s0, then add π′ to the action chain set Otherwise it is added back to the search sequence.
[0038] 33) For action chain collection Each action chain π a , calculate the action-path chain corresponding to each action chain.
[0039] 331) Calculate the action chain π a Related areas in is the resource consumed to execute action a, R m It is area W m The type of resource owned.
[0040] 332) Generate a search sequence Where W m It is W i The relevant area in i All regions in the π are constructed as a temporary graph, and the movement is attempted in the search sequence. a [j] can be in W m Execute in: #Plan the next action, and if the action chain is completed, [π a ,(π a [j],W m )]join in That is, the action-path chain is found, otherwise (W m ,j+1,[π m ,(π a [j],W m )])Join q and continue searching.
[0041] 4) After obtaining the action chain corresponding to the subtask, a local search algorithm based on simulated annealing is designed to update the task sequence. Compared with the classic simulated annealing algorithm, we design a local search algorithm based on the calculation time t c The variable temperature attenuation coefficient and cutoff condition of the convergence rate are calculated; and a special neighborhood calculation is performed based on the structural characteristics of the action chain. The search target of the local search algorithm based on simulated annealing is to minimize the comprehensive cost of the overall execution plan of the task z = αt m +βc r +γc t , where t m is the execution time, c r is the resource consumption, c t is the trajectory consumption, and αβγ are the proportional coefficients.
[0042] 41) Variable decay coefficient: A decay coefficient related to the target value is set, so that the local search algorithm based on simulated annealing can spend more search time when there is an opportunity to obtain a better target value. Consider the current planned calculation time t c The optimization objective function is:
[0043] z'=α(t m +t c )+βc r +γc t =z+αt c . (1)
[0044] No. The attenuation coefficient can be written as:
[0045]
[0046] Among them, z * is the current optimal value of the objective function z', αt c is the weighted computation time consumption. Therefore, when the computation consumption is αt c When αt is relatively small, the temperature will remain high to maintain a high exploration ability; on the contrary, when αt c Relative z * When it is large, the temperature will decay rapidly, so that the objective function z' of the local search algorithm based on simulated annealing converges to the local optimal solution in the current area.
[0047] 42) Set dynamic cutoff conditions, algorithm cutoff:
[0048] When the time consumption of continued calculation is greater than the benefit of the potential objective function decrease, the algorithm terminates:
[0049]
[0050] in, is the current average descent rate, and the update method is is the computation time of the current iteration. Combining the above dynamic discrimination methods, the algorithm can show good adaptability between different optimization objectives.
[0051] 43) Neighborhood calculation:
[0052] A random unfinished subtask was removed Partial solution of Finally, considering the structure of the main agent and the collaborative agent, as well as the partially ordered set Order constraints in and conflict constraints ≠, random assignment of tasks will generate large
[0053] There are solutions to the problem of task order conflicts. Will bring new timing constraints, namely, subtask ω j2 Need to be in ω j1 Then execute. Therefore, first calculate due to The temporal relationship caused by the distribution relationship in the
[0054]
[0055] in, Task-based Updated order constraints; j1, j2 are subtask numbers; n is the agent number;
[0056] Then in the task Search for all timing constraints that meet The insertable order [(n,k)] is:
[0057]
[0058] Among them, i represents the task sequence J to be inserted n The subtask number of j represents the subtask ω i The subtask numbers with sequential constraints, k1, k2 represent subtask ω j The position index of the corresponding action-path chain in the overall task sequence.
[0059] After randomly selecting the agent number and the position index n, k where the action-path chain can be inserted into the task sequence, Select Action-Path Chain Insert to J n The kth position of , completes the action-path chain and the selection of the main agent. The action in finds the cooperative agent from [(n,k)], and before each insertion action, updates it using formula (4) Update [(n,k)] using equation (5).
[0060] 44) The overall algorithm is as follows
[0061] 441) Input: updated quasi-posset p, unassigned subtask set Ω un , the set of executed subtasks Ω finish
[0062] 442) Update Ω based on environmental information u Action-path chain
[0063] 443) Initialize temperature T and set the initial node to the solution of the previous round
[0064] 444) while dynamic cutoff condition (3) is not satisfied:
[0065] 445) If Cancel a chain from the assigned tasks and join Ω un
[0066] 446) from Ω un Randomly select tasks and find the neighborhood solution according to formula (4) (5)
[0067] 447) Determine whether to accept the new solution based on the temperature T condition.
[0068] 448) According to formula (2), variable rate temperature decay is performed
[0069] 449)return the optimal solution
[0070] 5) Design a mechanism for online synchronous execution of subtasks and perform online adaptation.
[0071] Since for an action-path chain like Is able to satisfy the subtask ω i Medium i The core action needs to follow the timing constraints ≠, where Represents subtask ω i Should be in subtask ω j Before execution, (ω i ,ω j )∈≠ represents the subtask ω i and subtask ω j Cannot be executed simultaneously. At the same time, if multiple agents participate in each action in the chain, it takes a continuous time period to complete. Due to the uncertainty of the environment, agents usually cannot strictly reach the target area to perform actions according to the calculated time. This will lead to synchronization of collaborative agents, or execution time and timing constraints ≠ Conflict. Therefore, we proposed an online synchronization mechanism based on the difference between the main and collaborative agents in the chain. The online synchronization method follows the following method:
[0072] 51) The master agent does not need to consider the status of the collaborative agents;
[0073] 52) The main agent needs to consider whether the partial order conditions of the core actions are met;
[0074] 53) The collaborative agent does not consider the partial order condition and needs to wait for the main agent to start executing before executing the corresponding task;
[0075] 54) If the corresponding task has been completed when the collaborative agent arrives, it will directly proceed to execute the next task.
[0076] Through the above steps, the dynamic allocation decision of complex model cluster tasks based on formal methods is realized, which ensures that the collaborative agents wait for the main agent and achieves the highest execution efficiency in the presence of uncertainties such as disturbances.
[0077] In specific implementation, the present invention stores a computer program on a computer-readable storage medium, and when the computer program is executed by a processor, the complex model cluster dynamic allocation decision method based on formal methods provided by the present invention can be implemented.
[0078] In specific implementation, the present invention also provides a system for realizing the above-mentioned cluster collaborative online task dynamic decision-making based on formal methods, including: a partial order fusion module, a reasoning module, a task allocation and an online synchronization module.
[0079] Among them, the partial order fusion module is responsible for continuously abstracting the LTL task formulas published online into quasi-partial order sets and then fusing them with the previous partial order sets, so as to obtain the temporal relationship between subtasks and subtasks necessary to satisfy the updated total LTL formula; the task reasoning module is responsible for combining subtasks and calculating the pre- and post-action requirements of subtasks under complex action models; the task allocation module is responsible for combining the action chain module calculated by the task reasoning module with the functions, distribution and environmental characteristics of the multi-agent cluster, and using the simulated annealing algorithm to calculate an efficient task plan that meets the characteristics of the quasi-partial order set and the agent, and outputs it; finally, the online module can design a master-slave structure based on the action chain to deal with the uncertainty in the environment.
[0080] Compared with the prior art, the present invention has the following beneficial effects:
[0081] The present invention proposes a cluster collaborative online task dynamic decision method and system based on formal methods. First, for an LTL formula, the temporal relationship between subtasks is extracted by calculating the partial order set of each subformula; then, an effective partial order product algorithm of any time is proposed, and the result obtained can not only completely retain the partial order relationship of each partial order set, but also solve the potential conflict relationship between subtasks. More importantly, this method can avoid directly converting a long string of LTL formulas into the dimensionality curse problem existing in NBA. The computational complexity of the first solution only increases linearly with the increase of the length of the task formula; then, combined with the complex action model of the intelligent body, the pre- and post-requirements required by the subtask are calculated, which is expressed as an action-trajectory chain to realize a subtask; finally, an online task allocation algorithm based on simulated annealing is proposed to generate an efficient task allocation scheme under the premise of satisfying complex constraints such as temporal constraints, collaborative constraints and interactive object constraints. The present invention is particularly suitable for efficient task planning in large-scale clusters with complex action models, dynamic environments and complex tasks, and can be applied to scenes such as multi-missile coordinated strikes, multi-robot coordinated transportation, and multi-UAV coordinated reconnaissance. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 The present invention provides an overall flow chart of the method.
[0083] Figure 2 This is a flowchart of the online adaptive algorithm. DETAILED DESCRIPTION
[0084] The present invention will be further described below by way of embodiments in conjunction with the accompanying drawings, but the scope of the present invention is not limited in any way.
[0085] The present invention proposes a cluster collaborative online task dynamic decision-making method and system based on formal methods. It includes: ① using the partial order multiplication method to convert online formal tasks and existing tasks into task sets; ② using the action chain method to calculate the pre- and post-task requirements of subtasks under complex models; ③ using the simulated annealing algorithm to allocate online tasks; ④ performing online adaptation according to environmental uncertainty. Specifically including:
[0086] Common cluster task planning problems are described as follows:
[0087] A team of N agents of L types operates in a two-dimensional workspace W. There are M regions in the workspace, represented as: Each agent Belongs to only one category l=M type (n), where M type : Each type of agent can provide an action set represented as All action sets are represented as At the same time, these agents follow a common transition graph to navigate between regions, which is represented as: in Transfers between regions are allowed, Map each transfer to the time it takes.
[0088] ① Mobility model of intelligent agent:
[0089] Represented as a weighted transition system (WTS) is a tuple Where W is the state set; is a transfer relation; AP is an atomic proposition defined in the next section; L t :W→2 AP is the labeling function; Ct: is the transfer cost function.
[0090] ②Action model of agent n:
[0091] is represented as an automaton Where S n is the state set, is the initial state, A n is the action set of agent n, → n :S n ×A n ×S n is the action transfer relation, λ n is the labeling function of the action, C A :W m ,A n →R m It is the resource cost function when executing an action, indicating that executing an action will consume a certain resource in the corresponding area.
[0092] In addition, there are some interactive objects in the workspace, and their collection is Category is They are interactive and can be carried from one area to another by the agent. is described as a triple: in is the type of object, t u ∈R + Yes u The time at which it appears in the workspace W, Is a returnable object o u At time t≥tu is a function of the region where it is located, so is the initial position of the object. It should be noted that new objects will appear in the workspace and be added to the collection over time. In. We use Indicates the initial existing object, using represents objects that appear online during execution, i.e.
[0093] To interact with objects, agents can provide a range of collaborative behaviors A collaborative behavior Defined as a tuple: in are the objects that the behavior interacts with (the empty set if none), is the action required for the behavior, 0 <n i ≤N is the number of executions required for the behavior i The number of agents taking actions, L k Behavior C k The index set of the required actions, d k Indicates behavior C k Execution time.
[0094] An action can be executed only if the object it requires is in the specified area. Since objects can only be carried by the agent, it is necessary for the planning process to find the correct order of these carrying actions.
[0095] For specific tasks, we define two types of atomic propositions: (i) The value is true when any agent of type l∈L is in the region middle; The value is true when any of the types is Objects in the area in (ii) The value is true when the collaborative behavior C k With objects Interaction from area Start to area End; Remember Given these atomic propositions, a team task can be formalized into a sc-LTL formula via {p,c}: in are two sets of sc-LTL formulas, It was scheduled in advance, When the object ou Added to Generated online due to real-time triggering.
[0096] To meet the LTL formula The complete action sequence of all agents is defined as: in yes sequence, which means that agent n at time t k By providing actions To participate in the execution of behavior At the same time, for interactive objects, their action sequences are defined as: in is with the object Related tasks sequence, when object o u At time t k Add Behavior When we will Add to Assuming the formula The time interval from release to completion is D i , the average efficiency is defined as: That is, the percentage of time the action is performed.
[0097] In summary, the problem can be described as: Given the sc-LTL formula Generate and update action sequences for all agents and the sequence of actions of objects Satisfy And maximize the execution efficiency η.
[0098] In view of the above problems, the present invention proposes the following decision-making method:
[0099] When a new task formula When it was released, LTL2BA was first used to convert the formula into a nondeterministic Büchi automaton (NBA), and then its partial order set was calculated. Then, the partial order set product was incrementally calculated to update the partial order set, so that the new partial order set not only contains the partial order relationship of each partial order set, but also resolves the potential conflict relationship between subtasks. The core algorithm partial order product mainly includes the following two steps:
[0100] ① Subtask fusion
[0101] In this step, we generate all possible subtask combinations that can satisfy both Ω1 and Ω2. If a subtask set Ω′={ω′1,…,ω′n′} satisfies Ω1, then for each There is a subtask satisfy and Or simply written as Because Ω′=Ω1 at the beginning, Ω′ obviously satisfies Ω1, that is, Next, a mapping M Ω :Ω2→Ω′ is defined to store the satisfied relation from Ω2 to Ω′, and It is M Ω The domain of It is M Ω Since all relations are unknown, M Ω , and Initially, they are set to empty. That is, if We will store as well as And will Add in ω′ j Add in in It is M Ω Then we used depth-first search (DFS) to construct a set of all possible subtasks. And the corresponding In this process we first check all possible combinations of undocumented subtasks See if it is satisfied or If it exists, we can use Ω′ and M′ to Ω Create a new collection of subtasks And a new mapping function as follows:
[0102]
[0103] This means that the subtask ω′ in Ω′ j Can be executed to satisfy
[0104] This step is until the time budget t b When it is exhausted or the search sequence que is exhausted, it ends. Once |D(M Ω )|=|Ω2|, which means that Ω′ that satisfies Ω1, Ω2 has been found. In this case, the next step of partial order inheritance is triggered.
[0105] ② Partial order inheritance
[0106] In this step, we consider the given subtask set Ω and mapping function M Ω Calculate the partial order relationship between them and construct a product partial order set P. First, we construct the "less than or equal to" constraint ≤ as follows:
[0107]
[0108] It inherits the less-than-or-equal-to relation ≤1 in P1 and the less-than-or-equal-to relation ≤2 in P2. Next, we update Ω to take into account the constraints imposed by the self-loop labels in other subtasks. Specifically, if a new relation (ω i ,ω j )pass was added≤, at the same time ω i Will be required to j Although previously executed does not belong to P1 ≤ 1. In this case, we will perform σ j Previously, i and Update to ensure that the self-loop label is satisfied For each subtask ω i , the set of newly added suf-subtasks from ≤1, ≤2 is defined as Right now:
[0109]
[0110] Among them are those that should be in ω i There is no requirement for the subtasks to be executed later, although the number of subtasks is ≤1 or ≤2. The action label σ in i and self-loop tags Will be updated as follows:
[0111]
[0112] in and σ i Should be in the additional tag Execute to satisfy We then examine a subtask Is it related to another subtask σ j Conflict and If such a situation exists, an additional ordering (ω i ,ω j ) will be added to ≤ and Ω will be updated as before. For those subtask sets Ω that do not conflict in ≤, we can use a simple combination such as to generate its "not equal to" relation. Finally, the resulting partial order set Added to middle.
[0113] The specific algorithm is as follows:
[0114] Posset Product Algorithm
[0115] Input: Partially ordered set P1 and Partially ordered set P2
[0116] Output: Product of partially ordered sets
[0117] que=[(Ω=Ω1,M Ω )]
[0118] When |que|>0 and t <t b :
[0119] Ω′,M′ Ω =que.pop()
[0120] Traversal
[0121] Traversal
[0122] Construct a new mapping function and set of subtasks And add it to que to build the original mapping function and subtask set And add it to que
[0123] Traverse(Ω,M Ω )∈que and they satisfy |D(M Ω )|=|Ω2|
[0124] Remove (Ω,M from que Ω ), construct a new partial order relation
[0125] Update Ω by self-loop update formula
[0126] Find potential constraints and update ≤,Ω according to these potential constraints.
[0127] If the newly generated subtask set has no conflict with the timing constraint ≤,Ω
[0128] Calculate the inequality constraint ≠ and add the newly generated partial order to
[0129] Update time t, if t <t b , then return
[0130] return
[0131] The above process is recursively applied to each pair of posets to calculate the equation Φ b That is, the first product calculated in the first round is Next, according to the efficiency metric, we start from Take the best partial order set P from final . It is related to the next Multiply, such as This process is repeated until the time budget is exhausted or all possible posets are found. The arbitrary time property of the poset is guaranteed, so we can quickly get a set that satisfies all P final And then continue to get more partial ordered sets.
[0132] The overall design process of the algorithm is as follows Figure 1 As shown, the design process is given below:
[0133] 1. Continuously publish formal language task formulas online The total task formula is
[0134] 2. Calculate the partial order set of each sub-formula
[0135] 3. Incremental calculation of poset products
[0136] 4. Use the time-constrained contract network algorithm to generate task allocation plans
[0137] 5. Execute tasks to achieve task formula
[0138] Action chain calculation
[0139] ①(Action-path chain) An action-path chain is a sequence of actions and positions. Indicates that the corresponding actions are executed in the region in order. The resources consumed by the actions match the environment model, that is, It can change the state of the intelligent model from The final transfer back to s0 is called an action chain. n All that can be provided to ω i The action chain set is recorded as The set of action-path chains is denoted as
[0140] To execute subtask ω i =(i,{ai,m},{})∈Ω f , the agent needs to be in region W m Execute an action i To achieve the task required atomic proposition a i,m But for a person in a state and An agent that cannot directly execute action a i It must first perform other actions to move to a feasible state Similarly, when executing action a i Afterwards, if the current state Then additional actions are required to transfer back to the initial state However, due to resource constraints, these other actions may need to be performed in other areas. n and environment model T information, and can satisfy the subtask ω i The action-path sequence of n is called the action-path sequence that agent n can provide to satisfy the subtask ω i The action-path chain:
[0141] Input: Subtask Agent Model E n
[0142]
[0143]
[0144] Given a subtask ω i , and the agent model E n , the above gives the calculation method of the corresponding action-path chain. First, search the action chain set in lines 2-9 Using deep search from the initial state Start by searching all action sequences. In lines 6-7, select the action sequence that not only satisfies the subtask but also returns the agent state to the initial state. The action sequence π′ is used as the action chain. Then in lines 10-22, the algorithm calculates the action-path chain set As shown in line 11, the action-path chain only considers regions with action-related resources. It searches for movement paths in lines 15-17 and tries to implement the action chain in lines 18-22. Action in.
[0145] Task Assignment:
[0146] After obtaining the updated quasi-partial order set, the present invention proposes a task assignment algorithm based on simulated annealing to assign these updated tasks to the intelligent agents so that all partial order relations are satisfied and the comprehensive cost of task completion is minimized. The present invention defines the four components of simulated annealing as follows:
[0147] ①Solution space
[0148] A feasible solution consists of a set of action sequences of agents. These action sequences take into account the corresponding environmental resource constraints, motion constraints, synchronization constraints and quasi-posteriori constraints. In addition, considering that the action-path chain has logical coherence, there is a master agent that fully executes all action sequence constraints. Therefore, the action sequence of agent n1 is It can be defined as:
[0149]
[0150] in is an element in the chain, indicating that agent n1 will be in area W k Execute action a k , to assist in completing the action-path chain of agent n2 This means that agent n1 completely executes the action chain The space consisting of all feasible solutions is the solution space.
[0151] ②Optimization goal
[0152] To evaluate the action sequence We designed three dimensions of indicators to evaluate the execution effect, the maximum completion time t m , resource cost c r and the agent trajectory and c t The optimization objective function is defined as
[0153]
[0154] where α, β and γ are the corresponding coefficients.
[0155] ③Neighborhood search
[0156] By making small changes to the tasks in the current solution, other solutions are obtained, which are called the domain of this solution. Partial solution of After that, considering the structure of the main and collaborative agents, and ≠ f The timing constraints of an agent, randomly assigning tasks will generate a large number of solutions with task order conflicts. Will bring new timing constraints, namely, subtask ω j2 Need to be in ω j1 Then execute. Therefore, first calculate due to The temporal relationship caused by the distribution relationship in the
[0157]
[0158] Then in Search for all timing constraints that meet The insertable order [(n,k)] is:
[0159]
[0160] After randomly selecting n, k, Select Action-Path Chain Insert to J n The kth position of , completes the action-path chain and the selection of the main agent. The action in finds the cooperative agent from [(n,k)], and before each insertion action, updates it using formula (11) Update [(n, k)] using equation (12).
[0161] ③ Temperature
[0162] The temperature T represents the probability of choosing a worse solution after obtaining the neighborhood. After obtaining a neighborhood solution, the current solution is accepted with probability P.
[0163]
[0164] The higher the temperature, the more the algorithm will tend to choose a worse solution to avoid falling into the local optimum. The lower the temperature, the more the algorithm will try to choose a better solution than the current solution to complete convergence.
[0165] ④ Variable attenuation coefficient
[0166] We set a decay coefficient related to the optimization value, so that the algorithm can show different exploration behaviors based on the expectation of obtaining a better value. Consider the current replanning computation time t re The optimization objective function is:
[0167] z'=α(t m +t c )+βc r +γc t =z+αt c .
[0168] No. The attenuation coefficient can be written as:
[0169]
[0170] Among them, z * is the current optimal value, αt c is the weighted computation time consumption. Therefore, when the computation consumption is αt c When αt is relatively small, the temperature will remain high to maintain a high exploration ability; on the contrary, when αt c Relative z * When it is large, the temperature will decay rapidly, causing the result to converge to the local optimal solution in the current area.
[0171] ⑤ Dynamic cutoff condition
[0172] When the time consumption of continued calculation is greater than the benefit of the potential objective function decrease, the algorithm terminates:
[0173]
[0174] in, is the current average descent rate, and the update method is is the computation time of the current iteration. Combining the above dynamic discrimination methods, the algorithm can show good adaptability between different optimization objectives.
[0175] Input: updated quasi-posset p, unassigned subtask set Ω u , the set of executed subtasks Ω finish
[0176] 1 Update Ω according to environmental information u Action-path chain
[0177] 2 Initialize temperature T and set the initial node to the solution J of the previous round
[0178] 3 while dynamic cutoff condition (10) is not satisfied:
[0179] 4 If From the assigned task, cancel a chain and join Ω un
[0180] 5 From Ω un Randomly select a task and find the neighborhood solution J′ according to equations (11) and (12).
[0181] 6. Determine whether to accept the new solution based on the temperature T condition.
[0182] 7 Variable rate temperature decay according to formula (9)
[0183] 8 return optimal solution
[0184] Online Adaptation
[0185] For an action-path chain like Is able to satisfy the subtask ω i Medium i The core action needs to follow the timing constraints ≠, where Represents subtask ω i Should be in subtask ω j Before execution, (ω i ,ω j )∈≠ represents the subtask ω i and subtask ω j Cannot be executed simultaneously. At the same time, if multiple agents participate in each action in the chain, it takes a continuous time period to complete. Due to the uncertainty of the environment, agents usually cannot strictly reach the target area and perform actions according to the calculated time. This will lead to synchronization of collaborative agents, or execution time and timing constraints ≠ f Therefore, we propose an online synchronization mechanism based on the distinction between the main and cooperative agents in the chain.
[0186] The specific implementation methods and implementation cases described above fully and completely realize the dynamic allocation decision of complex model clusters based on formal methods.
[0187] It should be noted that the purpose of publishing the embodiments is to help further understand the present invention, but those skilled in the art can understand that various substitutions and modifications are possible without departing from the scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the contents disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined in the claims.
Claims
1. A dynamic allocation decision method for unmanned clusters based on formal methods, characterized in that: The steps include: 1) Convert the formula representing the linear temporal logic task of the unmanned cluster into a quasi-partial order set; the quasi-partial order set includes the temporal relationship of the subtask sequence of the linear temporal logic task; 2) Design a partial order set multiplication algorithm, which does not need to calculate long formulas, but incrementally calculates the product of the quasi-partial order set of the unmanned cluster dynamic linear temporal logic task proposed online and the original quasi-partial order set, and obtains a quasi-partial order set that fully satisfies all unmanned cluster dynamic linear temporal logic tasks; the partial order set multiplication algorithm is used to generate at least one valid partial order set product within a given time budget, and generate the most partial order set products that meet the conditions; include: 21) Perform subtask fusion to generate all subtask sets Ω that can simultaneously satisfy the subtask sets Ω1 and Ω2 of the two quasi-possets; 22) Partial order inheritance: Calculate the partial order relationship between each subtask in the new subtask set Ω, and construct a partial order set as an element in the partial order product set; 3) Calculate the action-path chain of each subtask in the obtained quasi-partial order set to obtain a complete action trajectory chain provided by the agent that satisfies the subtask; including: Input subtasks and agents; subtask ω i =(i,σ i ), where i is the subtask number, σ i is the set of actions required by the subtask, expressed as a set of atomic propositions; 31) Initialize the search sequence in is the current state of the agent, π is the corresponding action sequence, σ is the atomic proposition sequence; and initialize the action chain set Action-Path Chain Set 32) When the search sequence is not empty, start the search as follows: 321) Select a unit s,π,σ=q.pop() from the search sequence; where s represents a state in the agent model, π represents the corresponding action sequence, and σ represents the atomic proposition sequence; Starting from the current state, search for all feasible actions; record the chain of actions executed successively π′=[π,a], and the result of executing the actions σ′=σ∪λ n (a), λ n (a) is the corresponding result of the agent executing action a; if the current result meets the requirements of the task And state requirements Then π′ is added to the action chain set, otherwise it is added back to the search sequence; 33) Calculate the action-path chain corresponding to each action chain in the action chain set; 4) Design a local search algorithm based on simulated annealing to update the task sequence, including: designing a variable attenuation coefficient and cutoff condition based on the calculation time and convergence rate; performing neighborhood calculation based on the structural characteristics of the action chain; searching with the minimum comprehensive cost of the overall task execution plan as the search goal; and obtaining the optimal task sequence; 5) Design a mechanism for online synchronous execution of subtasks and perform online adaptation; Through the above steps, the dynamic allocation decision of unmanned cluster tasks based on formal methods is realized, so that the collaborative agents wait for the main agent and the task execution efficiency is maximized.
2. The unmanned cluster dynamic allocation decision method based on formal methods as claimed in claim 1, characterized in that: Step 1) includes: The unmanned swarm linear temporal logic tasks are represented as non-deterministic Büchi automata. Prune the non-deterministic Büchi automaton to obtain the unmanned cluster subtask sequence set; The partial order relationship between subtasks in the task sequence is analyzed to obtain the quasi-partial order set of unmanned cluster subtasks with the least timing constraints.
3. The unmanned cluster dynamic allocation decision method based on formal methods as claimed in claim 1, characterized in that: Step 21) performing subtask fusion, including: 211) Use a collection Record the product of the finally searched partial order set, and initialize the search subtask combination sequence que=[(Ω1,M Ω )], where Ω1 is the set of subtasks in the first poset P1, M Ω It is the task mapping function that maps the subtasks in Ω2 to a new subtask set, which is initially set to empty; 212) When the search sequence is not empty, that is, |que|>0 and the search time has not been used up, that is, t <t b , t is the current search time; start calculating a set of subtask combinations: From the search sequence, select a set of subtask combinations, including the subtask set Ω′ and the task mapping function M′ Ω ; Try to select two subtasks that have not been selected before from the first partial order set P1 and the second partial order set P2 If satisfied or in Represents a subtask The action set contains The action set of That is, the subtask in Ω2 The new subtask express; like or If not satisfied, record That is, the subtask in Ω2 The last element of the new subtask collection will be express.
4. The unmanned cluster dynamic allocation decision method based on formal methods as claimed in claim 3, characterized in that: Step 22) Partial order inheritance includes: 221) Find the combination sequence que of search subtasks that satisfies |D(M Ω )|=|Ω2| subtask combination (Ω,M Ω ), where D(M Ω ) represents the task mapping function M Ω The domain of |D(M Ω )|=|Ω2| means that the current subtask set Ω contains all subtask information in the subtask set Ω2; 222) Remove (Ω,M from que Ω )back: First, calculate the "≤" constraint, the calculation formula is: Among them, ≤1 represents the ≤ constraint set in the first poset P1, and ≤2 represents the ≤ constraint set in the second poset P2. Indicates that the ≤ constraint set in P2 is passed through the task mapping function M Ω Map it to the new subtask set Ω, then take the union with the ≤ constraint set in P1 to form a new "≤" constraint, and update the calculation Ω according to the new partial order relation (≤); Then calculate the "≠" constraint, the calculation formula is: Among them, ≠1 represents the ≠ constraint set in the first poset P1, and ≠2 represents the ≠ constraint set in the second poset P2. Indicates that the ≠ constraint set in P2 is passed through the task mapping function M Ω Map it to the new subtask set Ω, and then take the union with the ≠ constraint set in P1 to form a new "≠" constraint; Finally, add the newly generated partial order set P = (Ω, ≤, ≠) to middle; 223) Determine whether the search time exceeds the time budget. If t>t b , then stop the calculation, otherwise return to step 212); 224) Return 5. The unmanned cluster dynamic allocation decision method based on formal methods as claimed in claim 4, characterized in that: The calculation method of step 33) includes: 331) Calculate the action chain π a Related areas in is the resource consumed to execute action a, R m It is area W m The type of resources owned; 332) Generate a search sequence Where W m It is W i The relevant area in W i All regions in the π are constructed as a temporary graph, and the movement is attempted in the search sequence. a [j] can be in W m Execute in: #Plan the next action, and if the action chain is completed, [π m ,(π a [j],W m )]join in That is, the action-path chain is found, otherwise (W m ,j+1,[π a ,(π a [j],W m )])Join q and continue searching.
6. The unmanned cluster dynamic allocation decision method based on formal methods as claimed in claim 5, characterized in that: Step 4) The local search algorithm based on simulated annealing includes: 41) Setting the variable attenuation coefficient related to the target value; No. The variable attenuation coefficient for: in, For the Searches; * is the current planned computing time t c The optimization goal is the current optimal value of the objective function, αt c is the weighted computation time consumption; 42) Set the dynamic cutoff condition for the search: When the time consumption of continued calculation is greater than the benefit of the potential objective function decrease, the algorithm terminates: in, is the current average descent rate, and the update method is is the computation time of the current iteration; 43) Neighborhood calculation: For an agent's task sequence J n , First, calculate the timing relationship in task allocation: in, Task-based Updated order constraints; j1, j2 are subtask numbers; n is the agent number; Then search for all tasks that satisfy the timing constraints The insertable order [(n,k)] is: Among them, i represents the task sequence J to be inserted n The subtask number of j represents the subtask ω i The subtask numbers with sequential constraints, k1, k2 represent subtask ω j The position index of the corresponding action-path chain in the overall task sequence; After randomly selecting n, k, from the action-path chain set Select Action-Path Chain Insert to J n The kth position of ; Chain in order The actions in the search for cooperative agents from [(n,k)] and before each insertion action, update and [(n,k)].
7. The unmanned cluster dynamic allocation decision method based on formal methods as claimed in claim 1, characterized in that: Step 5) The method of performing online adaptation includes: 51) The master agent does not consider the status of the collaborative agents; 52) The main agent considers whether the partial order conditions of the core actions are met; 53) The collaborative agents do not consider partial order conditions and wait for the main agent to start executing before executing the corresponding task; 54) If the corresponding task has been completed when the collaborative agent arrives, it will directly proceed to execute the next task.
8. A system using the complex model cluster dynamic allocation decision method based on formal methods as claimed in claim 1, characterized in that: The system includes: partial order fusion module, task reasoning module, task allocation and online synchronization module; among them, The partial order fusion module is responsible for continuously abstracting the LTL task formulas published online into quasi-partial order sets and then merging them with the previous quasi-partial order sets, so as to obtain the temporal relationship between subtasks and subtasks that meets the requirements of the updated total LTL formula; The task reasoning module is responsible for combining subtasks and calculating the pre- and post-action requirements of subtasks under complex action models; The task allocation module is responsible for combining the action chain module calculated by the task reasoning module with the functions, distribution and environmental characteristics of the multi-agent cluster, and using the simulated annealing algorithm to calculate an efficient task plan that meets the characteristics of the quasi-possessed set and the agent, and outputs it; The online synchronization module designs a master-slave structure based on the action chain to deal with the uncertainty in the environment.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the complex model cluster dynamic allocation decision method based on formal methods described in claim 1 is implemented.
Citation Information
Patent Citations
LTL-A*-A*optimal path planning method applicable to dynamic environment
CN106500697A
Collaborative task allocation method based on simulated annealing-spot scattering hybrid algorithm
CN113313360A