Multi-subject task cooperative scheduling and execution method and device, equipment and medium

By initially allocating capability characteristics and execution conditions in unmanned equipment collaborative operations, and dynamic iterative optimization during task execution, the optimization problem caused by unreasonable task allocation in the existing technology is solved, and efficient and globally optimal task collaborative execution is achieved.

CN119990580APending Publication Date: 2025-05-13BEIJING GUOKE FUNDAMENTAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411898019.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13

Smart Images

  • Figure CN119990580A_ABST
    Figure CN119990580A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-subject task cooperative scheduling and execution method and device, equipment and a medium, and the method comprises the steps: carrying out the preliminary distribution of cooperative tasks according to the capability characteristics and execution conditions of a plurality of unmanned devices, and obtaining a task-subject distribution scheme; in the task-main body allocation scheme, planning execution time windows for executing the corresponding tasks on the plurality of unmanned devices to obtain execution information; during the task execution period of the unmanned equipment based on the task-subject allocation scheme and the corresponding execution information, performing dynamic iterative optimization on the task allocation execution information according to the evaluation index, and continuing to perform task scheduling execution based on the optimized task allocation execution information; wherein the step of performing dynamic iterative optimization on the task allocation execution information comprises at least one of the following sub-steps: performing dynamic iterative optimization on the task-main body allocation scheme and performing dynamic iterative optimization on the execution information. And improvement of optimization iteration efficiency and realization of global optimization are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of collaborative task operations, and in particular to a method, device, equipment and medium for collaborative scheduling and execution of multi-agent tasks. Background Art

[0002] With the rapid development of unmanned technology, unmanned equipment (such as drones, unmanned vehicles, unmanned submersibles, search and rescue robots, handling robots, delivery robots, etc.) has shown great potential in many fields such as terrain exploration, disaster relief, environmental monitoring, logistics and distribution. However, a single unmanned equipment is often difficult to fully meet the complex and changing task requirements, so it is considered to carry out the collaborative operation of multiple unmanned equipment of the same type or the collaborative operation of multiple unmanned equipment of different types.

[0003] However, in the process of realizing the concept disclosed in the present invention, the following technical problems were found in the related technology: multiple or multiple unmanned devices mainly focus on simple information sharing and collaborative execution of tasks. Generally, after the tasks are initialized and assigned, the execution efficiency of each unmanned device in performing the corresponding tasks is mainly improved through optimization methods, and the irrationality of the initial task allocation is not considered. This results in a long optimization process and the optimization result may be a local optimum rather than an overall optimum. Summary of the invention

[0004] In order to overcome the problems existing in the related art, the present disclosure provides a method, apparatus, device and medium for collaborative scheduling and execution of multi-agent tasks.

[0005] According to the first aspect of the embodiment of the present disclosure, a method for collaborative scheduling and execution of multi-agent tasks is provided. The above method includes: performing preliminary allocation of collaborative tasks according to the capability characteristics and execution conditions of multiple unmanned devices to obtain a task-agent allocation scheme; in the above task-agent allocation scheme, planning the execution time windows for multiple unmanned devices to execute corresponding tasks to obtain execution information; during the period when the unmanned devices perform tasks based on the above task-agent allocation scheme and the corresponding execution information, dynamically iteratively optimizing the task allocation execution information according to the evaluation indicators, and continuing to schedule and execute tasks based on the optimized task allocation execution information; wherein, dynamically iteratively optimizing the task allocation execution information includes at least one of the following: dynamically iteratively optimizing the above task-agent allocation scheme and dynamically iteratively optimizing the above execution information.

[0006] In some embodiments, for at least one of the following situations, at least one of task viscosity and conflict information is also used as a factor affecting task allocation: preliminary allocation of collaborative tasks is performed based on the capabilities and execution conditions of multiple unmanned devices, or, based on evaluation indicators, task allocation execution information is dynamically iterated and optimized. Among them, the above-mentioned task viscosity represents the difference in allocation priority between tasks and allocation subjects due to the correlation between tasks. The task viscosity value corresponding to the subsequent associated tasks allocated to the subject that pre-allocated one or more associated tasks is greater than the task viscosity value corresponding to the subsequent associated tasks allocated to other subjects. The above-mentioned conflict information is used to indicate at least one of the following execution conflict situations: resource synchronization application conflicts when different unmanned devices apply for the same resource, spatial trajectory intersection conflicts when different unmanned devices perform their respective tasks, and resource shortage conflicts caused by the limited resources of unmanned devices when performing their own tasks.

[0007] In some embodiments, based on a genetic algorithm integrated with reinforcement learning, preliminary allocation of collaborative tasks and dynamic iterative optimization of the above-mentioned task-subject allocation scheme are performed. Among them, based on a genetic algorithm integrated with reinforcement learning, preliminary allocation of collaborative tasks includes: generating an initial population based on a policy network, each individual in the above-mentioned initial population represents a combination of unmanned equipment and an allocation scheme for collaborative task allocation under the combination; the input of the above-mentioned policy network includes: the current state of each unmanned equipment in the unmanned equipment set and the information of the task to be assigned; the above-mentioned current state includes: the current capability characteristics and the current execution conditions; some or all of the unmanned equipment in the above-mentioned unmanned equipment set constitute the above-mentioned unmanned equipment combination; the allocation scheme corresponding to the above-mentioned initial population is determined as the preliminary allocated task-subject allocation scheme.

[0008] In some embodiments, the above evaluation index includes: at least one of the improvement of the quality of task completion, the improvement of efficiency, the reduction of cost, and the avoidance of execution conflicts. According to the evaluation index, based on the genetic algorithm integrated with reinforcement learning, the above task-subject allocation scheme is dynamically iteratively optimized, including: based on reinforcement learning, taking at least one of the improvement of the quality of task completion, the improvement of efficiency, the reduction of cost, and the avoidance of execution conflicts as the reward function, the above policy network is updated so that the above policy network learns a better task allocation scheme, and the initial population is evolved based on the updated policy network; the above evolutionary processing includes a preset round of selection-crossover-mutation operations: based on the updated policy network, new individuals in the population are generated, and the above new individuals correspond to a better task allocation scheme; for the individuals in the initial population and the above new individuals, the target individual is selected as the parent individual based on the preset fitness function; the above parent individual is cross-operated to generate a new offspring individual; the new offspring individual is mutated based on the preset probability; for the individuals in the population after the evolutionary processing, the individual with the highest function value corresponding to the fitness function is determined as the optimized target task-subject allocation scheme.

[0009] In some embodiments, the fitness function satisfies the following expression:

[0010]

[0011] Resourceconsumptionj(x)+Matchingscore(x)+feedback(x),

[0012] Where Fitness(x) represents the fitness function of individual x; Taskcompletioni(x) represents the completion of the i-th task in the task allocation scheme corresponding to individual x; i represents the task number, which ranges from 1 to N, and N represents the total number of tasks; w i represents the weight corresponding to the time distribution matching degree of the i-th task; Resourceconsumptionj(x) represents the consumption of the j-th resource in the task allocation scheme corresponding to individual x; j represents the resource type number, ranging from 1 to M, and M represents the total number of resource types; c jRepresents the weight corresponding to the resource conflict avoidance rate of the j-th resource, and the above-mentioned resource conflict avoidance rate is determined based on the following execution conflict situations: resource synchronization conflict when different unmanned devices apply for the same resource and resource shortage conflict caused by the limited resources of the unmanned device when performing its own task; Matchingscore(x) represents the matching score between the capability characteristics of the unmanned device and the task requirements; feedback(x) represents the feedback input information of the allocation plan corresponding to individual x during the execution process, and the feedback input information includes at least one of task viscosity and conflict information; the above-mentioned conflict information includes: spatial trajectory intersection conflict when different unmanned devices perform their respective tasks.

[0013] In some embodiments, the above reward function satisfies the following expression:

[0014] Reward(s,a)=Fitness(a)+γ a′ maxFitness(s′,a′),

[0015] Among them, Reward(s,a) represents the reward function corresponding to the current environment state s and the current action a; the current environment state includes task requirements and resource conditions, and the current action includes at least one of the following: adjusting the correspondence between tasks and subjects, and adjusting the resource allocation for task execution; a′ represents the next action; s′ represents the next environment state; γ a′ Represents the discount factor corresponding to the next action a′, which is 0≤γ a′ ≤1; maxFitness(s′,a′) represents the maximum value of the fitness function corresponding to the next action a′ in the next environment state s′.

[0016] In some embodiments, the evaluation index includes at least one of: improving the quality of task completion, improving efficiency, reducing costs, and avoiding execution conflicts. Based on the dynamic programming method, the execution information is dynamically iterated and optimized according to the evaluation index, including: dividing the total time window for multiple unmanned devices to execute corresponding assigned tasks into multiple time periods; during the period when multiple unmanned devices execute tasks according to the execution information, for each current time period, dynamically planning is performed according to the execution information, the external environment, and the real-time status of the device to determine the execution information of the next time period; until the task execution of the total time window is completed in a cyclic iteration.

[0017] According to the second aspect of the embodiment of the present disclosure, a device for collaborative scheduling and execution of multi-agent tasks is provided. The above-mentioned device includes: a task allocation module, an execution time window planning module and a dynamic optimization module. The above-mentioned task allocation module is used to perform preliminary allocation of collaborative tasks according to the capability characteristics and execution conditions of multiple unmanned devices to obtain a task-agent allocation scheme. The above-mentioned execution time window planning module is used to plan the execution time windows for multiple unmanned devices to execute corresponding tasks in the above-mentioned task-agent allocation scheme to obtain execution information. The above-mentioned dynamic optimization module is used to dynamically iteratively optimize the task allocation execution information according to the evaluation index during the period when the unmanned device performs tasks based on the above-mentioned task-agent allocation scheme and the corresponding execution information, and continue to schedule and execute tasks based on the optimized task allocation execution information. Among them, the dynamic iterative optimization of the task allocation execution information includes at least one of the following: dynamic iterative optimization of the above-mentioned task-agent allocation scheme, dynamic iterative optimization of the above-mentioned execution information.

[0018] According to the third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the method for collaborative scheduling and execution of multi-subject tasks provided by the first aspect of the present disclosure.

[0019] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the method for collaborative scheduling and execution of multi-agent tasks provided in the first aspect of the present disclosure are implemented.

[0020] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0021] By making preliminary assignments of tasks taking into account the capability characteristics and execution conditions of unmanned equipment, it is helpful to make reasonable assignments based on the matching of task requirements and the execution capabilities of unmanned equipment, and to improve the optimization iteration efficiency; at the same time, in the case of a fixed assignment correspondence between tasks and unmanned equipment, the execution time window for the unmanned equipment to execute the corresponding tasks is planned to obtain execution information; and during the execution of the corresponding assigned tasks based on the execution information, the task assignment execution information is dynamically iterated and optimized according to the evaluation indicators, which helps to dynamically adjust the assignment correspondence between specific tasks and subjects according to actual conditions, or adjust the execution information (such as whether the task is executed, the task execution period, the resources required for task execution, the task execution progress, etc.), etc., so as to optimize the overall task collaborative execution, realize global dynamic optimization adapted to the real-time execution environment and execution status, and promote the improvement of evaluation indicators.

[0022] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0024] Figure 1 The present invention is a flowchart of a method for collaborative scheduling and execution of multi-agent tasks according to an exemplary embodiment.

[0025] Figure 2 It is a schematic diagram of a process of performing preliminary allocation of collaborative tasks and dynamically iterative optimization of the above-mentioned task-agent allocation scheme based on a genetic algorithm integrated with reinforcement learning according to an exemplary embodiment.

[0026] Figure 3 It is a schematic diagram of an execution process of dynamically iteratively optimizing task allocation execution information and continuing task scheduling execution based on the optimized task allocation execution information according to an exemplary embodiment.

[0027] Figure 4 The invention is a block diagram of a device for collaborative scheduling and execution of multi-agent tasks according to an exemplary embodiment.

[0028] Figure 5 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0029] Exemplary embodiments will be described in detail below with reference to the accompanying drawings.

[0030] It should be pointed out that the relevant embodiments and drawings are only for describing exemplary embodiments provided by the present disclosure, rather than all embodiments of the present disclosure, and it should not be understood that the present disclosure is limited to the relevant exemplary embodiments.

[0031] It should be noted that the terms "first", "second", etc. used in the present disclosure are only used to distinguish different steps, devices or modules, etc. The related terms neither represent any specific technical meanings nor indicate the order or interdependence between them.

[0032] It should be noted that the modifications of the terms "one", "multiple", and "at least one" used in the present disclosure are illustrative rather than restrictive. Unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0033] It should be noted that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. Unless otherwise specified, the scope of the present disclosure is not limited by the order of description of the steps in the relevant embodiments.

[0034] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the device is located and with the authorization given by the owner of the corresponding device.

[0035] Exemplary Methods

[0036] Figure 1 The present invention is a flowchart of a method for collaborative scheduling and execution of multi-agent tasks according to an exemplary embodiment.

[0037] Reference Figure 1 As shown, the first exemplary embodiment of the present disclosure provides a method for collaborative scheduling and execution of multi-agent tasks, and the method includes the following steps: S110, S120 and S130.

[0038] The method of this embodiment can be executed by an electronic device with computing capabilities, and the electronic device can communicate with multiple unmanned devices. The multiple unmanned devices can report real-time status to the above-mentioned electronic device, and can execute tasks according to the task-subject allocation plan and execution information updated by the above-mentioned electronic device.

[0039] In other application scenarios, on the premise of satisfying the computing resources and communication-related resources required for scheduling, the method of this embodiment can also be applied to one or more of the above-mentioned multiple unmanned devices. In this application scenario, priority is given to ensuring that each unmanned device has the resources required to perform the corresponding tasks.

[0040] In step S110, preliminary allocation of collaborative tasks is performed based on the capabilities and execution conditions of multiple unmanned devices to obtain a task-subject allocation plan.

[0041] In some application scenarios, multiple unmanned devices of the same type may work together, or multiple unmanned devices of different types may work together.

[0042] The above-mentioned unmanned equipment may include but is not limited to the following equipment: unmanned vehicles, drones, robots, cruisers, etc. The meaning of unmanned equipment is that the equipment can perform corresponding tasks based on pre-written control programs (including execution logic, algorithms for dealing with various conditions and dynamic adjustment logic, etc.), and no manual operation of the unmanned equipment is required during the execution of the task.

[0043] As an example, collaborative operation scenarios include but are not limited to: disaster search and rescue scenarios (for example, search and rescue robots, drones and unmanned vehicles can collaborate to detect, excavate and rescue people trapped in geological disasters, and transport and deliver materials), exploration scenarios (for example, unmanned mining vehicles can be used for mine exploration and material transportation in mining areas), cruise scenarios (for example, a group of cruisers or a cluster of cruisers and drones can collaborate to patrol and monitor the airspace or territorial waters), target tracking scenarios (for example, a cluster of drones tracks a moving object as the tracking object; or a fleet of unmanned vehicles takes the route of the lead vehicle as the target, and the following vehicles follow), logistics and distribution scenarios (for example, unmanned vehicles perform distribution; handling robots perform unloading and transportation, etc.), traffic management scenarios (for example, unmanned vehicles or drones monitor traffic conditions, and provide guidance and evacuation after traffic accidents), and environmental monitoring scenarios (for example, drone groups can be used for air quality monitoring, water quality monitoring, forest fire monitoring, etc.).

[0044] Collaborative tasks refer to tasks that require multiple subjects to complete together. Collaborative tasks can be tasks corresponding to various collaborative operation scenarios. Such as but not limited to: disaster search and rescue tasks, logistics delivery tasks, cargo handling tasks, cruise tasks, target tracking tasks, geological exploration tasks, air quality monitoring tasks, etc. The collaborative completion method can be: dividing the collaborative task into multiple subtasks, each subtask is assigned to a different subject for cooperation and execution; or, there is a situation in which the same task or the same task link in the collaborative task is completed by multiple subjects in collaboration.

[0045] Since the capabilities and execution conditions corresponding to different unmanned equipment have different emphases, the preliminary allocation of collaborative tasks is performed according to the capabilities and execution conditions of multiple unmanned equipment, so that the initial allocation correspondence between tasks and subjects is reasonable and can serve as a better initial solution, thereby improving the iterative efficiency of obtaining the global optimal solution during dynamic iterative optimization during the execution of subsequent tasks, and solving the problems in related technologies that random initial task allocation makes the optimization process time-consuming and the optimization result may be a local optimum rather than an overall optimum.

[0046] The capability characteristics of unmanned equipment can cover one or more of the following dimensions: the hardware and software configuration required to complete the task (computing power, communication capability, task processing algorithm configuration), electromechanical structure settings (flight capability, acceleration capability, endurance, loading capacity), specific function settings (anti-interference capability, reconnaissance or counter-reconnaissance capability, etc.), etc.

[0047] The execution conditions of unmanned equipment can cover one or more of the following dimensions: available resources of the equipment, available operating time, available space range, etc.

[0048] A collaborative task, for example, is a task set that includes multiple tasks, each of which has specific requirements, such as time constraints, location requirements, required payload, endurance requirements, processing power requirements, etc.

[0049] In some embodiments, the collaborative tasks can be preliminarily allocated based on some matching algorithms or adaptive learning algorithms to obtain a task-subject allocation scheme, which includes: an allocation correspondence between tasks and unmanned equipment.

[0050] As an example, the geological disaster rescue collaborative task T1 includes: disaster area reconnaissance task T11, trapped person rescue task T12 and material delivery task T13. The unmanned equipment group includes drones, unmanned vehicles and robots, among which drone A has strong reconnaissance capability (an example of capability characteristics, similar in the future, no further explanation) and is currently in an idle and disposable state (an example of execution conditions, similar in the future, no further explanation); unmanned vehicle B has strong load-bearing capacity and endurance, and is in an idle and disposable state; drone C has strong endurance and is in an idle and disposable state; drone D has strong load-bearing capacity and is in a transporting state; robot E has strong rescue capability and is in an idle and disposable state; unmanned vehicle E has strong load-bearing capacity and endurance, and is currently in an occupied state; drone F has strong load-bearing and endurance, and is currently in an occupied state.

[0051] Based on a matching algorithm or an adaptive learning algorithm, tasks T11 to T13 in the above-mentioned collaborative task T1 are distributed among all or part of the devices in the unmanned equipment group. For example, an exemplary task-subject allocation scheme FP1 is: the disaster area reconnaissance task T11 is assigned to drone A; the trapped person rescue task T12 is assigned to robot E; and the material delivery task T13 is assigned to unmanned vehicle B and drone C.

[0052] In step S120, in the above-mentioned task-subject allocation scheme, the execution time windows for executing corresponding tasks of multiple unmanned equipment are planned to obtain execution information.

[0053] In some embodiments, in the above step S120, in the above task-subject allocation scheme, the execution time windows for executing corresponding tasks of multiple unmanned devices are planned to obtain execution information, including:

[0054] For multiple unmanned devices, single-agent time planning is performed based on their respective available time windows and the estimated duration of the corresponding assigned tasks to obtain the execution time window of each unmanned device for executing the corresponding assigned tasks; the multiple unmanned devices and the corresponding assigned tasks are determined by the task-agent allocation scheme;

[0055] Execution information is generated based on the execution time window of each unmanned device for executing the corresponding assigned task and the resources required for task execution.

[0056] For example, after the initial allocation of collaborative tasks to obtain the corresponding relationship between tasks and subjects, single-subject time planning is performed for the unmanned equipment assigned with tasks, and the execution time window for each unmanned equipment to execute the corresponding assigned tasks is determined. At the same time, the resources required for task execution are also determined, and execution information is generated based on the execution time window and the resources required for task execution. In the initial state, the execution information includes: execution time distribution information (specifically, the overall distribution information of the execution time window for the unmanned equipment assigned with tasks to execute the corresponding assigned tasks) and execution required resource information.

[0057] For example, the execution information ZX1 corresponding to the task-subject allocation scheme FP1 includes: execution time distribution information FB1 and execution required resource information ZY1; the execution time distribution information FB1 is, for example: the disaster area reconnaissance task T11 is executed by drone A during the time period t0-t9, and it is estimated that the specific locations of most of the people to be rescued can be searched within the time period t0-t1; the trapped person rescue task T12 is executed by robot E during the time period t2-t3 (t2 is after t1); the material delivery task T13 is executed collaboratively by unmanned vehicle B and drone C during the time period t4-t5 (t4 is after t1), wherein unmanned vehicle B and drone C respectively share the delivery of materials in different geographical areas.

[0058] Resource information required for task execution includes, but is not limited to: hardware and software resources (such as CPU, GPU, SOC, MUC, disk memory, mounted external storage, etc.), functional device configuration (such as search and rescue radar, anti-interference radar, etc.), required computing power resources, required software and algorithm resources, required power resources (such as engine horsepower, fuel reserves, power, etc.), task-related external mounts (such as first aid supplies), etc.

[0059] Planning the execution time window to obtain the execution information in the task-subject allocation plan obtained by the preliminary allocation of collaborative tasks is an operation when the task has not yet been executed. It belongs to the pre-planning process and can be used as the initial solution for the task allocation execution information in the subsequent step S130. The task is executed based on this initial solution and the task allocation execution information is dynamically iterated and optimized according to the actual execution environment and actual conditions of the task.

[0060] In step S130, while the unmanned equipment is executing tasks based on the above-mentioned task-subject allocation scheme and the corresponding execution information, the task allocation execution information is dynamically iteratively optimized according to the evaluation indicators, and the task scheduling execution is continued based on the optimized task allocation execution information; wherein, the dynamic iterative optimization of the task allocation execution information includes at least one of the following: dynamic iterative optimization of the above-mentioned task-subject allocation scheme and dynamic iterative optimization of the above-mentioned execution information.

[0061] In some embodiments, the evaluation index includes at least one of: improving the quality of task completion, improving efficiency, reducing costs, and avoiding execution conflicts. Various known algorithms or improved algorithms can be used for dynamic iterative optimization.

[0062] For example, during the task execution based on the execution time distribution information FB1 and the execution required resource information ZY1 corresponding to the task-subject allocation scheme FP1, since the actual environment and execution status are dynamically changing, the task allocation execution information is dynamically iteratively optimized according to the evaluation indicators and the task scheduling execution is continued based on the optimized task allocation execution information. This helps to adapt to the real-time execution environment and execution status to achieve global dynamic optimization, and realize at least one of improved quality, improved efficiency, reduced costs, or avoidance of execution conflicts in task completion.

[0063] As an example, during the execution of the task, the execution conditions of drone D have changed from being in a transporting state to an idle and available state. Then, with the goal of improving the efficiency of task completion, the task-subject allocation plan FP1 is dynamically iteratively optimized, and the optimized task-subject allocation plan FP2 adds a corresponding allocation relationship between drone D and the material delivery task T13. During the execution of the task, according to trajectory prediction analysis, drone A has a trajectory intersection conflict with drone C that performs the material delivery task T13 during the period t4-t5 when performing the disaster area reconnaissance task T11. In order to avoid this conflict, the result of the dynamic iterative optimization of the execution information is, for example: the period t4-t5 when drone C performs the material delivery task T13 is adjusted to a period when there is no trajectory conflict with drone A, or a partial route of at least one of drone A or drone C is changed to prevent the two from colliding.

[0064] In the embodiment including the above steps S110 to S130, by considering the capability characteristics and execution conditions of the unmanned equipment for preliminary task allocation, it is helpful to make reasonable corresponding allocations based on the matching of task requirements and the execution capabilities of the unmanned equipment, and to improve the optimization iteration efficiency; at the same time, in the case of a fixed allocation correspondence between tasks and unmanned equipment, the execution time window for the unmanned equipment to execute the corresponding tasks is planned to obtain execution information; and during the execution of the corresponding assigned tasks based on the execution information, the task allocation execution information is dynamically iterated and optimized according to the evaluation indicators, which helps to dynamically adjust the allocation correspondence between specific tasks and subjects according to actual conditions, or adjust the execution information (such as whether the task is executed, the task execution period, the resources required for task execution, the task execution progress, etc.), etc., so that the overall task collaborative execution is optimized, and global dynamic optimization adapted to the real-time execution environment and execution status is achieved, thereby promoting the improvement of evaluation indicators.

[0065] In some embodiments, during the process of preliminary task allocation and dynamic iterative optimization, in addition to considering the matching of the capabilities and execution conditions of the unmanned equipment with the tasks, the influence of factors such as task viscosity and conflict information are also considered.

[0066] Among them, task viscosity is a concept proposed in the embodiment of the present disclosure. The meaning of the above task viscosity is: it indicates the difference in allocation priority between tasks and allocation subjects due to the association between tasks. The task viscosity value of the subject to which one or more associated tasks are pre-assigned corresponding to the subsequent associated tasks is greater than the task viscosity value corresponding to the subsequent associated tasks assigned by other subjects. For example, there are four tasks X, Y, Z, and R, and the subjects are subjects 1 to 4 respectively. There is an execution association between task X and task Y. Task X is assigned to subject 1. Then, when allocating task Y, due to the association between task Y and task X, task Y is assigned to subject 1 first. That is, when determining which subject among the four subjects 1 to 4 is assigned to task Y, since subject 1 has been assigned task X, the task viscosity value of subject 1 corresponding to task Y is greater than the task viscosity value of subject 2 to 4 corresponding to task Y. In this way, the execution efficiency and execution quality of the previous and subsequent tasks can be improved, or the transmission cost or information synchronization cost required for the execution of associated tasks in different subjects can be saved.

[0067] The above-mentioned conflict information is used to indicate but is not limited to at least one of the following execution conflict situations: resource synchronization conflict when different unmanned devices apply for the same resource, spatial trajectory intersection conflict when different unmanned devices perform their respective tasks, resource shortage conflict caused by the limited resources of the unmanned device when performing its own tasks, etc.

[0068] Therefore, for at least one of the following situations: the preliminary allocation of collaborative tasks based on the capability characteristics and execution conditions of multiple unmanned devices, or the dynamic iterative optimization of task allocation execution information based on evaluation indicators, at least one of task viscosity and conflict information is used as a factor affecting task allocation.

[0069] For example, if UAV A completes the geographic reconnaissance mission first, it is beneficial to perform the rescue mission next, because UAV A has already obtained geographic information through the reconnaissance mission, so it does not need to obtain geographic information from other UAVs or the cloud for destination search and route planning of the rescue mission. When allocating subsequent rescue missions, from the perspective of mission viscosity, giving priority to UAV A to perform the rescue mission next can reduce the resource occupation and consumption caused by information transmission and analysis between different terminals during the information synchronization process, saving information synchronization costs.

[0070] For related tasks that have contextual or logical associations, allocating the correspondence between tasks and subjects based on task viscosity can help improve the efficiency and quality of execution when the related tasks are executed.

[0071] With respect to the process of preliminary task allocation and the process of dynamic iterative optimization, the disclosed embodiment also provides a highly efficient initialization and dynamic optimization algorithm.

[0072] Figure 2 It is a schematic diagram of a process of performing preliminary allocation of collaborative tasks and dynamically iterative optimization of the above-mentioned task-agent allocation scheme based on a genetic algorithm integrated with reinforcement learning according to an exemplary embodiment. Figure 3 It is a schematic diagram of an execution process of dynamically iteratively optimizing task allocation execution information and continuing task scheduling execution based on the optimized task allocation execution information according to an exemplary embodiment.

[0073] In some embodiments, in combination Figure 2 and Figure 3 As shown, in the above steps S110 and S130, based on the genetic algorithm integrated with reinforcement learning, the preliminary allocation of collaborative tasks and the dynamic iterative optimization of the above task-agent allocation scheme are performed.

[0074] For example, refer to Figure 2 As shown, in the above step S110, based on the genetic algorithm integrated with reinforcement learning, the preliminary allocation of collaborative tasks includes:

[0075] Based on the policy network 210, an initial population is generated, each individual in the initial population represents an unmanned equipment combination and an allocation scheme for collaborative task allocation under the combination; the input of the policy network includes: the current state of each unmanned equipment in the unmanned equipment set and the information of the task to be assigned; the current state includes: the current capability characteristics and the current execution conditions; some or all of the unmanned equipment in the unmanned equipment set constitute the unmanned equipment combination;

[0076] The allocation scheme corresponding to the above initial population is determined as the preliminary allocated task-subject allocation scheme.

[0077] For example, refer to Figure 2 As shown, the generated initial population is indicated by a single-point dashed box, and the individuals in the initial population are indicated by triangles. The scheme corresponding to a randomly selected individual in the above initial population or the scheme corresponding to an individual with a higher ability characteristic matching degree (for example, corresponding to the Matchingscore(x) in the subsequent formula (1-1) or formula (1-2)) (for example, the allocation scheme corresponding to the second individual) can be used as the preliminary assigned task-subject allocation scheme, for example, the task-subject allocation scheme FP1 in the aforementioned example is obtained.

[0078] In the above reinforcement learning-based genetic algorithm, the problem definition and modeling are first performed. The capability characteristics of the unmanned equipment (such as different acceleration capabilities, deflection capabilities, endurance capabilities, load capacity, etc.) and execution conditions are defined; task requirements are defined: each task has specific requirements, such as time limits, location requirements, required loads, etc.; and optimization goals are defined, such as: maximizing task completion efficiency, minimizing resource consumption, etc. The optimization goals are reflected in the evaluation indicators of the subsequent step S130.

[0079] When generating the initial population, binary or real number coding is used to represent each individual, where each gene bit represents whether an unmanned device is assigned to a certain task. The task type and functional characteristics are encoded as chromosomes, and each chromosome represents a possible task-subject allocation scheme. A new generation of populations and corresponding chromosomes are generated through operations such as selection, crossover, and mutation. Among them, the quality of each chromosome is evaluated based on the fitness function, and the fitness function is generally designed based on the optimization goal. For example, the above fitness function is a weighted sum of at least one of the following evaluation indicators: quality improvement of task completion, efficiency improvement, cost reduction, and avoidance of execution conflicts.

[0080] Continue to refer to Figure 2 and Figure 3 As shown, in the above step S130, according to the evaluation index, based on the genetic algorithm integrated with reinforcement learning, the above task-subject allocation scheme is dynamically iteratively optimized, including:

[0081] Based on reinforcement learning, the policy network 210 is updated with at least one of the following reward functions: quality improvement, efficiency improvement, cost reduction, and avoidance of execution conflicts in task completion, so that the policy network learns a better task allocation scheme, and the initial population is evolved based on the updated policy network. Figure 2 The angular arrows are used to indicate the process of updating the policy network and the process of evolving the initial population based on the updated policy network, and the double-dotted line frame is used to indicate the evolved population. In order to indicate the difference between the individuals obtained after multiple rounds of iterations and the individuals in the initial population, the individuals in the evolved population are indicated by quadrilaterals, pentagons and hexagons.

[0082] The above evolution process includes a preset round of selection-crossover-mutation operations:

[0083] Generate new individuals in the population based on the updated policy network, and the new individuals correspond to a better task allocation solution;

[0084] For the individuals in the initial population and the new individuals, the target individuals are selected as parent individuals based on the preset fitness function;

[0085] Perform a crossover operation on the above-mentioned parent individuals to generate new offspring individuals;

[0086] Perform mutation operations on new offspring individuals based on preset probabilities;

[0087] For the individuals in the population after evolution, the individuals with the highest function value corresponding to the fitness function are determined as the optimized target task-subject allocation scheme, for example, Figure 2 As shown, the individuals corresponding to the hexagons are taken as examples of individuals with the highest values ​​of the fitness function.

[0088] Since reinforcement learning is integrated into the genetic algorithm, the optimization of individuals in the population is achieved through reinforcement learning of the policy network, thereby obtaining the optimal task-agent allocation solution.

[0089] In some embodiments, reference Figure 2 As shown, a policy network 210 is constructed to generate a preliminary plan for unmanned system task allocation. The policy network can be constructed based on, but not limited to, the following neural networks: deep neural network (DNN), convolutional neural network (CNN) or recurrent neural network (RNN), etc. The input of the above policy network includes: the current state of each unmanned device in the unmanned device set and the task information to be assigned; the output includes: the unmanned device combination and one or more allocation schemes for collaborative task allocation under the combination (for example, the task allocation and functional characteristic configuration scheme for the unmanned device).

[0090] In some embodiments, for example Figure 2 The dotted box is used to illustrate the influencing factors of task allocation. The above-mentioned strategy network is used to use the capability characteristics, execution conditions and conditions required for task completion of the unmanned equipment as the influencing factors of task allocation; in some embodiments, at least one of task viscosity and conflict information can also be used as the influencing factors of task allocation.

[0091] Design reward function: reward is given according to the optimization goal (taking task completion and resource consumption as examples). For example, high rewards are given when task completion is high and resource consumption is low. The policy network selects actions (such as selecting tasks, adjusting speed, load, etc.) according to the current environment state (such as task requirements, resource conditions, etc.); and after performing an action, it receives a reward value from the environment feedback (set based on the reward function, which can be a positive reward or a negative penalty, so that the policy network makes decisions on corresponding actions in the direction of the optimization goal), and updates and determines subsequent actions based on this reward value, so that the rewards corresponding to all action paths tend to be maximized.

[0092] The training process of the policy network: By interacting with the environment (i.e., the set of corresponding relationships between unmanned equipment and task allocation), data is collected and the policy network is trained so that it gradually learns to generate a better task allocation solution. The value function is updated using a gradient-based update method (such as Q-learning, deep Q network, etc.). The specific decision strategy π(s) is updated based on the reward signal r provided by the environment and the transfer of the environment state s.

[0093] For example, in some embodiments, the capability characteristics, execution conditions, and conditions required for task completion of unmanned equipment are used as factors affecting task allocation, and the fitness function in the genetic algorithm based on reinforcement learning satisfies the following expression:

[0094]

[0095] Where Fitness(x) represents the fitness function of individual x; Taskcompletioni(x) represents the completion of the i-th task in the task allocation scheme corresponding to individual x; i represents the task number, which ranges from 1 to N, and N represents the total number of tasks; w i represents the weight corresponding to the time distribution matching degree of the i-th task; Resourceconsumptionj(x) represents the consumption of the j-th resource in the task allocation scheme corresponding to individual x; j represents the resource type number, ranging from 1 to M, and M represents the total number of resource types; c jRepresents the weight corresponding to the resource conflict avoidance rate of the j-th resource. The resource conflict avoidance rate is determined based on the following execution conflict situations: resource synchronization conflict when different unmanned devices apply for the same resource and resource shortage conflict caused by the limited resources of the unmanned device when executing its own task; Matchingscore(x) represents the matching score between the capability characteristics of the unmanned device and the task requirements.

[0096] In the case where at least one of task viscosity and conflict information is used as a factor affecting task allocation, at least one of the task viscosity and conflict information is also used as an input of the fitness function to affect an output result of the fitness function.

[0097] For example, in other embodiments, not only the capability characteristics, execution conditions and conditions required for task completion of unmanned equipment are considered, but also at least one of task viscosity and conflict information is used as a factor affecting task allocation. Then, the fitness function in the genetic algorithm based on reinforcement learning satisfies the following expression:

[0098]

[0099] Where Fitness(x) represents the fitness function of individual x; Taskcompletioni(x) represents the completion of the i-th task in the task allocation scheme corresponding to individual x; i represents the task number, which ranges from 1 to N, and N represents the total number of tasks; w i represents the weight corresponding to the time distribution matching degree of the i-th task; Resourceconsumptionj(x) represents the consumption of the j-th resource in the task allocation scheme corresponding to individual x; j represents the resource type number, ranging from 1 to M, and M represents the total number of resource types; c j Represents the weight corresponding to the resource conflict avoidance rate of the j-th resource, and the above-mentioned resource conflict avoidance rate is determined based on the following execution conflict situations: resource synchronization conflict when different unmanned devices apply for the same resource and resource shortage conflict caused by the limited resources of the unmanned device when performing its own task; Matchingscore(x) represents the matching score between the capability characteristics of the unmanned device and the task requirements; feedback(x) represents the feedback input information of the allocation plan corresponding to individual x during the execution process, and the feedback input information includes at least one of task viscosity and conflict information; the above-mentioned conflict information includes: spatial trajectory intersection conflict when different unmanned devices perform their respective tasks.

[0100] In some embodiments, in the fitness function of the above embodiments (for example, the fitness function corresponding to formula (1-1) and formula (1-2)), the weight w corresponding to the time distribution matching degree is iIn the first round of iteration, the initial value is preset; in the subsequent iterations, the value is taken from the result of iterative optimization of the execution information. For example, the corresponding weight of the time distribution matching degree can be assigned according to the result of iterative optimization using a subsequent dynamic programming method.

[0101] In some embodiments, in the fitness function of the above embodiments (for example, the fitness function corresponding to formula (1-1) and formula (1-2)), the weight c corresponding to the resource conflict avoidance rate is j It can be pre-set according to the actual situation of the task. For example, if the same resource is required by multiple tasks, the weight of the resource conflict avoidance rate under the task allocation scheme for the same resource corresponding to the execution entities of different tasks is determined based on at least one dimension such as the task importance, task urgency, or whether the resources required for the task are adjustable for multiple different tasks. Generally speaking, for the same resource, due to differences in the task importance, task urgency, or whether the resources required for the task are adjustable for the tasks that require the resource, the weight of the resource conflict avoidance rate corresponding to the same resource will be different when the same resource is allocated to the execution entities of different tasks, that is, when corresponding to different task allocation schemes.

[0102] As an example, if a resource ZY1 is required by three tasks 1 to 3 at the same time, the three tasks 1 to 3 have different task importance levels, for example, task 1 is very important, task 2 is generally important, and task 3 is generally important, and the resources required by tasks 2 and 3 can be adjusted, while the resources required by task 1 cannot be adjusted. By considering these two dimensions, it can be determined that tasks 2 and 3 can avoid competing with task 1 for resource ZY1 by adjusting resources. Then, for the task allocation scheme that allocates resource ZY1 to the corresponding execution subject of task 1, the value of the resource conflict avoidance rate is expressed as: x1-c ZY1 ; Among them, x1 represents the task allocation scheme corresponding to allocating resource ZY1 to the execution subject of task 1, c ZY1 The weight corresponding to the resource conflict avoidance rate of resource ZY1 is x1-c ZY1 It must be smaller than the resource conflict avoidance rate in the task allocation scheme corresponding to the allocation of resource ZY1 to task 2 (or task 3), expressed as x2-c ZY1 ; Among them, x2 represents the task allocation scheme corresponding to allocating resource ZY1 to the execution subject of task 2 or task 3, that is, by setting the value of resource conflict avoidance rate to have the following relationship: x1-c ZY1 <x2-c ZY1 , corresponding to Fitness(x1)>Fitness(x2), which helps to improve the fitness value of the allocation scheme corresponding to resource ZY1 allocated to task 1 during dynamic iterative optimization of the task allocation scheme, thereby increasing the probability of resource ZY1 being allocated to task 1.

[0103] In some embodiments, the above reward function satisfies the following expression:

[0104] Reward(s,a)=Fitness(a)+γ a′ maxFitness(s′,a′), (2)

[0105] Among them, Reward(s,a) represents the reward function corresponding to the current environment state s and the current action a; the current environment state includes task requirements and resource conditions, and the current action includes at least one of the following: adjusting the correspondence between tasks and subjects, and adjusting the resource allocation for task execution; a′ represents the next action; s′ represents the next environment state; γ a′ Represents the discount factor corresponding to the next action a′, which is 0≤γ a′ ≤1; maxFitness(s′,a′) represents the maximum value of the fitness function corresponding to the next action a′ in the next environment state s′. The Fitness() function in formula (2) can use formula (1-1) or formula (1-2).

[0106] In some embodiments, reference Figure 3 As shown, an elliptical frame is used to illustrate a dynamic programming method (specifically, it may be a time distribution optimization algorithm), and an arrow is used to illustrate a process of dynamically iterating and optimizing execution information based on the dynamic programming method.

[0107] In the above step S130, based on the dynamic programming method and according to the evaluation index, the above execution information is dynamically iterated and optimized, including:

[0108] Divide the total time window for multiple unmanned devices to perform corresponding assigned tasks into multiple time periods;

[0109] While multiple unmanned devices are executing tasks according to the above execution information, dynamic planning is performed for each current time period based on the execution information, external environment and real-time status of the equipment to determine the execution information for the next time period; until the loop iteration completes the task execution of the total time window.

[0110] For each current time period, dynamic planning is performed based on the execution information, external environment, and real-time status of the device to determine the execution information for the next time period, including:

[0111] For the current time period, obtain current status information; the current status information includes: real-time status of the device, task execution status and external environment; the task execution status includes the execution information, and the execution information includes one or more of the following: whether the task is executed, the task execution period, the resources required for task execution, and the task execution progress; the external environment may include one or more of the following: wind speed, the presence or absence of obstacles, temperature, humidity, visibility, the presence or absence of interference, etc.;

[0112] Based on the decision variables, the current state information in the current time period is processed to generate multiple predicted state information for the next time period, and the subsequent time period decisions are recursively made based on the multiple predicted state information to generate the final predicted state information;

[0113] According to the evaluation index, determine the optimal decision strategy that satisfies the evaluation index or maximizes the benefit of the evaluation index from the decision strategies composed of the predicted state information of the future time period after the current time period;

[0114] The task execution status in the optimal prediction status information of the next time period in the above optimal decision strategy is determined as the execution information of the next time period.

[0115] The above dynamic programming method can be used to optimize the execution information as a time distribution optimization algorithm, which mainly optimizes the time distribution of task execution according to the time requirements and resource constraints of the task. The optimization goal can be one or a combination of the following: minimizing the task completion time, reducing energy consumption costs, etc.

[0116] The objective function of the time distribution optimization algorithm may involve multiple variables and constraints, such as task start time, task duration, and resource constraints. Through the optimization iteration of the time distribution optimization algorithm, the time for each unmanned device to perform the corresponding assigned task is optimized by global dynamic programming to ensure that multiple unmanned devices can perform different tasks in different time periods to achieve at least one of the following optimization goals: avoid resource conflicts and task overlaps, minimize task completion time, reduce energy consumption costs, and evenly distribute tasks among multiple agents.

[0117] The above dynamic planning method has the ability to adaptively adjust and can dynamically adjust the time plan according to the actual environment and execution status during the task execution (such as weather changes, traffic conditions, route conflicts, resource conflicts, etc.). At the same time, by obtaining the status information of unmanned equipment and the task execution status in real time, it can make adjustments quickly to ensure the continuity and efficiency of task execution.

[0118] In some embodiments, a dynamic programming problem can be modeled. For example, (a) define unmanned devices and tasks: each unmanned device has its specific capabilities and states, and each task has a specific time window, resource requirements, and execution conditions. The entire time range is divided into a series of discrete time periods (such as hours, minutes, or seconds) to form a timeline. State representation: define a state space for each unmanned device and task, including position, speed, remaining energy, etc., as well as whether the task has been assigned, execution progress, etc.

[0119] (b) Define the dynamic programming state: This includes the definition of state variables and decision variables. Define a multidimensional state variable, including time, the state of the unmanned equipment, the state of the task, etc. In each time period, the decision variable indicates whether the unmanned equipment performs a task, which task it performs, and the duration of the task.

[0120] (c) Define the state transition equation. State transition of unmanned equipment: Calculate the state of the next time period based on the current state of the unmanned equipment, decision variables, and external conditions (such as wind speed, obstacles, etc.). Task state transition: Update the task status (such as completion degree, required resources, etc.) based on the execution progress of the task and decision variables.

[0121] (d) Define the objective function and constraints: Define the global optimization objective function corresponding to the time optimization algorithm, such as maximizing the number of completed tasks, minimizing resource consumption, balancing the workload of multiple unmanned devices, etc.

[0122] Constraints include: avoiding resource conflicts (e.g., two unmanned devices cannot occupy the same resource at the same time), avoiding task overlaps (some tasks can only be performed by one unmanned device alone; some tasks can be performed by multiple devices together, such as the material delivery task in the aforementioned example can be performed collaboratively by multiple unmanned devices), unmanned system capability limitations (e.g., energy, speed, etc.), subject allocation constraints of task viscosity, and execution time constraints caused by task execution timing dependencies.

[0123] In the process of solving dynamic programming, the state space is divided into a series of sub-state spaces, each of which corresponds to a time period and a set of states of unmanned equipment and tasks. Define a value function to evaluate the total benefit or total cost that can be obtained by taking a certain decision in each sub-state space. Using the Bellman optimality principle, a recursive relationship between value functions is established, that is, the value of the current state can be calculated by the value of the future state. Optimal strategy solution: Through reverse iteration (starting from the last time period, gradually calculating the optimal decision for each time period), the global optimal time planning allocation strategy is finally obtained. This time planning allocation strategy is an execution time allocation optimization scheme based on the task-subject allocation scheme obtained by the genetic algorithm based on reinforcement learning.

[0124] As an example, assuming that the evaluation indicator is the improvement of task completion efficiency, the corresponding objective function is: minimize the total completion time of all tasks, expressed as follows:

[0125]

[0126] Where f(x) represents the objective function, which is used to improve the efficiency of task completion; x is the vector corresponding to the decision variable; t i (x) represents the execution time of the i-th task under the strategy corresponding to the decision variable x.

[0127] Specific constraints include, but are not limited to: timing dependencies of task execution, task execution duration constraints caused by functional limitations, time window constraints, resource conflict avoidance constraints, task viscosity subject allocation constraints, etc. Some specific constraints are exemplified below. It is understandable that some constraints are not required for certain scenarios, and there may be constraints that are not exemplified for certain scenarios.

[0128] Constraint 1: The timing dependency of task execution, for example, the kth task must be started after the ith task is completed:

[0129] t i-start (x)+Δt ik ≤t k-start (x), (4)

[0130] Among them, t i-start (x) represents the start time of the i-th task under the strategy corresponding to the decision variable x; t k-start (x) represents the start time of the kth task under the strategy corresponding to the decision variable x; Δt ik Indicates the minimum time interval between the i-th task and the k-th task.

[0131] Constraint 2: Due to the functional characteristics of unmanned equipment (such as endurance), the execution time of the task cannot exceed a certain upper limit:

[0132] t i (x)≤threshold_duration, (5)

[0133] Among them, threshold_duration represents the execution time threshold of the i-th task.

[0134] Constraint 3: Time window restriction: the task can only be executed within a specific time period (for example, to avoid trajectory conflicts with other unmanned equipment, it needs to be executed within a specific time period; or the task requirement is that it can only be executed within a specific time period):

[0135] start_window_min≤t i-start (x)≤start_window_max, (6)

[0136] end_window_min≤t i-end (x)≤end_window_max, (7)

[0137] Among them, start_window_min represents the earliest start time of the start time window of the i-th task; start_window_max represents the latest start time of the start time window of the i-th task; end_window_min represents the earliest end time of the end time window of the i-th task; end_window_max represents the latest end time of the end time window of the i-th task; t i-end (x) represents the end time of the i-th task.

[0138] Constraint 4: Resource conflict avoidance constraint, to prevent two different unmanned devices from using the same resource at the same time:

[0139]

[0140] Among them, Resourceconsumptionj(t) means that the jth resource is consumed at time t; match indicates the matching relationship between resources and unmanned equipment; Equipmenti1 indicates the i1th unmanned equipment; Equipmenti2 indicates the i2th unmanned equipment; i1 and i2 are different equipment serial numbers.

[0141] Constraint 5: Task viscosity subject allocation constraint: when allocating related tasks, give priority to subjects that have been pre-assigned related tasks:

[0142] if task related-1 match Equipmenti1 not Equipmenti2, taski related-1 andtaski related-2 are related tasks,

[0143] then TaskViscosity(task related-2 ,Equipmenti1)>TaskViscosity(taski related-2 ,Equipmenti2),(9)

[0144] Among them, task related-1 Indicates the first task in the associated task, task related-2Indicates the second task in the associated task. There is an associated relationship between the first task and the second task. related-1 match Equipmenti1 notEquipmenti2 means that the first task has been assigned to the i1th unmanned equipment, but not to the i2th unmanned equipment;

[0145] TaskViscosity(task related-2 ,Equipmenti1) represents the task viscosity of the second task for the i1th unmanned equipment when it is assigned; TaskViscosity(task related-2 ,Equipmenti2) represents the task viscosity of the second task for the i2th unmanned equipment when it is assigned.

[0146] It is understandable that there are more constraints that are not illustrated. Since the time distribution optimization algorithm (the process of dynamically iteratively optimizing execution information using dynamic programming) is based on the task-subject allocation scheme obtained by the genetic algorithm based on reinforcement learning, it optimizes the distribution of execution time and execution process under a certain task-subject allocation scheme. Therefore, the task allocation influencing factors involved in the previous step can also be used as constraints of the time distribution optimization algorithm and will not be explained one by one.

[0147] It should also be noted that in the process of dynamic iterative optimization of task allocation execution information, possible conflicts (such as resource competition between tasks, time overlap, trajectory intersection conflicts, etc.) will also be monitored in real time, and corresponding avoidance measures will be taken according to the type and degree of conflict, such as reallocating tasks, adjusting task priorities, etc.; the results corresponding to the above-mentioned conflict avoidance strategies will be re-passed as feedback information to the genetic algorithm integrated with reinforcement learning, such as the feedback(x) item in the fitness function represented by formula (1-2) of the aforementioned embodiment example, so that the task allocation execution information continues to be dynamically iteratively optimized according to the conflict information.

[0148] In some embodiments, the motion trajectories of multiple unmanned devices can be analyzed and predicted based on the deep learning model by inputting the current state of the unmanned device (such as current position, speed, acceleration, heading, power, etc.) and environmental information (obstacle position, wind speed, weather conditions, etc.), and identifying potential trajectory intersection conflict points. The optimal conflict avoidance strategy can also be formulated based on the above deep learning model, so as to achieve safe and efficient collaborative operation between unmanned devices through a collaborative control mechanism.

[0149] In some embodiments, the relative positions and speeds between the unmanned devices are calculated based on the prediction results to identify potential conflict points. A conflict point can be defined as an area where two unmanned devices may meet or approach each other in a certain period of time in the future. Risk assessment is performed on the identified conflict points to determine the possibility and severity of the conflict, which can be achieved by calculating parameters such as the distance, speed difference, and heading difference of the conflict points.

[0150] Exemplary Devices

[0151] Figure 4 is a block diagram of a device for collaborative scheduling and execution of multi-agent tasks according to an exemplary embodiment. Figure 4 As shown, the present disclosure secondly exemplarily provides a device 400 for collaborative scheduling and execution of multi-agent tasks, and the device 400 includes: a task allocation module 410, an execution window planning module 420 and a dynamic optimization module 430.

[0152] The task allocation module 410 is used to perform preliminary allocation of collaborative tasks according to the capability characteristics and execution conditions of multiple unmanned devices to obtain a task-subject allocation plan.

[0153] The execution time window planning module 420 is used to plan the execution time windows for executing corresponding tasks of multiple unmanned equipment in the task-subject allocation scheme to obtain execution information.

[0154] The dynamic optimization module 430 is used to dynamically iteratively optimize the task allocation execution information according to the evaluation index during the period when the unmanned equipment performs the task based on the task-subject allocation scheme and the corresponding execution information, and continue to perform the task scheduling execution based on the optimized task allocation execution information. The dynamic iterative optimization of the task allocation execution information includes at least one of the following: dynamically iteratively optimizing the task-subject allocation scheme and dynamically iteratively optimizing the execution information.

[0155] For more details and beneficial effects of this embodiment, please refer to the relevant description of the first embodiment, which will not be repeated here.

[0156] Exemplary Electronic Devices

[0157] Figure 5 1 is a block diagram of an electronic device 500 according to an exemplary embodiment. The electronic device 500 may be a vehicle controller, a vehicle terminal, a vehicle computer or other types of electronic devices.

[0158] Reference Figure 5As shown, the electronic device 500 may include at least one processor 510 and a memory 520. The processor 510 may execute instructions stored in the memory 520. The processor 510 is communicatively connected to the memory 520 via a data bus. In addition to the memory 520, the processor 510 may also be communicatively connected to an input device 530, an output device 540, and a communication device 550 via a data bus.

[0159] The processor 510 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processor (Graphic Process Unit, GPU), a field programmable gate array (Field Programmable Gate Array, FPGA), a system on chip (System on Chip, SOC), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC) or a combination thereof.

[0160] The memory 520 may be implemented by any type of volatile or nonvolatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0161] In an embodiment of the present disclosure, executable instructions are stored in the memory 520, and the processor 510 can read the executable instructions from the memory 520 and execute the instructions to implement all or part of the steps of the method for collaborative scheduling and execution of multi-agent tasks described in any of the above exemplary embodiments.

[0162] Exemplary computer-readable storage media

[0163] In addition to the above methods and devices, the exemplary embodiments of the present disclosure may also be a computer program product or a computer-readable storage medium storing the computer program product. The computer product includes computer program instructions that can be executed by a processor to implement all or part of the steps described in any method in the above exemplary embodiments.

[0164] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and also conventional procedural programming languages, such as "C" language or similar programming languages ​​and scripting languages ​​(e.g., Python). The program code may be executed entirely on the user computing device, partially on the user computing device, as an independent software package, partially on the user computing device and partially on the remote computing device, or entirely on the remote computing device or server.

[0165] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of readable storage media include: a static random access memory (SRAM) with one or more wires electrically connected, an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk, or any suitable combination of the above.

[0166] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0167] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for collaborative scheduling and execution of multi-agent tasks, characterized in that: include: Perform preliminary allocation of collaborative tasks based on the capabilities and execution conditions of multiple unmanned devices to obtain a task-subject allocation plan; In the task-subject allocation scheme, execution time windows for multiple unmanned devices to execute corresponding tasks are planned to obtain execution information; During the period when the unmanned equipment performs the task based on the task-subject allocation scheme and the corresponding execution information, dynamically iteratively optimize the task allocation execution information according to the evaluation index, and continue to perform task scheduling execution based on the optimized task allocation execution information; The dynamically iterative optimization of the task allocation execution information includes at least one of the following: dynamically iterative optimization of the task-subject allocation scheme and dynamically iterative optimization of the execution information.

2. The method according to claim 1, characterized in that For at least one of the following situations, at least one of the task viscosity and conflict information is used as a factor affecting task allocation: Perform preliminary allocation of collaborative tasks based on the capabilities and execution conditions of multiple unmanned devices, or dynamically iterate and optimize the task allocation execution information based on evaluation indicators; The task viscosity indicates the difference in allocation priority between tasks and allocation subjects due to the association between tasks. The task viscosity value of the subject to which one or more associated tasks are pre-assigned and the subsequent associated tasks are allocated is greater than the task viscosity value of other subjects to which the subsequent associated tasks are allocated. The conflict information is used to indicate at least one of the following execution conflict situations: resource synchronization conflict when different unmanned devices apply for the same resource, spatial trajectory intersection conflict when different unmanned devices perform their respective tasks, and resource shortage conflict caused by the limited resources of the unmanned device when performing its own task.

3. The method according to claim 1 or 2, characterized in that: Based on a genetic algorithm integrated with reinforcement learning, preliminary allocation of collaborative tasks and dynamic iterative optimization of the task-agent allocation scheme are performed; Among them, the preliminary allocation of collaborative tasks is carried out based on the genetic algorithm integrated with reinforcement learning, including: Based on the policy network, an initial population is generated, wherein each individual in the initial population represents an unmanned equipment combination and an allocation scheme for collaborative task allocation under the combination; the input of the policy network includes: the current state of each unmanned equipment in the unmanned equipment set and information on the task to be assigned; the current state includes: current capability characteristics and current execution conditions; some or all of the unmanned equipment in the unmanned equipment set constitute the unmanned equipment combination; The allocation scheme corresponding to the initial population is determined as the preliminary allocated task-subject allocation scheme.

4. The method according to claim 3, characterized in that The evaluation indicators include: at least one of: improving the quality of task completion, improving efficiency, reducing costs, and avoiding execution conflicts; According to the evaluation indicators, based on the genetic algorithm integrated with reinforcement learning, the task-agent allocation scheme is dynamically iteratively optimized, including: Based on reinforcement learning, the policy network is updated with at least one of the improvement of task completion quality, efficiency improvement, cost reduction, and avoidance of execution conflicts as a reward function, so that the policy network learns a better task allocation plan, and the initial population is evolved based on the updated policy network; the evolutionary process includes preset rounds of selection-crossover-mutation operations: new individuals in the population are generated based on the updated policy network, and the new individuals correspond to a better task allocation plan; for the individuals in the initial population and the new individuals, a target individual is selected as a parent individual based on a preset fitness function; a crossover operation is performed on the parent individual to generate a new child individual; and a mutation operation is performed on the new child individual based on a preset probability; For the individuals in the population after evolution processing, the individual with the highest function value corresponding to the fitness function is determined as the optimized target task-subject allocation scheme.

5. The method according to claim 4, characterized in that The fitness function satisfies the following expression: Where Fitness(x) represents the fitness function of individual x; Taskcompletioni(x) represents the completion of the i-th task in the task allocation scheme corresponding to individual x; i represents the task number, which ranges from 1 to N, and N represents the total number of tasks; w i represents the weight corresponding to the time distribution matching degree of the i-th task; Resourceconsumptionj(x) represents the consumption of the j-th resource in the task allocation scheme corresponding to individual x; j represents the resource type number, ranging from 1 to M, and M represents the total number of resource types; c j represents the weight corresponding to the resource conflict avoidance rate of the j-th resource, and the resource conflict avoidance rate is determined based on the following execution conflict situations: resource synchronization conflict when different unmanned devices apply for the same resource and resource shortage conflict caused by the limited resources of the unmanned device when performing its own task; Matchingscore(x) represents the matching score between the capability characteristics of the unmanned device and the task requirements; feedback(x) represents the feedback input information of the allocation plan corresponding to the individual x during the execution process, and the feedback input information includes at least one of task viscosity and conflict information; the conflict information includes: spatial trajectory intersection conflict when different unmanned devices perform their respective tasks.

6. The method according to claim 4, characterized in that The reward function satisfies the following expression: Reward(s,a)=Fitness(a)+γ a′ ·maxFitness(s′,a′), Among them, Reward(s,a) represents the reward function corresponding to the current environment state s and the current action a; the current environment state includes task requirements and resource conditions, and the current action includes at least one of the following: adjusting the correspondence between tasks and subjects, and adjusting the resource allocation for task execution; a′ represents the next action; s′ represents the next environment state; γ a′ Represents the discount factor corresponding to the next action a′, which is 0≤γ a′ ≤1; maxFitness(s′,a′) represents the maximum value of the fitness function corresponding to the next action a′ in the next environment state s′.

7. The method according to claim 1 or 6, characterized in that: The evaluation indicators include: at least one of: improving the quality of task completion, improving efficiency, reducing costs, and avoiding execution conflicts; The execution information is dynamically iterated and optimized based on the dynamic programming method and the evaluation index, including: Divide the total time window for multiple unmanned devices to perform corresponding assigned tasks into multiple time periods; During the period when multiple unmanned devices perform tasks according to the execution information, for each current time period, dynamic planning is performed according to the execution information, the external environment and the real-time status of the device to determine the execution information of the next time period; until the task execution of the total time window is completed in a loop iteration; For each current time period, dynamic planning is performed based on the execution information, external environment, and real-time status of the device to determine the execution information for the next time period, including: For the current time period, obtain current status information; the current status information includes: real-time status of the device, task execution status and external environment; the task execution status includes the execution information, and the execution information includes one or more of the following: whether the task is executed, the task execution period, the resources required for task execution, and the task execution progress; Based on the decision variables, the current state information in the current time period is processed to generate multiple predicted state information for the next time period, and the subsequent time period decisions are recursively made based on the multiple predicted state information to generate the final predicted state information; According to the evaluation index, determine the optimal decision strategy that satisfies the evaluation index or maximizes the benefit of the evaluation index from the decision strategies composed of the predicted state information of the future time period after the current time period; The task execution status in the optimal prediction status information of the next time period in the optimal decision strategy is determined as the execution information of the next time period.

8. A device for collaborative scheduling and execution of multi-agent tasks, characterized in that: include: The task allocation module is used to perform preliminary allocation of collaborative tasks according to the capabilities and execution conditions of multiple unmanned devices to obtain a task-subject allocation plan; An execution time window planning module is used to plan the execution time windows for executing corresponding tasks of multiple unmanned equipment in the task-subject allocation scheme to obtain execution information; A dynamic optimization module is used to dynamically iteratively optimize the task allocation execution information according to the evaluation index during the period when the unmanned equipment performs the task based on the task-subject allocation scheme and the corresponding execution information, and continue to schedule the task based on the optimized task allocation execution information; The dynamically iterative optimization of the task allocation execution information includes at least one of the following: dynamically iterative optimization of the task-subject allocation scheme and dynamically iterative optimization of the execution information.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1-7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.