Unmanned aerial vehicle group cooperative task planning method based on adaptive search teaching and learning algorithm

Through the adaptive search teaching and learning algorithm (TLBO) combined with neighborhood search operator, high-quality initial solution generation strategy and multimodal neighborhood search operator are designed, which solves the problem of inefficiency of traditional drone task planning algorithms in complex battlefield environments, and achieves efficient and accurate task planning and resource allocation.

CN120276494AActive Publication Date: 2025-07-08DALIAN UNIV OF TECH
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510748207.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional UAV mission planning algorithms are difficult to achieve efficient task allocation and path planning under the satisfaction of multiple constraints. Especially in complex battlefield environments, existing algorithms have high computational complexity, slow convergence speed, and are prone to fall into local optimality, and cannot effectively handle multiple tasks and timing relationships.

Method used

The drone cluster collaborative task planning method based on adaptive search teaching and learning algorithm (TLBO) is adopted, and combined with neighborhood search operators, a high-quality initial solution generation strategy and multimodal neighborhood search operators are designed to optimize the UAV resource allocation and task execution order.

Benefits of technology

It significantly improves the efficiency and accuracy of drone cluster mission planning, can quickly converge to high-quality solutions, meet multiple constraints, realize dynamic task allocation and global optimization, and is suitable for complex battlefield environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276494A_ABST
    Figure CN120276494A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle group cooperative task planning method based on an adaptive search teaching and learning algorithm, and belongs to the field of unmanned aerial vehicle task planning. The method comprises the steps that firstly, key factors of actual constraints are extracted, decision variables and constraints are defined, and a combinatorial optimization model is constructed; constructing a coding form, and transforming a target function into a fitness function suitable for coding; secondly, designing a high-quality initial solution generation scheme based on a code setting mode, providing a high-quality initial solution for a teaching and learning algorithm, finally summarizing each code to obtain an initial population, and determining a code updating rule to obtain an evolution population; finally, a neighborhood search operator is introduced in the evolution process, neighborhood search operation is carried out in a self-adaptive mode, and a solving scheme is obtained. According to the method, the real combat process in the actual combat is considered, the risk that the carrier is found by an enemy can be reduced, and the safety of the carrier and smooth execution of subsequent tasks are guaranteed; efficient distribution of unmanned aerial vehicle resources and optimization of task planning can be realized in a complex battlefield environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of UAV mission planning, and relates to a method for collaborative mission planning of UAV swarms based on an adaptive search teaching and learning algorithm. Background Art

[0002] With the rapid development of UAV technology, UAV swarms are increasingly widely used in military, civilian and other fields. Especially in modern warfare, the collaborative operation of UAV swarms has become an important means to improve combat effectiveness. However, with the expansion of the battlefield scale and the increase in task complexity, the task planning problem of UAV swarms has become increasingly complex. Traditional task planning methods often have difficulty in achieving efficient task allocation and path planning while satisfying multiple constraints (such as UAV range, ammunition loading, task sequence, etc.). The task planning problem of UAV swarms is essentially a complex multi-objective optimization problem, which usually involves the reasonable allocation of UAV resources, the optimization of task execution order, and the minimization of task completion time. Due to the NP-hard nature of this problem, traditional optimization algorithms (such as genetic algorithms, particle swarm algorithms, etc.) often face problems such as high computational complexity, slow convergence speed, and easy to fall into local optima when dealing with large-scale tasks. Therefore, how to design an efficient and stable task planning algorithm has become a key problem in the current research on UAV swarm collaborative combat.

[0003] Most existing UAV mission planning algorithms are based on a single optimization objective (such as mission completion time or the number of UAVs used), and rarely consider the temporal relationship between tasks and the dynamic allocation of UAV resources. For example, a Chinese invention patent provides a mission planning method for cooperative ground combat between manned and unmanned aerial vehicles (CN112733421A). This method does not consider the situation where there are multiple task types for the same target during modeling, and cannot perform calculations when facing the need for dual tasks of strike and assessment for a target in actual combat. Another Chinese invention patent provides a mission planning method for cooperative reconnaissance of multi-base heterogeneous UAVs (CN107330588A), which provides a mission planning method for cooperative reconnaissance of multi-base heterogeneous UAVs. This method considers the mission planning problem under multiple bases and uses the cuckoo algorithm for solution. Similarly, it only considers the situation of a single task and cannot respond to the actual combat requirements. In addition, traditional heuristic algorithms have deficiencies in generating initial solutions and local search capabilities, resulting in poor solution efficiency and stability of the algorithms in complex task scenarios. For example, a Chinese invention patent provides a multi-UAV task allocation method based on a multi-objective quantum krill swarm mechanism (CN112926825A). The encoding method of this method is not applicable to multi-task problems. If the corresponding encoding form is modified, a large amount of search cost will be generated, which cannot meet the timeliness required for actual combat and provide high-quality solutions. Therefore, there is an urgent need for a mission planning algorithm that can comprehensively consider various constraint conditions, has the ability to generate efficient initial solutions and strong local search capabilities, to meet the complex UAV swarm cooperative combat requirements in modern warfare.

[0004] To address the above problems, the present invention proposes a UAV mission planning algorithm based on the Teaching-Learning-Based Optimization (TLBO) algorithm. By introducing a heuristic initial solution generation strategy and a multi-modal neighborhood search operator, the present invention significantly improves the global search ability and local optimization ability of the algorithm, and can achieve efficient allocation of UAV resources and optimization of task execution order on the premise of meeting various constraint conditions. Summary of the Invention

[0005] To solve the above-mentioned technical problems, the present invention proposes a method for cooperative mission planning of UAV swarms based on an adaptive search teaching and learning algorithm, which is mainly applied to the combat mission planning of UAV swarms. In the problem background of the present invention, for each target point, the UAV will first perform an attack mission on it, and then the UAVs that do not attack this target point will perform an assessment mission. After all tasks are completed, all UAVs return to the delivery point to wait for recovery. The present invention will construct a combinatorial optimization model around the actual situation and use the TLBO algorithm combined with a neighborhood search operator during solution to achieve fast search of the results.

[0006] To achieve the above object, the specific technical solution adopted by the present invention is as follows:

[0007] A method for collaborative mission planning of unmanned aerial vehicle (UAV) swarms based on an adaptive search teaching and learning algorithm, the method for collaborative mission planning of UAV swarms comprising the following steps:

[0008] Step 1, extract the key factors of the actual constraints, clarify the decision variables, and construct a combinatorial optimization model for the constraints;

[0009] Step 2, construct an encoding form around the decision variables and constraint settings in the model. Then, according to the encoding form, transform the objective function into a fitness function applicable to the encoding for subsequent evaluation of the quality of the solution corresponding to the encoding;

[0010] Step 3, based on the encoding setting method, design a high-quality initial solution generation scheme to provide a high-quality initial solution for the teaching and learning algorithm. Generate multiple encodings through this strategy, and finally summarize each encoding to obtain an initial population;

[0011] Step 4, based on the initial population, determine the encoding update rule, that is, the evolution method of the encoding in the teaching and learning stages of the teaching and learning algorithm, to achieve the self-evolution of the algorithm and obtain an evolved population;

[0012] Step 5, introduce a neighborhood search operator during the evolution process, adaptively perform neighborhood search operations, further improve the search efficiency of the algorithm, enhance the search ability of the population, and finally obtain a solution.

[0013] Further, the specific content of Step 1 is as follows:

[0014] First, extract the key factors and problem-solving requirements in the actual constraints, clarify the decision variables, and construct a combinatorial optimization model for the constraints. Let the task objective set be , represent the number of assignable UAVs, represent the number of target points, represent the UAV set, represent the target point set. At this time, all the positions passed by the UAVs are . Among them, represent the th UAV; represent the th target point. Then, carry out the model representation.

[0015] Step 1.1: Construct decision variables and related representations according to the problem-solving requirements;

[0016] In combinatorial optimization problems, decision variables usually represent the task plans to be obtained, which need to include all the information of the solutions to be found. In this problem, the information to be solved includes the numbers of the drones used and the tasks executed by each drone. Therefore, the decision variables are set as follows:

[0017] (1)

[0018] In formula (1), represents the decision variable of this problem, represents the task type, represents the drone starting from the target point and going to the target point to execute the task. This formula indicates that each task at each target point needs to be completed. Correspondingly, the following formula can be obtained from the decision variable:

[0019] (2)

[0020] Formula (2) is used to represent the flight route of the drone. When it means the drone starting from the target point and going to the target point Obviously, when there must be In addition, there is also:

[0021] (3)

[0022] Formula (3) is used to represent the state of the drone. When , the drone needs to be dispatched to execute the task.

[0023] Step 1.2: Construct the task completion guarantee constraint according to the task requirements;

[0024] (4)

[0025] Formula (4) indicates that each task at each target point needs to be completed, where represents that there are two types of task types, namely attack tasks and assessment tasks, and there are two types of tasks that need to be executed by drones at each target point.

[0026] Step 1.3: Construct the task execution requirement constraint according to the task execution requirements;

[0027] (5)

[0028] Formula (5) indicates that for the drone , at most only one task for the same target point can be executed can be carried out.

[0029] Step 1.4: Construct flight constraints according to the flight conditions of the UAV;

[0030] (6)

[0031] Formula (6) indicates that for any UAV , starting from the delivery point and finally returning to the delivery point, where and the subscript 0 in represents the delivery point. Similarly, there is the following formula:

[0032] (7)

[0033] Formula (7) indicates that for any UAV , after traveling from any target point to target point , the subsequent journey must start from target point , that is, the out-degree of each target point is equal to the in-degree.

[0034] Step 1.5: Construct task completion constraints according to the task execution requirements;

[0035] (8)

[0036] Formula (8) indicates that each target point needs to complete 2 types of tasks and must be executed.

[0037] Step 1.6: Construct task execution time constraints according to the actual flight rules and execution capabilities of the UAV;

[0038] (9)

[0039] Formula (9) indicates that there is a sequence of tasks for each target point, and there is a minimum interval time, where is the task execution time, is the minimum interval time between two tasks, and respectively represent the start time of the attack task and the start time of the evaluation task at target point .

[0040] Step 1.7: Construct UAV capacity constraints (maximum flight time constraint, ammunition loading constraint) according to the UAV's own capacity limitations;

[0041] (10)

[0042] Formula (10) represents the flight time constraint of the UAV, that is, the flight time of any UAV should be less than its maximum flight time, where Indicates the flight time of the UAV , Indicates the UAV 's maximum flight time. In addition, since the UAV consumes a certain amount of ammunition when performing attack missions, it is also necessary to set the ammunition constraint of the UAV. In this problem, it is assumed that the UAV consumes one ammunition each time it performs an attack mission. The formula is as follows:

[0043] (11)

[0044] Formula (11) represents the UAV ammunition loading constraint, that is, the number of times the UAV performs an attack mission needs to be less than its ammunition loading upper limit, where represents the maximum ammunition loading quantity of the UAV.

[0045] Step 1.7: Comprehensively consider the mission requirements and complete the expected construction to combine and optimize the objective function;

[0046] (12)

[0047] The objective function of this problem is shown in Equation (12), where represents the estimated maximum number of UAVs used, represents the UAV 's maximum flight time, and are used to adjust the weights of each item and satisfy . The optimization index of this objective function is to minimize the number of UAVs used and the mission completion time. And is obtained from the following formula, and the specific form is as follows:

[0048] (13)

[0049] Among them, represents the straight-line distance from the delivery point to the target point, represents the number of target points, represents the UAV range, represents the target point, represents the set of target points.

[0050] Furthermore, the specific steps of Step 2 are as follows:

[0051] Step 2.1: Construct an algorithm encoding according to the decision variables;

[0052] In intelligent algorithms, encoding is a digital string used to store decision variables and can be recognized by a computer. Therefore, in the present invention, it is also necessary to construct the corresponding Teaching-Learning Based Optimization (TLBO) algorithm encoding around the decision variables and constraints in the model. For the convenience of description, the Teaching-Learning Based Optimization algorithm is hereinafter collectively referred to as the TLBO algorithm.

[0053] The encoded information includes the targets and Unmanned Aerial Vehicles (hereinafter collectively referred to as UAVs), and the task sequences of the same UAV are executed in the encoding order from front to back. The encoding contains two lines of information. The first line stores the target points, and the second line stores the UAV numbers. The length of each line is twice the number of target points, which can be regarded as a 2-row column matrix. In the first line of code, when a target appears for the first time, it is represented as an attack task for that point. Therefore, the second appearance of the target represents an evaluation task. And in the second line, it represents the UAV that executes the task, and the task it executes is the task indicated in the first line encoding in the same column. In summary, the decision variables of the combinatorial optimization model proposed in the present invention are stored through such encoding. And the decision variables corresponding to the encoding determine the complete execution plan of a round of tasks. The task execution plan obtained from the encoding is generally called the decoded plan, and this term will also be used subsequently.

[0054] Step 2.2: According to the constraint settings and the encoding form, transform the objective function into a fitness function;

[0055] In algorithms similar to TLBO, it is necessary to transform the objective function in the combinatorial optimization model into a corresponding fitness function to describe the quality of the solution corresponding to a certain encoding. Since the problem studied in the present invention has relatively high requirements for the solution time, the present invention also needs to transform the objective function of the combinatorial optimization model proposed in Step 1 to obtain a fitness function that can quickly evaluate the quality of the solution corresponding to the encoding. On the other hand, since in the TLBO solution process, the generated encoding cannot fully guarantee meeting the constraints, it is necessary to introduce a penalty term to indicate whether the decoded solution of the current encoding can be implemented. The specific form is as follows:

[0056] (14)

[0057] *(15)

[0058] (16)

[0059] Among them, represents the fitness function; and are the weights for adjusting the two main indicators, and satisfy ; represents the penalty term for the part of the decoded solution that exceeds the flight time constraint; The penalty term for the part where the decoding scheme exceeds the ammunition loading constraint; Represents the estimated maximum number of drones to be used; Represents the maximum flight time of the drone; Represents the drone in the decoding scheme Flight time; Represents the penalty coefficient of the penalty term for the flight time constraint part; Represents the penalty coefficient of the penalty term for the ammunition loading constraint; Represents the drone Number of times to perform attack missions; Represents the maximum ammunition loading capacity of the drone.

[0060] In formula (14), the fitness function consists of three parts. The first part (the first two terms in the formula) represents the objective function given in step 1, that is, minimizing the number of drones used and the task completion time. The second part Represents that in the decoding scheme of the encoding, penalties are imposed on all drones that exceed the flight time limit. The third part Represents that in the decoding scheme of the encoding, penalties are imposed on all drones that exceed the maximum ammunition loading capacity.

[0061] Furthermore, the specific steps of step 3 are as follows:

[0062] Considering the temporal relationship among tasks in the same target point, which makes the corresponding high-quality feasible solution space narrow, in order to improve the search efficiency of the algorithm, the present invention designs a generation strategy for high-quality initial solutions, and the specific process is as follows:

[0063] First, according to the target set to be attacked, copy and generate an attack mission target set and an evaluation mission target set, where the two generated sets are exactly the same as the targets in the target set. And take the ceiling of the number obtained by dividing the number of targets by the ammunition load as the initial number of drones, and store it in a specific set, named the motivation set here. Without loss of generality, in the following process, the attack mission target set is taken as an example, and the allocation process of the evaluation mission target set is the same as the following process.

[0064] Step 3.1, randomly select a target from the attack mission target set;

[0065] Step 3.2, calculate the time required for each drone in the motivation set to perform the mission of the selected target in step 3.1, including the navigation time from the previous target to the selected target and the mission execution time;

[0066] Step 3.3, on the premise of satisfying Formula (11) and Formula (17), select the UAV with the minimum time required to execute the task. If no executable UAV can be found, add a UAV to the UAV set and assign the task to the newly added UAV. Among them, Formula (17) is as follows:

[0067] (17)

[0068] Among them, represents the current flight time of the UAV ; is a manually set parameter used to balance the time for a UAV to execute offensive tasks and evaluation tasks; represents the maximum flight time of the UAV; the above formula means that the time for the UAV to execute the attack task shall not exceed a certain ratio of its own maximum flight time. Different will have different impacts on the initial solution.

[0069] Step 3.4, delete the target randomly selected in Step 3.1 from the offensive task target set, and convert the corresponding task execution information into encoded information and store it on the encoding, that is, a column on the encoding introduced in Step 2.1;

[0070] Step 3.5, repeat Steps 3.1 to 3.4 until the offensive task target set becomes an empty set;

[0071] Step 3.6, perform a process similar to Steps 3.1 to 3.4 on the evaluation task target set. The difference is that the iterative process for the offensive task target set in Steps 3.1 to 3.4 is replaced by the evaluation task target set. In addition, when the UAV with the minimum time required to execute the evaluation task is also the UAV for executing the attack task of this target, then select the UAV with the second smallest required time or add a UAV again. Finally, when performing the evaluation task assignment, in Formula (17) takes the value of 1.

[0072] After determining the initial solution generation strategy, use the above strategy to generate several initial solutions as the search starting point of TLBO, where the number of generated initial solutions is determined by the preset parameter . We call several generated initial solutions combined into a population, and is the population size. Each solution in the population is called an individual, and the above names will also be used subsequently. Finally, for each generated individual, use the fitness function proposed in Step 2 to calculate the fitness of each individual and save it. From the setting of the fitness function, the lower the fitness of an individual, the more excellent the corresponding execution plan.

[0073] Furthermore, the specific steps of Step 4 are as follows:

[0074] The evolution of individuals in the population is essentially the exchange of encoded information between individuals and the independent exploration of new encodings by individuals. In step 4, the method of information exchange in the encoding between individuals is mainly given, that is, the education stage and the learning stage in the TLBO algorithm. Before the start of the solution by TLBO, an iteration number will be preset, that is, after generating the initial population, the number of times the entire population evolves. In one iteration, each individual in the population will execute the education stage, the learning stage mentioned in step 4, and the neighborhood search proposed in the subsequent step 5. After all individuals complete the above three processes, the fitness function value of each individual is solved using the fitness function proposed in step 2, and the individual with the lowest function value is selected as the optimal individual of the population, that is, the teacher.

[0075] Step 4.1: Education stage;

[0076] Since the evolutionary method for encoding in the traditional TLBO algorithm is not applicable to the integer encoding proposed in step 2, for this reason, the present invention proposes an evolutionary method between individuals applicable to this problem, and the specific update method is as follows:

[0077] Based on the encoding form set in step 2, assume that there are two individuals, individual and individual . First, randomly select target points from each target point, and subsequently, the selected target points will be used to guide the information exchange between the two individuals. Without loss of generality, if individual learns from individual , locate the positions of the previously selected targets in the encoding of individual , and use this to guide individual to perform encoding position replacement. The specific operation is as follows: First, create an empty encoding with the same size as individual . Traverse the selected target points one by one. Let the currently traversed target point be . At this time, access the encoding of individual , and find the position where appears in the encoding of (it appears twice in total. Once represents the attack task, assuming its column in the encoding at this time is . Once represents the evaluation task, assuming its column in the encoding at this time is ). Copy all the information (including the target point and the drone number corresponding to the task execution) in the column of the encoding of to the th column of In the column. Similarly, encoded Copy all the information of the column to the column. At this point, the information exchange of the target point is completed. Repeat the above process until the selected target points have all the information on encoded and copied to

[0078] At this time, the remaining target points' task information has not been stored in . At this time, access the encoding, and traverse its encoding backward from the first column. If the target point recorded in the first row of the current column of the encoding belongs to one of the remaining target points, then copy the information of this column to the blank column with the earliest position in . When there are no more blank columns, the information exchange process ends. is the individual obtained after learning is completed.

[0079] In each round of iteration, the teacher will conduct an education process for each individual in the population. Assume is the teacher in the population, and a certain individual in the population is , then the process of the student individual learning from the teacher is shown in formula (18):

[0080] (18)

[0081] Among them, represents the individual encoding obtained after learning. In formula (18), represents the individual described above learning from the individual process.

[0082] Step 4.2: Mutual learning stage;

[0083] In the learning stage, individuals recognize their gaps through pairwise interaction and communication, and improve their knowledge level accordingly. This stage is carried out after all individuals complete the education stage, and each individual has to conduct a mutual learning once. The mechanism is as follows:

[0084] Let the currently selected student individual be , randomly select an individual in the population for communication and learning. represents the fitness function of the individual. The student individual The communication process is as shown in formula (19):

[0085] (19)

[0086] When the individual fitness is greater than then directly retain the result after learning;

[0087] When the individual fitness is less than first obtain and then calculate the fitness of If its fitness does not develop towards a lower value, then regard this learning as ineffective learning and degrade the individual to before learning, that is ; if the fitness successfully develops towards a lower value, then retain the effect of this learning.

[0088] When each individual in the population completes the education stage and the mutual learning stage, then carry out step 5.

[0089] Furthermore, the specific steps of step 5 are as follows:

[0090] Since step 4 focuses on the information exchange within the population, it is also necessary to supplement the "exploration" ability of TLBO in the solution space. After the process of education between teachers and students and the mutual learning among students, the present invention endows individuals with autonomous thinking and learning behaviors on this basis, that is, individuals can expand to a certain extent based on the current knowledge, further improving the efficiency of finding the optimal solution.

[0091] It includes a total of three search behaviors, specifically: (1) randomly select a column in the encoding and randomly select a drone from the replaceable drone pool to replace the drone number at this gene position; (2) select a certain column of the encoding and randomly select a certain gene position within the corresponding feasible region for exchange; (3) select a certain column of the encoding and randomly select any valuable insertion position within the corresponding feasible region for insertion. In each round of iteration, each individual will carry out neighborhood search, and all three above-mentioned search behaviors are executed with a certain probability. After all are completed, output the search result and retain the corresponding result.

[0092] The beneficial effects of the present invention are:

[0093] (1) High - efficiency task planning ability: The present invention combines the adaptive search Teaching - Learning - Based Optimization (TLBO) algorithm with a neighborhood search operator, significantly improving the efficiency and accuracy of the task planning for unmanned aerial vehicle (UAV) swarms. Its innovation lies in the design of a high - quality initial solution generation strategy and a multi - modal neighborhood search operator, enabling the algorithm to quickly converge to high - quality solutions. In the specific implementation, the initial solution generation strategy in step 3 avoids the inefficient search of random initial solutions, and the neighborhood search operator in step 5 enhances the local search ability of the algorithm, thus realizing the optimization of task allocation under complex constraint conditions.

[0094] (2) Global optimization under multiple constraints: The present invention comprehensively considers various constraint conditions such as UAV range, ammunition loading, and task timing. Through the combinatorial optimization model constructed in step 1 and the fitness function designed in step 2, the multi - objective optimization problem is transformed into a solvable form. The penalty terms (such as flight - time constraint and ammunition constraint) in the fitness function ensure that the generated solutions meet the actual requirements, thus achieving global optimization.

[0095] (3) Dynamic task allocation ability: The present invention can handle the requirements of multiple tasks (attack and assessment tasks) at the same target point, and through the constraint design in steps 1.2 and 1.3, ensures the timing and uniqueness of task execution. The education phase and the mutual - learning phase in step 4 dynamically adjust the task allocation scheme through information exchange between individuals, further enhancing the flexibility and adaptability of the algorithm.

[0096] (4) Fast convergence and stability: Through the design of the TLBO algorithm in step 4, combined with the teacher - guidance and student - mutual - learning mechanisms, the algorithm efficiently explores the solution space, avoiding the problem that traditional heuristic algorithms are prone to falling into local optima. The neighborhood search operator in step 5 further enhances the search ability of the algorithm, enabling it to quickly converge to a stable solution within a limited number of iterations. The convergence speed and convergence of the present invention are superior to traditional algorithms.

[0097] (5) Wide applicability: The technical solution of the present invention is not only applicable to military combat scenarios, but can also be extended to the collaborative task planning of UAVs in the civilian field, such as disaster relief, logistics distribution, etc. Its core innovations (such as coding design, constraint handling, and adaptive search mechanism) are universal and can adapt to task requirements of different scales and complexities.

[0098] In summary, the present invention takes into account the actual combat process in real combat. By optimizing the delivery quantity of unmanned aerial vehicles (UAVs), shortening the delivery time of the vehicle, reducing the risk of the vehicle being detected by the enemy, and ensuring the safety of the vehicle and the smooth execution of subsequent tasks. At the same time, through the high-quality initial solution generation strategy and neighborhood search operator, the search efficiency and solution quality of the algorithm are significantly improved. It can achieve the efficient allocation of UAV resources and the optimization of mission planning in a complex battlefield environment, ensuring the efficient execution and safety of the mission, and has important practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 is a flowchart of the invention.

[0100] Figure 2 is a coding schematic diagram.

[0101] Figure 3 is a decoding schematic diagram.

[0102] Figure 4 is an individual learning schematic diagram.

[0103] Figure 5 is a generated graph of the neighborhood search optional UAV pool.

[0104] Figure 6 is a schematic diagram of neighborhood search exchangeable and insertable.

[0105] Figure 7 is a schematic diagram of valuable insertion positions for the first type of neighborhood search.

[0106] Figure 8 is a schematic diagram of valuable insertion positions for the second type of neighborhood search.

[0107] Figure 9 is a flowchart of algorithm execution.

[0108] Figure 10 is a Gantt chart of Example 1.

[0109] Figure 11 is a Gantt chart of Example 2.

[0110] Figure 12 is a Gantt chart of Example 3.

[0111] Figure 13 is a convergence curve graph of Example 1.

[0112] Figure 14 is a convergence curve graph of Example 2.

[0113] Figure 15 is a convergence curve graph of Example 3. DETAILED DESCRIPTION OF THE INVENTION

[0114] The present invention will be further described below in conjunction with specific embodiments. To ensure the rationality of the discussion, we compared the different performances under six groups of test cases. They respectively contain 16, 13, 8, 13, 10, and 15 target points. The remaining UAVs and algorithm-related parameters are shown in the table:

[0115] Table 1 shows the UAV-related parameters

[0116]

[0117] Table 2 shows the related parameters of the underlying UAV mission planning

[0118]

[0119] The specific steps of Step 1 are as follows:

[0120] First, extract the key factors and problem-solving requirements in the actual constraints, clarify the decision variables, and construct a combinatorial optimization model. Let the task objective set be , represents the number of assignable UAVs, represents the number of target points, represents the set of UAVs, represents the set of target points. At this time, all the positions passed by the UAVs are . Among them, represents the th UAV; represents the th target point. After that, the model representation is carried out.

[0121] Step 1.1: Construct decision variables and related representations according to the solution requirements;

[0122] In combinatorial optimization problems, decision variables usually represent the task plans to be obtained, which need to include all the information of the plans to be solved. In this problem, the information to be solved includes the UAV numbers used and the tasks executed by each UAV. Therefore, the decision variables are set as follows:

[0123] (1)

[0124] In formula (1), represents the decision variable of this problem, represents the task type, represents the UAV starting from the target point and going to the target point to execute tasks. This formula means that each task at each target point needs to be completed. Correspondingly, the following formula can be obtained from the decision variables:

[0125] (2)

[0126] Formula (2) is used to represent the flight path of the UAV. When it represents that the UAV starts from the target point and heads to the target point . Obviously, when it is certain that . Additionally, there is also:

[0127] (3)

[0128] Formula (3) is used to represent the state of the UAV. When the UAV needs to be dispatched to perform a mission.

[0129] Step 1.2: Construct the task completion guarantee constraints according to the mission requirements;

[0130] (4)

[0131] Formula (4) indicates that each mission at each target point needs to be completed, where indicates that there are two types of mission types, namely attack missions and assessment missions, and there are two types of missions that the UAV needs to execute at each target point.

[0132] Step 1.3: Construct the task execution requirement constraints according to the mission execution requirements;

[0133] (5)

[0134] Formula (5) indicates that for the UAV , at most only one mission at the same target point can be executed.

[0135] Step 1.4: Construct the flight constraints according to the UAV flight conditions;

[0136] (6)

[0137] Formula (6) indicates that for any UAV , it starts from the delivery point and finally returns to the delivery point, where and the subscript 0 in represents the delivery point. Similarly, there is also the following formula:

[0138] (7)

[0139] Formula (7) indicates that for any UAV , from any target point to the target point After that, the subsequent itinerary must start from the target point and leave, that is, the out-degree of each target point is equal to the in-degree.

[0140] Step 1.5: Construct task completion constraints according to the task execution requirements;

[0141] (8)

[0142] Formula (8) means that two types of tasks need to be executed for each target point to be completed.

[0143] Step 1.6: Construct task execution time constraints according to the actual flight laws and execution capabilities of the UAVs;

[0144] (9)

[0145] Formula (9) means that there is a sequence of tasks for each target point, and there is a minimum interval time, where is the task execution time, is the minimum interval time between two tasks, and respectively represent the start time of the attack task and the start time of the assessment task for the target point .

[0146] Step 1.7: Construct UAV capability constraints (maximum flight time constraint, ammunition loading constraint) according to the UAV's own capability limitations;

[0147] (10)

[0148] Formula (10) represents the flight time constraint of the UAV, that is, the flight time of any UAV should be less than its maximum flight time, where represents the flight time of the UAV , represents the UAV . In addition, since the UAV consumes a certain amount of ammunition when performing an attack task, it is also necessary to set the ammunition constraint of the UAV. In this problem, it is assumed that the UAV consumes one ammunition each time it performs an attack task. The formula is as follows:

[0149] (11)

[0150] Formula (11) represents the ammunition loading constraint of the UAV, that is, the number of times the UAV performs an attack task should be less than its ammunition loading upper limit, where represents the maximum ammunition loading quantity of the UAV.

[0151] Step 1.7: Comprehensively consider the task requirements and complete the expected construction of the combined optimization objective function;

[0152] (12)

[0153] The objective function of this problem is shown in formula (12), where Indicates the estimated maximum number of drones in use, Indicates drone The maximum flight time, and To adjust the weight of each item, and satisfy The optimization index of this objective function is to minimize the number of drones used and the time to complete the task. It is obtained by the following formula, the specific form is as follows:

[0154] (13)

[0155] in, Represents the straight-line distance from the delivery point to the target point. represents the number of target points, Indicates the range of the drone. represents the target point, Represents a set of target points.

[0156] Step 2: Construct the algorithm code around the decision variables and constraint settings in the model. Then, according to the code form, transform the objective function into a fitness function suitable for the code, which is used to evaluate the quality of the solution corresponding to the code;

[0157] Step 2.1: Construct algorithm coding based on decision variables;

[0158] In intelligent algorithms, codes are strings of numbers that are used to store decision variables and can be recognized by computers. Therefore, in the present invention, it is also necessary to construct corresponding teaching and learning algorithm codes around the decision variables and constraint settings in the model. For the convenience of description, the teaching and learning algorithm is collectively referred to as the TLBO algorithm below.

[0159] The encoded information includes the target and the UAV (hereinafter referred to as UAV), and the mission sequence of the same UAV is executed in the order of encoding from front to back. The encoding contains two lines of information, the first line stores the target point, and the second line stores the UAV number. The length of each line is twice the number of target points, which can be regarded as a 2-line A matrix of columns. In the first line of code, when the target first appears, it is represented as an attack task at that point. Therefore, the second appearance of the target represents an evaluation task. And in the second line, it represents the UAV that executes the task, and the task it executes is the task indicated in the first line of coding in the same column. In summary, the decision variables of the combined optimization model proposed by the present invention are stored through such coding. And the decision variables corresponding to the coding determine the complete execution plan of a round of tasks. The task execution plan obtained from the coding is generally called the decoded plan, and this term will also be used subsequently. For a more intuitive display, the specific form of the coding and the decoding process are respectively as Figure 2 and Figure 3 shown. Figure 2 In , is the UAV going to the target point to execute the task Figure 3 The decoding process of the coding is shown. That is, how to obtain a specific task execution plan from a coding. The upper part of the figure shows the coding, and the lower part is the decoded plan. Among them, UAV refers to the UAV that executes the task, going to the target point to execute the task .

[0160] Step 2.2: According to the constraint settings and the coding form, transform the objective function into a fitness function;

[0161] In algorithms similar to TLBO, the objective function in the combined optimization model needs to be transformed into the corresponding fitness function to describe the quality of the solution corresponding to a certain coding. Since the problem studied in the present invention has relatively high requirements for the solution time, the present invention also needs to transform the objective function of the combined optimization model proposed in step 1 to obtain a fitness function that can quickly evaluate the quality of the solution corresponding to the coding. On the other hand, since in the TLBO solution process, the generated coding cannot fully guarantee to meet the constraints, it is necessary to introduce a penalty term to indicate whether the decoded solution of the current coding can be implemented. The specific form is as follows:

[0162] (14)

[0163] (15)

[0164] (16)

[0165] Among them, represents the fitness function; and To adjust the weights of two main indicators and satisfy ; represents the penalty term for the part of the decoding scheme that exceeds the flight time constraint; is the penalty term for the part of the decoding scheme that exceeds the ammunition loading constraint; represents the estimated maximum number of drones in use; represents the maximum flight time of the drone; represents the drone in the decoding scheme flight time; represents the penalty coefficient of the penalty term for the flight time constraint part; represents the penalty coefficient of the penalty term for the ammunition loading constraint; represents the drone number of attack missions executed; represents the maximum ammunition loading capacity of the drone.

[0166] In formula (14), the fitness function consists of three parts. The first part (the first two terms in the formula) represents the objective function given in step 1, that is, minimizing the number of drones used and the task completion time. The second part represents the penalty for all drones that exceed the flight time limit in the encoded decoding scheme. The third part represents the penalty for all drones that exceed the maximum ammunition loading capacity in the encoded decoding scheme.

[0167] Step 3: Based on the encoding setting method, design a high-quality initial solution generation scheme to provide a high-quality initial solution for the teaching and learning algorithm. Generate multiple encodings through this strategy and finally aggregate each encoding to obtain the initial population;

[0168] Considering the temporal relationship between tasks in the same target point, which makes the corresponding high-quality feasible solution space narrow. To improve the search efficiency of the algorithm, the present invention designs a generation strategy for high-quality initial solutions, and the specific process is as follows:

[0169] First, according to the set of targets to be struck, copy and generate an offensive mission target set and an evaluation mission target set, where the two generated sets are exactly the same as the targets in the target set. And take the ceiling of the number obtained by dividing the number of targets by the ammunition load as the initial number of drones and store it in a specific set, named the starting set here. Without loss of generality, the subsequent process takes the offensive mission target set as an example, and the allocation process of the evaluation mission target set is the same as the following process.

[0170] Step 3.1, randomly select a target from the offensive mission target set;

[0171] Step 3.2, calculate the time required for each unmanned aircraft in the motive set to perform the task of the selected target in Step 3.1, including the navigation time from the previous target to the selected target and the task execution time;

[0172] Step 3.3, on the premise of satisfying Formula (11) and Formula (17), select the unmanned aircraft with the minimum time required to perform the task. If no executable unmanned aircraft can be found, add an unmanned aircraft to the motive set and assign this task to the newly added unmanned aircraft. Among them, Formula (17) is as follows:

[0173] (17)

[0174] Among them, represents the current navigation time of the unmanned aircraft ; is a manually set parameter used to balance the time for an unmanned aircraft to perform an offensive task and an evaluation task; represents the maximum flight time of the unmanned aircraft; the above formula means that the time for the unmanned aircraft to perform an attack task shall not exceed a certain ratio of its own maximum flight time. Different will have different impacts on the initial solution.

[0175] Step 3.4, delete the target randomly selected in Step 3.1 from the offensive task target set, and convert the corresponding task execution information into encoded information and store it on the encoding, that is, a column on the encoding introduced in Step 2.1;

[0176] Step 3.5, repeat Steps 3.1 to 3.4 until the offensive task target set becomes an empty set;

[0177] Step 3.6, perform a process similar to Steps 3.1 to 3.4 on the evaluation task target set. The difference is that the iterative process for the offensive task target set in Steps 3.1 to 3.4 is replaced by the evaluation task target set. In addition, when the unmanned aircraft with the minimum time required to perform the evaluation task is also the unmanned aircraft for performing the attack task of this target, then select the unmanned aircraft with the second smallest required time or add an unmanned aircraft again. Finally, when performing the evaluation task assignment, in Formula (17) takes the value of 1.

[0178] After determining the initial solution generation strategy, use the above strategy to generate several initial solutions as the search starting point of TLBO, where the number of generated initial solutions is determined by the preset parameter ; We call several generated initial solutions combined into a population, and That is the population size. Each solution in the population is called an individual, and the above terms will be used subsequently. Finally, for each generated individual, the fitness function proposed in Step 2 is used to calculate the fitness of each individual and save it. According to the setting of the fitness function, the lower the fitness of an individual, the better the corresponding execution plan.

[0179] Step 4: Based on the initial population, determine the encoding update rule, that is, the evolution method of individuals in the teaching and learning stages of the TLBO algorithm, to achieve the self-evolution of the algorithm and obtain the evolved population.

[0180] The evolution of individuals in the population is essentially the exchange of information in the encoding between individuals and the independent exploration of new encodings by individuals. In Step 4, the way of exchanging information in the encoding between individuals is mainly given, that is, the education stage and the learning stage in the TLBO algorithm. Before the TLBO starts to solve, an iteration number will be preset, that is, after the initial population is generated, the number of times the entire population evolves. In one iteration, each individual in the population will execute the education stage and the learning stage mentioned in Step 4 once, and the neighborhood search proposed in the subsequent Step 5. After all individuals complete the above three processes, the fitness function value of each individual is solved using the fitness function proposed in Step 2, and the individual with the lowest function value is selected as the optimal individual of the population, that is, the teacher.

[0181] Step 4.1: Education stage;

[0182] Since the evolution method for encoding in the traditional TLBO algorithm is not applicable to the integer encoding proposed in Step 2, the present invention proposes an evolution method between individuals applicable to this problem. The specific update method is as follows:

[0183] Based on the encoding form set in Step 2, assume there are two individuals, individual and individual . First, randomly select target points from each target point, and subsequently, the selected target points will be used to guide the information exchange between the two individuals. Without loss of generality, if individual learns from individual , locate the positions of the previously selected targets in the encoding of individual , and use this to guide individual to perform the replacement of the encoding positions. The specific operation is as follows: First, create an empty encoding with the same size as individual . Traverse the selected target points one by one. Let the currently traversed target point be . At this time, access individual Encoding, find Among The position where it appears in the encoding (appears twice in total, once representing an attack mission. Assume its position in the encoding at this time is in the column. Once representing an evaluation mission. Assume its position in the encoding at this time is in the column). Copy All the information in the column of the encoding (including the target point and the drone number corresponding to the mission execution) to The column of Similarly, copy all the information in the column of the encoding to The column of At this time, the information exchange of the target point is completed. Repeat the above process until all the information of the selected target points on the encoding is copied to

[0184] At this time, there are still target points whose mission information has not been stored in . At this time, access the encoding, and traverse it backward from the first column. If the target point recorded in the first row of the current column of the encoding belongs to one of the remaining target points, copy the information of this column to the blank column with the earliest position in . When there are no more blank columns, the information exchange process ends. Is The individual obtained after learning. The above process is visually demonstrated in Figure 4

[0185] In each round of iteration, the teacher will conduct an education process for each individual in the population. Assume is the teacher in the population, and a certain individual in the population is , then the process of the student individual learning from the teacher is shown in formula (18):

[0186] (18)

[0187] Among them, represents the individual encoding obtained after learning. In formula (18), represents the individual learning from the individual described above.

[0188] ​Step 4.2: Mutual learning stage;

[0189] In the learning stage, individuals recognize their own gaps through pairwise interaction and communication, and improve their knowledge levels accordingly. This stage is carried out after all individuals have completed the education stage, and each individual has to conduct mutual learning once. The mechanism is as follows:

[0190] Let the currently selected student individual be , randomly select an individual in the population for communication and learning. denotes the fitness function of the individual. The communication process of the student individual is shown in Equation (19):

[0191] (19)

[0192] When the fitness of the individual is greater than , then directly retain the result after learning;

[0193] When the fitness of the individual is less than , first obtain , and then calculate the fitness of . If its fitness does not develop towards a lower value, then regard this learning as ineffective learning, and degrade the individual to before learning, that is ; if the fitness successfully develops towards a lower value, then retain the effect of this learning.

[0194] When each individual in the population has completed the education stage and the mutual learning stage, then carry out Step 5.

[0195] Step 5: Introduce a neighborhood search operator during the evolution process, adaptively carry out neighborhood search operations, further improve the search efficiency of TLBO, enhance the search ability of the population, and finally obtain a solution;

[0196] Since Step 4 focuses on information exchange within the population, it is also necessary to supplement the "exploration" ability of TLBO in the solution space. After the process of education between the teacher and students and the mutual learning among students, the present invention endows individuals with autonomous thinking and learning behaviors on this basis, that is, individuals can expand to a certain extent based on the current knowledge, further improving the efficiency of finding the optimal solution.

[0197] There are three search behaviors in total, specifically: (1) randomly select a column in the encoding, and randomly select a drone from the replaceable drone pool to replace the drone number at this gene position; (2) select a column of the encoding, and randomly select a gene position within the corresponding feasible region for exchange; (3) select a column of the encoding, and randomly select any valuable insertion position within the corresponding feasible region for insertion. In each round of iteration, each individual will perform a neighborhood search, and the above three search behaviors are all executed with a certain probability. After all are completed, the search results are output and the corresponding results are retained. Some specific information about the above three search behaviors is as shown in Figure 5 、 6 、7, and 8.

[0198] Figure 10 、 11 、12 show the Gantt charts of drone tasks under three examples. Among them, the light yellow rectangles in the figure represent the flight time of the drones, the blue rectangles with the number 1 attached after the number represent the time for the drones to perform attack tasks, the dark yellow rectangles with the number 2 attached after the number represent the time for the drones to perform evaluation tasks, and the light gray rectangles without numbers represent the waiting time of the drones. Through the solution of this part, we can calculate the number of drones required for each delivery point and obtain the task execution order of each drone. As shown by the Gantt chart, the task execution time of each aircraft and the number of attack tasks executed obviously satisfy the maximum flight time constraint and ammunition loading constraint of the drones. At the same time, the task execution time of each drone is generally balanced and the task allocation scheme is reasonable. In addition, we can obtain the task completion time of the target points with time constraints, and then we can calculate the task time window of each delivery point. Figure 13 、 14 、15 respectively show the convergence curves under three examples, and it can be seen that the algorithm achieves rapid convergence in the initial stage of iteration.

[0199] In summary, the simulation results show that the method proposed in the present invention comprehensively analyzes the location information of the delivery points, the target points to which the delivery points belong, and the resource information of the drones, and can reasonably allocate tasks to each drone. The obtained task allocation scheme takes into account both the minimum number of drones used and the highest utilization rate of drone resources. On this basis, it also ensures the shortest execution time of the tasks. It fully demonstrates the efficiency and rationality of TLBO.

[0200] The above embodiments only represent the implementation manners of the present invention, but should not be construed as limiting the scope of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A collaborative task planning method for unmanned aerial vehicle swarms based on an adaptive search teaching and learning algorithm, characterized in that The method for planning a collaborative mission of a drone swarm comprises the following steps: Step 1: Extract the key factors of actual constraints, clarify the decision variables and constraints to build a combined optimization model; Step 2: Construct a coding form around the decision variables and constraint settings in the model; then, based on the coding form, transform the objective function into a fitness function suitable for the coding, which is used to evaluate the quality of the solution corresponding to the coding; Step 3: Based on the encoding setting method, a high-quality initial solution generation scheme is designed to provide a high-quality initial solution for the teaching and learning algorithm. Multiple encodings are generated through this strategy, and finally the initial population is obtained by summarizing the encodings. Step 4: Based on the initial population, determine the coding update rule, that is, the evolutionary method of the coding in the teaching and learning phases of the teaching and learning algorithm, realize the self-evolution of the algorithm, and obtain the evolutionary population; Step 5: Introduce the neighborhood search operator in the evolutionary process and adaptively perform neighborhood search operations to further improve the search efficiency of the algorithm, enhance the search capability of the population, and ultimately obtain a solution.

2. The method for collaborative mission planning of an unmanned aerial vehicle swarm based on an adaptive search teaching and learning algorithm according to claim 1, wherein The step 1 is specifically as follows: First, extract the key factors and problem-solving requirements in the actual constraints, clarify the decision variables, and construct a combinatorial optimization model with constraints. Let the task objective set be , represent the number of drones that can be allocated, represent the number of target points, represent the set of drones, represent the set of target points. At this time, all the positions passed by the drones are ; among them, represent the th drone; represent the th target point, and then carry out model representation; Step 1.1: Construct decision variables and related representations according to solution requirements; In the combinatorial optimization problem, the decision variable is the desired task solution, which needs to include various information of the solution to be solved. In this problem, the information to be solved includes the drone number used and the execution task of each drone, so the decision variables are set as follows: (1) In formula (1), represents the decision variable of this problem, represents the task type, represents the drone starting from the target point and going to the target point to execute the task. This formula indicates that each task at each target point needs to be completed. Correspondingly, the following formula is obtained from the decision variable: (2) Formula (2) is used to represent the flight route of the drone. When it indicates that the drone starts from the target point and heads to the target point . Obviously, when it must be that . In addition, there is also: (3) Formula (3) is used to represent the state of the drone. When , the drone needs to be dispatched to perform a mission. Step 1.2: Construct task completion guarantee constraints based on task requirements; (4) Formula (4) indicates that each task for each target point needs to be completed, where it indicates that there are two types of tasks, namely attack tasks and assessment tasks, and there are two types of tasks that need to be executed by the UAVs for each target point; Step 1.3: Construct task execution requirement constraints based on task execution requirements; (5) Formula (5) indicates that for the drone , at most only one task for the same target point can be executed; Step 1.4: Construct navigation constraints according to the UAV’s navigation status; (6) Formula (6) represents that for any drone , starting from the delivery point and finally returning to the delivery point, where and the 0 in the subscript represents the delivery point; similarly, there is also the following formula: (7) Formula (7) indicates that for any UAV , after traveling from any target point to target point , the subsequent journey must start from target point , that is, the out-degree of each target point is equal to the in-degree; Step 1.5: Construct task completion constraints based on task execution requirements; (8) Formula (8) indicates that both types of tasks must be executed when each target point completes them; Step 1.6: Construct the task execution time constraint based on the actual flight rules and execution capabilities of the drone; (9) Formula (9) indicates that there is a sequence of tasks for each target point and there is a minimum interval time, where is the task execution time, is the minimum interval time between two tasks, and respectively represent the start time of the attack task and the start time of the evaluation task for the target point . Step 1.7: Construct the UAV capability constraints (maximum flight time constraints, ammunition loading constraints) according to the UAV's own capability limitations; (10) Equation (10) represents the flight time constraint of the UAV, that is, the flight time of any UAV should be less than its maximum flight time, where represents the flight time of the UAV , represents the maximum flight time of the UAV ; In addition, the ammunition constraint of the drone needs to be set. Assume that the drone consumes one ammunition each time it performs an attack mission. The formula is as follows: (11) Formula (11) represents the ammunition loading constraint of the UAV, that is, the UAV The number of attack missions executed needs to be less than its ammunition loading limit, where represents the maximum ammunition loading quantity of the UAV; Step 1.7: Comprehensively consider the task requirements and complete the expected construction of the combined optimization objective function; (12) The objective function of this problem is shown in Equation (12), where represents the estimated maximum number of drones used, represents the drone 's maximum flight time, and are used to adjust the weights of each item and satisfy ; the optimization index of this objective function is to minimize the number of drones used and the task completion time.

3. A method for collaborative mission planning of a drone swarm based on an adaptive search teaching and learning algorithm according to claim 2, wherein In the formula (12), is obtained by the following formula, and the specific form is as follows: (13) Among them, represents the straight-line distance from the delivery point to the target point, represents the number of target points, represents the flight range of the drone, represents the target point, represents the set of target points.

4. A method for collaborative mission planning of a drone swarm based on an adaptive search teaching and learning algorithm according to claim 2, characterized in that The step 2 is specifically as follows: Step 2.1: Construct algorithm coding based on decision variables; Based on the decision variables and constraint settings, the corresponding teaching and learning algorithm coding is constructed, and the teaching and learning algorithm is collectively referred to as the TLBO algorithm; The encoded information includes the target and the unmanned aerial vehicle, which is generically described as UAV, and the task sequence of the same UAV is executed in the encoded order from front to back; the encoding contains two lines of information, the first line stores the target points, and the second line stores the UAV numbers. The length of each line is twice the number of target points, which is regarded as a 2-row matrix of columns; In the first line of code, when the target appears for the first time, it is represented as the attack task of the point; therefore, the second appearance of the target represents the evaluation task; and the second line represents the drone that performs the task, and the task it performs is the task indicated by the first line of code in the same column; the decision variables of the combined optimization model proposed above are stored through such codes; and the decision variables corresponding to the code determine the complete execution plan of a round of tasks, and the task execution plan obtained by the code is called the decoded plan; Step 2.2: According to the constraint setting and encoding form, transform the objective function into a fitness function; Modify the objective function of the combinatorial optimization model proposed in Step 1 to obtain a fitness function that can quickly evaluate the quality of the encoding corresponding scheme; a penalty term also needs to be introduced to indicate whether the decoding scheme of the current encoding can be implemented, and the specific form is as follows: (14) (15) (16) Among them, represents the fitness function; and is used to adjust the weights of two main indicators and satisfies ; represents the penalty term for the part of the decoding scheme that exceeds the flight time constraint; is the penalty term for the part of the decoding scheme that exceeds the ammunition loading constraint; represents the estimated maximum number of drones used; represents the maximum flight time of the drone; represents the drone in the decoding scheme flight time; represents the penalty coefficient of the penalty term for the flight time constraint part; represents the penalty coefficient of the penalty term for the ammunition loading constraint; represents the drone number of attack missions executed; represents the maximum ammunition loading capacity of the drone; In formula (14), the fitness function consists of three parts. The first two terms in the formula are the first part, which represents the objective function given in step 1, that is, minimizing the number of drones used and the task completion time; the second part represents the penalty for all drones that exceed the flight time limit in the decoding scheme of the encoding; the third part represents the penalty for all drones that exceed the maximum ammunition load in the decoding scheme of the encoding.

5. The method for collaborative mission planning of an unmanned aerial vehicle swarm based on an adaptive search teaching and learning algorithm according to claim 4, wherein The specific content of Step 3 is as follows: First, copy and generate an offensive mission target set and an evaluation mission target set according to the target set to be attacked, where the above two generated sets are exactly the same as the targets in the target set; and take the ceiling of the number obtained by dividing the number of targets by the ammunition load as the initial number of UAVs, and store it in a specific set, named the motivation set here; taking the offensive mission target set as an example, the allocation process of the evaluation mission target set is the same as this process; Step 3.1, randomly select a target from the offensive mission target set; Step 3.2, calculate the time required for each UAV in the motivation set to execute the mission of the target selected in Step 3.1, including the navigation time from the previous target to the selected target and the mission execution time; Step 3.3, on the premise of satisfying Formula (11) and Formula (17), select the UAV with the minimum mission execution time. If no executable UAV can be found, add a UAV to the motivation set and assign this mission to the newly added UAV; where Formula (17) is shown as follows: (17) Among them, represents the current flight time of the drone ; is a manually set parameter used to balance the time for a drone to perform offensive tasks and evaluation tasks; represents the maximum flight time of the drone; the above formula indicates that the time for the drone to perform attack tasks shall not exceed a certain ratio of its own maximum flight time, and different will have different impacts on the initial solution; Step 3.4, delete the target randomly selected in Step 3.1 from the offensive mission target set, and convert the corresponding mission execution information into encoded information and store it on the encoding, that is, a column on the encoding introduced in Step 2.1; Step 3.5, repeat Steps 3.1 to 3.4 until the offensive mission target set becomes an empty set; Step 3.6, perform a process similar to Steps 3.1 to 3.4 on the set of evaluation task objectives, except that the iterative process for the set of offensive task objectives in Steps 3.1 to 3.4 is replaced by the set of evaluation task objectives; in addition, when the UAV with the minimum time required to perform the evaluation task is also the UAV for performing the target attack task, then select the UAV with the second smallest required time or add UAVs again; finally, when performing the evaluation task allocation, in Equation (17), takes a value of 1; After determining the initial solution generation strategy, use the above strategy to generate a number of initial solutions as the search starting point of TLBO, where the number of generated initial solutions is determined by the preset parameter ; A number of generated initial solutions are combined into a population, and is the population size. Each solution in the population is called an individual; Finally, for each generated individual, use the fitness function proposed in step 2 to calculate the fitness of each individual and save it.

6. A method for collaborative mission planning of a drone swarm based on an adaptive search teaching and learning algorithm according to claim 5, characterized in that, The specific content of Step 4 is as follows: In Step 4, mainly give the information exchange method between individuals on the encoding, that is, the education stage and the learning stage in the TLBO algorithm; in the stage before the start of the TLBO solution, a preset number of iterations is set, that is, after generating the initial population, the number of times the entire population evolves; in one round of iteration, each individual in the population will execute the education stage, learning stage mentioned in Step 4, and the neighborhood search proposed in subsequent Step 5 once; after all individuals complete the above three processes, use the fitness function proposed in Step 2 to solve the fitness function value of each individual, and select the individual with the lowest function value as the optimal individual of the population, that is, the teacher; Step 4.1: Education stage; Propose an individual evolution method applicable to this problem; Step 4.2: Mutual learning stage; In the learning stage, individuals understand their own gaps through pairwise interaction and communication, and improve their knowledge level accordingly; this stage is carried out after all individuals complete the education stage, and each individual has to conduct mutual learning once.

7. A method for collaborative mission planning of unmanned aerial vehicle swarms based on an adaptive search teaching and learning algorithm according to claim 6, wherein In Step 4.1, the specific update method of the individual evolution method is as follows: Based on the coding form set in step 2, assume there are two individuals, individual and individual ; First, randomly select target points from each target point, and subsequently use the selected target points to guide the information exchange between the two individuals; If individual learns from individual , locate the positions of the previously selected targets in the coding of individual , and use this to guide individual to perform coding position replacement; The specific operation is as follows: First, create an empty coding with the same size as individual ; Traverse the selected target points one by one. Let the currently traversed target point be . At this time, access the coding of individual , and find the position where appears in the coding of ; Copy all the information in the column of the coding of to the th column of ; Similarly, copy all the information in the column of the coding of to the th column of ; At this time, the information exchange for the target point is completed. Repeat the above process until all the information of the selected target points on the coding is copied to ; At this time, the remaining task information of the target points has not been stored in . At this time, access 's encoding, and traverse its encoding from the first column backward. If the target point recorded in the first row of the current column of the encoding belongs to one of the remaining target points, then copy the information of this column to the blank column with the earliest position; when there is no blank column anymore, the information exchange process ends; That is the individual obtained after learning is completed; In each round of iteration, the teacher will conduct an educational process for each individual in the population. Suppose is the teacher in the population, and a certain individual in the population is , then the process of the student individual learning from the teacher is shown in formula (18): (18) Among them, represents the individual coding obtained after learning. In formula (18), represents the individual described above towards the individual learning process.

8. A method for collaborative mission planning of an unmanned aerial vehicle swarm based on an adaptive search teaching and learning algorithm according to claim 7, characterized in that In Step 4.2, the mechanism of the mutual learning stage is as follows: Let the currently selected student individual be , randomly select an individual in the population for communication and learning. represents the fitness function of the individual. The communication process of the student individual is shown in Equation (19) as follows: (19) When the individual fitness is greater than then directly retain the result after learning; When the individual has a fitness less than , first obtain , and then calculate the fitness of . If its fitness does not develop towards a lower value, consider this learning as ineffective learning and degrade the individual to its state before learning, that is ; If the fitness successfully develops to a lower level, retain the learning effect of this time; When each individual in the population completes the education stage and the mutual learning stage, Step 5 is carried out.

9. The method for collaborative mission planning of an unmanned aerial vehicle swarm based on an adaptive search teaching and learning algorithm according to claim 8, wherein The specific content of Step 5 is as follows: After the process of education between teachers and students and mutual learning among students, the individual expands based on the current knowledge, including three specific search behaviors, namely: (1) randomly select a column in the encoding, and randomly select a drone from the replaceable drone pool to replace the drone number at this gene position; (2) select a column of the encoding, and randomly select a gene position within the corresponding feasible region for exchange; (3) select a column of the encoding, and randomly select any valuable insertion position within the corresponding feasible region for insertion; In each round of iteration, each individual will conduct a neighborhood search, perform the above search behaviors, output the search results after all are completed, and retain the corresponding results.

Citation Information

Patent Citations

  • Task planning method of multistatic heterogeneous unmanned aerial vehicle cooperative reconnaissance

    CN107330588A

  • Task planning method for cooperative ground battle of manned and unmanned aerial vehicles

    CN112733421A

  • Multi-unmanned aerial vehicle task allocation method based on multi-target quantum krill group mechanism

    CN112926825A

  • Improved-firefly-algorithm-based multi-unmanned-aerial-vehicle cooperative coupling task distribution method

    CN107219858A

  • Multi-satellite task scheduling planning method based on improved teaching optimization method

    CN112668930A