An autonomous learning method for a large model for crowd wisdom strategy generation of an unmanned system
By constructing a large-scale model for generating unmanned system swarm intelligence strategies and an autonomous learning method, the problem of low efficiency and slow response of traditional unmanned swarms in complex task scenarios has been solved. This has enabled efficient collaboration and resource optimization of unmanned swarms in complex environments, promoting intelligent and autonomous collaborative task execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2025-06-10
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional unmanned swarm collaborative planning methods are difficult to adapt to complex and ever-changing task scenarios, resulting in low efficiency, slow response, inability to effectively deal with emergencies, serious waste of resources, and cumbersome human-machine collaboration, making it difficult to meet the needs of time-sensitive tasks.
We construct a large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems. By establishing relationship models, path planning models, resource consumption models, and task completion models among intelligent agents, and combining them with a multi-objective particle swarm optimization algorithm, we can achieve autonomous collaboration and resource optimization of unmanned swarms.
It significantly improves the flexibility and adaptability of unmanned swarms in path planning in complex environments, optimizes resource utilization efficiency, increases task completion rate, and enables intelligent and autonomous collaborative task execution.
Smart Images

Figure CN120595597B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) data allocation and processing technology, specifically to a method, device, and storage medium for autonomous learning of large models for generating swarm intelligence strategies for unmanned systems. Background Technology
[0002] Unmanned swarms are widely used in logistics, emergency rescue, and industrial manufacturing. The effectiveness and reliability of their collaborative planning are crucial, but the actual mission scenarios are complex and varied, and full of uncertainties.
[0003] For example, in logistics scenarios, urban low-altitude delivery needs to cope with high-density airspace management, sudden weather changes and traffic congestion, while also facing temporary changes in customer demand or urgent orders; traditional path planning algorithms are mainly based on fixed routes and cannot flexibly adjust to cope with emergencies; the stand-alone operation mode is prone to path conflicts or resource waste, and does not fully consider the drone's endurance limitations and charging scheduling issues, resulting in low mission efficiency.
[0004] For example, in emergency rescue scenarios, the disaster site environment is harsh (such as thick smoke, strong winds, and high temperatures), and secondary disasters (such as secondary collapses and fire spread) threaten the stability of equipment. In addition, multiple tasks need to be carried out in parallel (search and rescue, material delivery, and communication relay). Traditional task allocation algorithms (such as the Hungarian algorithm) are difficult to handle task priority conflicts, environmental perception relies on a single sensor or manually labeled data, and the response is lagging. Human-machine collaboration requires frequent human intervention, which is difficult to meet the rapid response requirements of time-sensitive tasks.
[0005] For example, in industrial manufacturing scenarios, unmanned transport vehicles (AGVs), robotic arms and other equipment need to work together to complete complex production processes, but supply chain fluctuations (such as raw material delays and order changes) and equipment failures occur frequently; traditional production scheduling systems (such as JIT) are mainly based on rigid process design and are difficult to adapt to sudden adjustments; local optimization path planning does not consider global resource competition, which can easily lead to task congestion or equipment idleness; human-machine interaction is cumbersome and the level of automation is limited, affecting production efficiency and safety.
[0006] Therefore, in summary, traditional collaborative planning methods are difficult to adapt to these changes, and have exposed problems such as low efficiency and slow response in various fields, affecting task execution.
[0007] To address these technical issues, this application proposes a large-scale autonomous learning method for generating swarm intelligence strategies in unmanned systems. Summary of the Invention
[0008] The main objective of this invention is to provide a large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems, which is used to perform dynamic modeling operations of generative large-scale models for intelligent packaging and testing scenarios, in order to solve the technical problems mentioned in the background art.
[0009] The present invention solves the above-mentioned technical problems by adopting the following technical solutions:
[0010] An autonomous learning method for large-scale models of swarm intelligence strategy generation in unmanned systems includes:
[0011] S1. Divide tasks into multiple categories, construct a relationship model between any number of intelligent agents, and construct a coordination model based on the cooperation relationship between intelligent agents when performing different tasks, as well as the efficiency relationship between intelligent agents and different tasks. This model is used to describe the efficiency relationship, resource consumption relationship, and path relationship when any n intelligent agents cooperate to complete a certain task.
[0012] S2. Based on the collaborative model, establish a path planning model for the agent when performing tasks, and realize real-time update of path planning when the agent performs tasks to adapt to complex task environments;
[0013] S3. Based on the collaborative model, statistical analysis of the resource consumption of intelligent agents when performing tasks is conducted, a resource consumption model is constructed, and an optimization algorithm is designed to optimize resource utilization efficiency and reduce resource waste.
[0014] S4. Based on the collaborative model, statistical analysis of the task completion status of the agent when performing tasks is conducted, a task completion model is constructed, and an optimization algorithm is designed to optimize the task completion rate and improve task execution efficiency.
[0015] S5. Combine the path planning model, resource consumption model and task completion model to construct an optimization model, and solve the model through a multi-objective particle swarm optimization algorithm to ensure that the unmanned swarm can coordinate efficiently when performing tasks.
[0016] S6. Based on the optimization model, by preparing task data, agent data, resource consumption data, and task completion data, the model is trained and evaluated to realize the autonomous evolution and autonomous collaborative task execution of unmanned swarm collaborative intelligence.
[0017] Preferably, the process of constructing the relationship model between agents in step S1 includes:
[0018] S11. The agents use drones. For n agents, construct the attribute vector:
[0019]
[0020] in, Indicates the first The comprehensive capabilities of an intelligent agent Indicates the first Maximum flight speed of an intelligent agent Indicates the first Maximum load weight of each intelligent agent No. The largest energy reserve, Representing the Sensor efficiency of individual agents;
[0021] S12. For task class m, define the task feature matrix:
[0022]
[0023] in, Indicates task The degree of urgency, Indicates task The accuracy requirements, Task Energy consumption ;
[0024] S13. Construct cooperative relationships between agents, whose cooperative efficiency tensor is represented as:
[0025]
[0026] in, Indicates the first The first intelligent agent and the first Tasks are executed between the intelligent agents. The probability of adaptation, Indicates task The collaboration weight matrix, This represents the Sigmoid activation function;
[0027] S14. Construct the task adaptation function between the agent and the task, which is:
[0028]
[0029] in, Indicates the first The first intelligent agent and the first The degree of suitability for each task Indicates the first Each task feature vector matrix Indicates the attenuation coefficient;
[0030] S15. Establish the relationship between the number of completed tasks and the number of agents; the group cooperation efficiency function is:
[0031]
[0032] in, Represents a subset of intelligent agents and the tasks performed The relationship of synergistic effectiveness Indicates the cooperative gain coefficient. Indicates the number of agents in the subset. This represents the scale penalty coefficient.
[0033] Preferably, in step S1, the coordination model is constructed using a dynamic programming value function, the calculation formula of which is:
[0034]
[0035] in, Indicates in At that moment, the state was The maximum expected return, Represents path cost, Indicates the cost of resources. and These represent the weights of path cost and resource cost, respectively. Represents the expected value of a future state;
[0036] There are constraints:
[0037]
[0038] in, This indicates the minimum threshold for collaborative effectiveness. This indicates the amount of key resource reserves.
[0039] Preferably, the process of establishing the path planning model in step S2 includes:
[0040] S21. Parameterize the agent's path trajectory; its formula is expressed as:
[0041]
[0042] in, Represents intelligent agents At any moment The three-dimensional position, , , Indicates in The three-dimensional coordinates of the agent at any given moment Represents the orbital parameter vector. No. One basis function;
[0043] S22. Establish the path cost function, the expression of which is:
[0044]
[0045] in, , , These represent path smoothing weight, resource distance weight, and resource sensitivity weight, respectively. , Representing intelligent agents respectively , The location.
[0046] Preferably, the mathematical expression for the resource consumption model in step S3 is:
[0047]
[0048]
[0049]
[0050] in, Indicates completion of the task Energy consumption in , The correlation coefficient of the drone is represented. Indicates the speed of the drone. The initial velocity, For the acceleration of the drone, It is air density. It is the zero-lift drag coefficient. It is the wing area. It is the Oswald efficiency factor. Wing aspect ratio, It's about the quality of the drone. It is the acceleration due to gravity. This represents the total time spent on a drone mission.
[0051] There are constraints:
[0052]
[0053] in, and These represent the initial and final positions of the drone, respectively. and Indicates maximum speed and maximum acceleration. , Indicates the initial velocity and the final velocity.
[0054] Preferably, the formula for constructing the task completion model T in step S4 is:
[0055]
[0056]
[0057] in, This represents the set of subtasks for the k-th type of task. This represents the i-th subtask of the k-th task. A collection of drones assigned to perform sub-tasks. For spatiotemporal constraint coefficients, Remaining time for the task Indicates the maximum speed. As initial energy, This indicates the energy consumption of the task.
[0058] Preferably, the optimization model construction process in step S5 includes:
[0059] S51. Integrating the state variables of the collaboration, path, resource, and task completion model, we have:
[0060]
[0061] There exists a multi-objective value function as an optimization model. :
[0062]
[0063]
[0064]
[0065] in; For task completion rate, For learning rate, For activation function, For the task change rate, For resource consumption, This serves as a baseline value for resource consumption. Hyperbolic tangent function, For location variance, For path smoothness, To achieve minimum collaborative efficiency, For safety reasons, For maximum energy distance, Urgency level coefficient Key indicators Weighting coefficients representing task completion rate The weighting coefficients representing energy consumption. Weighting coefficients representing path smoothness.
[0066] S25. Solve using improved multi-objective particle swarm optimization:
[0067]
[0068]
[0069]
[0070]
[0071]
[0072] in; Represents particles The Vi in the The speed of updates after each iteration Represents particles The Dimensional speed The speed of updates after each iteration Represents particles The Vidi The position updated after the next iteration Represents particles The Vidi The position updated after the next iteration Represents dynamic inertia weight. , Represents the adaptive learning factor. , These represent the minimum and maximum values of the cognitive learning factor, respectively. , These represent the minimum and maximum values of the social learning factor, respectively. , It is a random number. This represents the optimal position of an individual. Indicates the globally optimal position. , This represents the upper and lower limits of the inertia weight. Indicates the current iteration number. This indicates the maximum number of iterations.
[0073] Preferably, the specific operation process of model training in step S6 includes:
[0074] S61. Constructing a multimodal fusion formula:
[0075]
[0076] in This represents the large model of the base. For cross-model attention mechanisms, Encode historical decisions;
[0077] S62. An attention mechanism exists:
[0078]
[0079] S63. Building Strategy - Dual-End Value Output:
[0080]
[0081] in, , For strategy header parameters, For value head parameters, As a discount factor, It expresses an expectation of the future state;
[0082] S64. Calculate the loss function L
[0083]
[0084] in, It is an expectation. To predict returns for the model, To optimize overall returns, This refers to the KL divergence. This refers to the probability distribution of the current strategy. This refers to the probability distribution of expert strategies. For random variables variance The distribution coefficient, The hidden coefficient;
[0085] S65. Implement gradient collaborative update:
[0086]
[0087] in, For gradient projection, For collaborative learning rate, Indicates time The synergistic effectiveness matrix.
[0088] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0089] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0090] As can be seen from the above technical solution, this invention provides a large-scale autonomous learning method for generating swarm intelligence strategies in unmanned systems. Compared with existing technologies, this invention has the following advantages:
[0091] 1. By analyzing the collaborative mechanisms under different tasks, this invention can establish a basic collaborative framework for multi-agent systems, which facilitates the design of dynamic adjustment mechanisms by combining task changes and agent states, and establishes a path planning model that can be updated in real time, significantly improving the flexibility and adaptability of unmanned swarms in path planning in complex environments.
[0092] 2. This invention constructs a resource consumption model and a task completion model, quantifies resource utilization efficiency and optimizes task execution based on experimental / simulation data, and can form a complete optimization system by combining comprehensive path planning, resource consumption and task completion models. It also adopts an improved multi-objective particle swarm optimization algorithm to achieve multi-dimensional collaborative optimization, ensuring efficient collaboration while taking into account resource utilization and task completion rate, and comprehensively improving the task execution efficiency of unmanned swarms.
[0093] 3. This invention innovatively designs a strategy-value dual-head output architecture by training a large model through multimodal fusion and combining task data and agent data. This promotes the autonomous evolution of unmanned swarm collaborative intelligence, realizes intelligent and autonomous collaborative task execution, and provides an innovative solution with environmental adaptability and resource optimization capabilities for complex scenarios.
[0094] It should be understood that the descriptions in this section are not intended to identify key or essential features of embodiments of the invention, nor are they intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Of course, implementing any product of the invention does not necessarily require achieving all of the advantages described above simultaneously. Attached Figure Description
[0095] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0096] Figure 1 This is a schematic diagram of the overall operation process of the present invention. Detailed Implementation
[0097] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0098] For details in the embodiments, please refer to Figure 1 .
[0099] like Figure 1 As shown. The autonomous learning method for large-scale models of swarm intelligence strategy generation for unmanned systems proposed in this embodiment of the invention includes the following steps:
[0100] Step 1: Classify tasks into m categories, such as emergency tasks, strike tasks, and reconnaissance tasks. Construct a relationship model among any n agents to analyze their cooperation when performing different tasks. For different tasks, study the efficiency relationship between multiple agents and the task, including resource consumption rate and task completion rate. Finally, construct a coordination model to describe the efficiency, resource consumption, and path relationships when any n agents cooperate to complete a task.
[0101] Constructing the drone attribute vector:
[0102]
[0103] in, Indicates the first The overall capabilities of a drone Indicates the first Maximum flight speed of each drone Indicates the first Maximum payload weight of each drone No. The largest energy reserve;
[0104] For task m, define the task feature matrix:
[0105]
[0106] in, Indicates task The degree of urgency, Indicates task The accuracy requirements, Task Energy consumption;
[0107] The collaborative relationship between drones can be represented by the collaborative efficiency tensor as follows:
[0108]
[0109] in , indicating the first The drone and the first The compatibility probability between individual drones Indicates task The collaboration weight matrix, This represents the Sigmoid activation function;
[0110] Construct a task adaptation function between the drone and the task:
[0111]
[0112] in, Indicates the first The drone and the first The degree of suitability for each task Indicates the first Each task feature vector matrix Indicates the attenuation coefficient;
[0113] The relationship between the completed task and the number of drones is established, and the group cooperation efficiency function is:
[0114]
[0115] in, Represents a subset of drones and the tasks performed The relationship of synergistic effectiveness Indicates the cooperative gain coefficient. This indicates the number of drones in the subset. Indicates the size penalty coefficient;
[0116] Dynamic Programming Value Function
[0117]
[0118] in, Indicates in At that moment, the state was The maximum expected return, Represents path cost, Indicates the cost of resources. It represents the expected value of a future state, and the expected return after considering the uncertainty of state transitions;
[0119] Constraint formula:
[0120]
[0121] in, This indicates the minimum threshold for collaborative effectiveness. Indicates the reserves of key resources;
[0122] At this point, by dividing the task into m categories in the initial stage of task planning and constructing a relationship model between any n agents, we can deeply analyze the cooperation mechanism of agents under different tasks, thereby clarifying the basic framework of multi-agent collaboration and finally providing support for the generation of collaborative strategies in complex task scenarios.
[0123] Step 2: Based on the collaborative model, establish a path planning model for the agent when performing tasks. By selecting or designing a suitable path planning algorithm, and combining task changes and agent states, design a dynamic adjustment mechanism to update the path planning in real time to adapt to complex task environments;
[0124] The drone's path trajectory is parameterized, and its formula is expressed as:
[0125]
[0126] in, Indicates drone At any moment The three-dimensional position, , , Indicates in The three-dimensional coordinates of the drone at any time Represents the orbital parameter vector. No. One basis function;
[0127] The path cost function is defined as follows:
[0128]
[0129] in, , , These represent path smoothing weight, resource distance weight, and resource sensitivity weight, respectively. , Indicates drone , Location;
[0130] At this point, by establishing a path planning model based on the collaborative model and designing a dynamic adjustment mechanism that combines task changes and agent states, the path planning can be updated in real time, thereby adapting to complex and ever-changing task environments and improving the flexibility and environmental adaptability of unmanned swarm path planning.
[0131] Step 3: Based on the collaborative model, analyze the resource consumption of the agent when performing tasks. Collect resource consumption data through experiments or simulations, construct a resource consumption model using statistical methods or machine learning algorithms, and design optimization algorithms to optimize resource utilization efficiency and reduce resource waste;
[0132] Energy consumption model:
[0133]
[0134]
[0135]
[0136] in, This indicates the energy consumption during task completion. , The correlation coefficient of the drone is represented. Indicates the speed of the drone. The initial velocity, For the acceleration of the drone, It is air density. It is the zero-lift drag coefficient. It is the wing area. It is the Oswald efficiency factor. Wing aspect ratio, It's the weight of the drone. It is the acceleration due to gravity. This represents the total time spent on a drone mission;
[0137]
[0138] in, and These represent the initial and final positions of the drone, respectively. and Indicates maximum speed and maximum acceleration. , Indicates the initial velocity and the final velocity;
[0139] At this point, by analyzing resource consumption based on a collaborative model, constructing a resource consumption model using experimental or simulation data, and designing optimization algorithms, it is possible to quantify resource consumption, optimize resource utilization efficiency, thereby reducing resource waste and improving the resource utilization efficiency of unmanned clusters.
[0140] Step 4: Based on the collaborative model, analyze the task completion performance of the agent. Collect task completion data through experiments or simulations, construct a task completion model using statistical methods or machine learning algorithms, and design optimization algorithms to optimize the task completion rate and improve task execution efficiency;
[0141]
[0142]
[0143] in, This represents the set of subtasks for the k-th type of task. This represents the i-th subtask of the k-th task. A collection of drones assigned to perform sub-tasks. For spatiotemporal constraint coefficients, Remaining time for the task Indicates the maximum speed. As initial energy, Indicates the energy consumption of the task;
[0144] Step 5: Integrate the path model, resource consumption model, and task completion model to form a complete optimization model. Use an improved multi-objective particle swarm optimization algorithm to solve the model, ensuring efficient collaboration among unmanned swarms during task execution, while balancing resource consumption and task completion rate.
[0145] Integrate the state variables of the collaboration, path, resources, and task completion model:
[0146]
[0147] Multi-objective value function:
[0148]
[0149]
[0150]
[0151] in; For task completion rate, For learning rate, For activation function, For the task change rate, For resource consumption, This serves as a baseline value for resource consumption. Hyperbolic tangent function, For location variance, For path smoothness, To achieve minimum collaborative efficiency, For safety reasons, For maximum energy distance, Urgency level coefficient Key indicators Weighting coefficients representing task completion rate The weighting coefficients representing energy consumption. Weighting coefficients representing path smoothness.
[0152] Solving the problem using an improved multi-objective particle swarm optimization method:
[0153]
[0154]
[0155]
[0156]
[0157]
[0158] in; Represents particles The Vi in the The speed of updates after each iteration Represents particles The Dimensional speed The speed of updates after each iteration Represents particles The Vidi The position updated after the next iteration Represents particles The Vidi The position updated after the next iteration Represents dynamic inertia weight. , Represents the adaptive learning factor. , These represent the minimum and maximum values of the cognitive learning factor, respectively. , These represent the minimum and maximum values of the social learning factor, respectively. , It is a random number. This represents the optimal position of an individual. Indicates the globally optimal position. , This represents the upper and lower limits of the inertia weight. Indicates the current iteration number. This indicates the maximum number of iterations.
[0159] At this point, a complete optimization model is formed by integrating the path model, resource consumption model, and task completion model, and the improved multi-objective particle swarm optimization algorithm is used to solve it. This enables multi-objective collaborative optimization, thereby achieving efficient collaboration while taking into account resource consumption and task completion rate, and improving the overall task execution efficiency of the unmanned swarm.
[0160] Step Six: Based on the above models and algorithms, train a large model to achieve autonomous evolution of unmanned swarm collaborative intelligence. By preparing task data, agent data, resource consumption data, and task completion data, model training and evaluation are performed to ultimately achieve intelligent and autonomous collaborative task execution by the unmanned swarm.
[0161] Multimodal fusion of input data:
[0162]
[0163] in This represents the large model of the base. For cross-model attention mechanisms, Encode historical decisions;
[0164]
[0165] Building Strategy - Value Dual-End Output:
[0166]
[0167] in, , For strategy header parameters, For value head parameters, As a discount factor, It expresses an expectation of the future state;
[0168]
[0169] in, It is an expectation. To predict returns for the model, To optimize overall returns, This refers to the KL divergence. This refers to the probability distribution of the current strategy. This refers to the probability distribution of expert strategies. For random variables variance The distribution coefficient, The hidden coefficient;
[0170] Gradient-coordinated update:
[0171]
[0172] in; For gradient projection, For collaborative learning rate, Indicates time The synergistic effectiveness matrix;
[0173] In summary, by using task data and agent data to train a large model and designing multimodal fusion and a policy-value dual-head output, we can achieve autonomous evolution of unmanned swarm collaborative intelligence, thereby promoting intelligent and autonomous collaborative task execution and ultimately providing innovative solutions for the application of unmanned systems in complex scenarios.
[0174] Based on the collaborative model, path planning, resource consumption, and task completion models are established separately, then combined into an optimization model and solved using a multi-objective particle swarm optimization algorithm. Finally, based on the optimization model, a large-scale model with multimodal fusion is trained and evaluated, which significantly improves the flexibility and adaptability of unmanned swarms in path planning in complex environments, optimizes resource utilization efficiency, increases task completion rate, promotes the autonomous evolution and autonomous task execution of unmanned swarm collaborative intelligence, and solves the problems of low efficiency, slow response, and poor adaptability of traditional collaborative planning methods in scenarios such as logistics, emergency rescue, and industrial manufacturing.
[0175] In one specific embodiment, this application includes the following evaluation table to illustrate the assessment effect:
[0176]
[0177] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0178] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0179] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the above embodiments of the large model autonomous learning method for generating swarm intelligence strategies for unmanned systems.
[0180] It is understood that the system provided in the embodiments of the present invention corresponds to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.
[0181] This application also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus.
[0182] Memory, used to store computer programs;
[0183] When the processor executes the program stored in memory, it implements the above-mentioned autonomous learning method for generating large-scale models of swarm intelligence strategies for unmanned systems.
[0184] The communication bus mentioned in the above-mentioned electronic devices can be a standard bus for interconnecting peripheral components or an extended industrial standard structure bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0185] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0186] The memory may include random access memory or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0187] The processors mentioned above can be general-purpose processors, including central processing units, network processors, etc.; they can also be digital signal processors, application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0188] It should also be noted that electronic devices include terminal devices, which can also be called terminals, user equipment, mobile stations, mobile terminals, etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality terminal devices, augmented reality terminal devices, wireless terminals in industrial control, wireless terminals in autonomous driving, wireless terminals in remote surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and so on. The embodiments of this application do not limit the specific technologies or device forms used in the terminal devices.
[0189] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0190] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0191] Furthermore, it should be noted that if any directional indication (such as up, down, left, right, front, back, etc.) is involved in the embodiments of the present invention, the directional indication is only used to explain the relative positional relationship and movement of each component in a specific posture. If the specific posture changes, the directional indication will also change accordingly.
[0192] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, in the embodiments of this invention, "multiple" refers to two or more. Moreover, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
Claims
1. A large-scale autonomous learning method for generating swarm intelligence strategies in unmanned systems, characterized in that, include: S1. Divide tasks into multiple categories, construct a relationship model between any number of agents, and construct a collaborative model based on the cooperation relationship between agents when performing different tasks and the efficiency relationship between agents and different tasks. S2. Based on the collaborative model, establish a path planning model for the agent when performing tasks; S3. Based on the collaborative model, statistically analyze the resource consumption of the intelligent agent when performing tasks, and construct a resource consumption model; S4. Based on the collaborative model, statistically analyze the task completion status of the intelligent agent when performing tasks, and construct a task completion model; S5. Combine the path planning model, resource consumption model, and task completion model to construct an optimization model, and solve the model using a multi-objective particle swarm optimization algorithm; S6. Based on optimized model training and evaluation, realize the autonomous evolution and autonomous collaborative task execution of unmanned swarm collaborative intelligence; The process of constructing the relationship model between agents in step S1 includes: S11. The agents use drones. For n agents, construct the attribute vector: in, Indicates the first The comprehensive capabilities of an intelligent agent Indicates the first Maximum flight speed of an intelligent agent Indicates the first Maximum load weight of each intelligent agent No. The largest energy reserve, Representing the Sensor efficiency of individual agents; S12. For task class m, define the task feature matrix: in, Indicates task The degree of urgency, Indicates task The accuracy requirements, Task Energy consumption ; S13. Construct cooperative relationships between agents, whose cooperative efficiency tensor is represented as: in, Indicates the first The first intelligent agent and the first Tasks are executed between the intelligent agents. The probability of adaptation, Indicates task The collaboration weight matrix, This represents the Sigmoid activation function; S14. Construct the task adaptation function between the agent and the task, which is: in, Indicates the first The first intelligent agent and the first The degree of suitability for each task Indicates the first Each task feature vector matrix Indicates the attenuation coefficient; S15. Establish the relationship between the number of completed tasks and the number of agents; the group cooperation efficiency function is: in, Represents a subset of intelligent agents and the tasks performed The relationship of synergistic effectiveness Indicates the cooperative gain coefficient. Indicates the number of agents in the subset. This represents the scale penalty coefficient.
2. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems as described in claim 1, characterized in that, In step S1, the coordination model is constructed using a dynamic programming value function, the formula for which the dynamic programming value function is calculated is: in, Indicates in At that moment, the state was The maximum expected return, Represents path cost, Indicates the cost of resources. and These represent the weights of path cost and resource cost, respectively. It represents the expected value of a future state.
3. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems as described in claim 1, characterized in that, The process of establishing the path planning model in step S2 includes: The agent's path trajectory is parameterized, and a path cost function is established, the formula of which is expressed as: in, , , These represent path smoothing weight, resource distance weight, and resource sensitivity weight, respectively. , Representing intelligent agents respectively , The location.
4. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems as described in claim 3, characterized in that, The mathematical expression for the resource consumption model in step S3 is: in, Indicates completion of the task Energy consumption in , The correlation coefficient of the drone is represented. Indicates the speed of the drone. The initial velocity, For the acceleration of the drone, It is air density. It is the zero-lift drag coefficient. It is the wing area. It is the Oswald efficiency factor. Wing aspect ratio, It's about the quality of the drone. It is the acceleration due to gravity. This represents the total time spent on a drone mission.
5. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems as described in claim 3, characterized in that, The task completion model in step S4 The construction calculation formula is: in, This represents the set of subtasks for the k-th type of task. This represents the i-th subtask of the k-th task. A collection of drones assigned to perform sub-tasks. For spatiotemporal constraint coefficients, Remaining time for the task Indicates the maximum speed. As initial energy, This indicates the energy consumption of the task.
6. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems as described in claim 3, characterized in that, The optimization model construction process in step S5 includes: S51. Integrating the state variables of the collaboration, path, resource, and task completion models, a multi-objective value function exists as the optimization model. : in; For task completion rate, For resource consumption, For path smoothness, Weighting coefficients representing task completion rate The weighting coefficients representing energy consumption. Weighting coefficients representing path smoothness; S52. Solve using improved multi-objective particle swarm optimization: in; Represents particles The Vi in the The speed of updates after each iteration Represents particles The Dimensional speed The speed of updates after each iteration Represents particles The Vidi The position updated after the next iteration Represents particles The Vidi The position updated after the next iteration Represents dynamic inertia weight. , Represents the adaptive learning factor. , These represent the minimum and maximum values of the cognitive learning factor, respectively. , These represent the minimum and maximum values of the social learning factor, respectively. , It is a random number. This represents the optimal position of an individual. Indicates the globally optimal position. , This represents the upper and lower limits of the inertia weight. Indicates the current iteration number. This indicates the maximum number of iterations.
7. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems as described in claim 1, characterized in that, The specific operational procedures for model training in step S6 include: S61. Constructing a multimodal fusion formula: in This represents the large model of the base. For cross-model attention mechanisms, Encode historical decisions; S62. An attention mechanism exists: S63. Building Strategy - Dual-End Value Output: in, , For strategy header parameters, For value head parameters, As a discount factor, It expresses an expectation of the future state; S64. Calculate the loss function L: in, It is an expectation. To predict returns for the model, To optimize overall returns, This refers to the KL divergence. This refers to the probability distribution of the current strategy. This refers to the probability distribution of expert strategies. For random variables variance The distribution coefficient, The hidden coefficient; S65. Implement gradient collaborative update: in, For gradient projection, For collaborative learning rate, Indicates time The synergistic effectiveness matrix.