Large model autonomous learning method oriented to unmanned system crowd-sourcing strategy generation
By constructing a large-model autonomous learning method for generating swarm intelligence strategies for unmanned systems, the problems of low efficiency and slow response of unmanned swarms in complex mission scenarios are solved, and efficient and autonomous collaborative task execution and resource optimization are achieved.
Patent Information
- Application Number
- CN202510770525.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Traditional unmanned swarm collaborative planning methods are difficult to adapt to complex and changeable mission scenarios, resulting in low efficiency, slow response, inability to effectively deal with emergencies, serious waste of resources, and cumbersome human-machine collaboration, making it difficult to meet time-sensitive mission requirements.
A large-model autonomous learning method for generating swarm intelligence strategies for unmanned systems is adopted. By constructing a relationship model, path planning model, resource consumption model and task completion model among intelligent agents, combined with a multi-objective particle swarm optimization algorithm, autonomous collaborative task execution of unmanned clusters is achieved.
It significantly improves the path planning flexibility and adaptability of unmanned swarms in complex environments, optimizes resource utilization efficiency, improves task completion rate, and realizes intelligent and autonomous collaborative task execution.
Smart Images

Figure CN120595597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) data distribution and processing, and in particular to a large-model autonomous learning method, device, and storage medium for generating swarm intelligence strategies for unmanned systems. Background Art
[0002] Unmanned swarms are widely used in logistics, emergency rescue, industrial manufacturing and other fields. The effectiveness and reliability of their collaborative planning are crucial, but actual mission scenarios are complex, changeable and full of uncertainty.
[0003] For example, in logistics scenarios, urban low-altitude delivery needs to deal with high-density airspace management, sudden weather changes and traffic congestion, while also facing temporary changes in customer demand or urgent orders; traditional path planning algorithms are mainly based on fixed routes and cannot be flexibly adjusted to deal with emergencies; the stand-alone operation mode of a single machine is prone to path conflicts or waste of resources, and does not fully consider the drone's endurance limitations and charging scheduling issues, resulting in low mission efficiency.
[0004] For example, in emergency rescue scenarios, the disaster site environment is harsh (such as thick smoke, strong winds, and high temperatures), secondary disasters (such as secondary collapse and fire spread) threaten equipment stability, and multiple tasks need to be carried out in parallel (search and rescue, material delivery, and communication relay); traditional task allocation algorithms (such as the Hungarian algorithm) have difficulty handling task priority conflicts, environmental perception relies on a single sensor or manually labeled data, and response is delayed; human-machine collaboration requires frequent manual intervention, which makes it difficult to meet the rapid response requirements of time-sensitive tasks.
[0005] For example, in industrial manufacturing scenarios, automated guided vehicles (AGVs), robotic arms and other equipment need to work together to complete complex production processes, but supply chain fluctuations (such as raw material delays and order changes) and equipment failures are frequent; traditional production scheduling systems (such as JIT) are mainly based on rigid process design and are difficult to adapt to sudden adjustments; local optimization path planning does not consider global resource competition, which can easily lead to task congestion or equipment idleness; human-computer interaction is cumbersome and the level of automation is limited, affecting production efficiency and safety.
[0006] Therefore, in summary, traditional collaborative planning methods are difficult to adapt to these changes, exposing problems such as low efficiency and slow response in various fields, affecting task execution.
[0007] To this end, this application specifically proposes a large-model autonomous learning method for generating swarm intelligence strategies for unmanned systems to solve the above technical problems. Summary of the Invention
[0008] The main purpose of the present invention is to provide a large-model autonomous learning method for generating swarm intelligence strategies for unmanned systems, which is used to perform generative large-model dynamic modeling operations for intelligent packaging and testing scenarios to solve the technical problems raised in the background technology.
[0009] The present invention adopts the following technical solutions to solve the above technical problems: A large-model autonomous learning method for generating swarm intelligence strategies for unmanned systems, including: S1. Classify tasks into multiple categories and construct a relationship model between any number of agents. Based on the coordination relationships between agents when performing different tasks, and the performance relationships between agents and different tasks, a coordination model is constructed to describe the performance relationships, resource consumption relationships, and path relationships when any n agents collaborate to complete a task. S2. Based on the collaborative model, a path planning model is established for the agent when performing tasks, enabling real-time updates of the path planning during task execution to adapt to complex task environments. S3. Based on the collaborative model, calculate the resource consumption of the agent when performing tasks, build a resource consumption model, and design an optimization algorithm to optimize resource utilization efficiency and reduce resource waste; S4. Based on the collaborative model, count the task completion status of the intelligent agent when performing tasks, build a task completion model, and design an optimization algorithm to optimize the task completion rate and improve task execution efficiency; S5. Combine the path planning model, resource consumption model, and task completion model to construct an optimization model. This model is then solved using a multi-objective particle swarm optimization algorithm to ensure efficient collaboration among the unmanned swarm during task execution. S6. Based on the optimization model, model training and evaluation are carried out by preparing task data, agent data, resource consumption data, and task completion data to achieve autonomous evolution and autonomous collaborative task execution of unmanned swarm intelligence.
[0010] Preferably, the process of constructing the relationship model between agents in step S1 includes: S11. The intelligent agent uses a drone. For n intelligent agents, construct the attribute vector:
[0011] in, Indicates the The comprehensive capabilities of an intelligent agent, Indicates the The maximum flight speed of an agent, Indicates the The maximum load weight of an agent, No. The largest energy reserve, Representative Sensor efficiency of each agent; S12. For m types of tasks, define the task feature matrix:
[0012] in, Indicates a task The urgency of Indicates a task The accuracy requirements, Task Energy consumption, ; S13. Construct a collaborative relationship between agents. The collaborative effectiveness tensor is expressed as:
[0013] in, Indicates the Agent and Execute tasks between agents The probability of adaptation, Indicates a task The collaboration weight matrix, Represents the Sigmoid activation function; S14. Construct the task adaptation function between the agent and the task, which is:
[0014] in, Indicates the Agent and The degree of suitability for a task, Indicates the The task feature vector matrix, represents the attenuation coefficient; S15. Construct the relationship between the completed task and the number of agents. The group collaborative efficiency function is:
[0015] in, Represents a subset of agents and tasks performed The synergistic relationship represents the collaboration gain coefficient, represents the number of agents in the subset, represents the scale penalty coefficient. Preferably, the coordination model in step S1 is constructed by executing a dynamic programming value function, and the calculation formula of the dynamic programming value function is:
[0016] in, Indicates At this moment, the status is The maximum expected return under represents the path cost, Represents resource cost, and Represent the weights of path cost and resource cost respectively, Indicates the expected value of future states; There are constraints:
[0017] in, Indicates the minimum threshold of synergistic effectiveness, Indicates the reserve amount of key resources.
[0018] Preferably, the process of establishing the path planning model in step S2 includes: S21. Parameterize the path trajectory of the agent, and its formula is expressed as:
[0019] in, Representing an agent At the moment The three-dimensional position of , , Indicates The three-dimensional coordinates of the agent at each moment, represents the orbital parameter vector, No. basis functions; S22. Establish a path cost function, which is expressed as:
[0020] in, , , Represent path smoothing weight, resource distance weight, and resource sensitivity weight respectively. , Represents the intelligent agent , location.
[0021] Preferably, the mathematical expression of the resource consumption model in step S3 is:
[0022]
[0023]
[0024] in, Indicates that the task is completed Energy consumption in , represents the correlation coefficient of the UAV, Indicates the speed of the drone, is the initial velocity, is the acceleration of the drone, is the air density, is the zero-lift drag coefficient, is the wing area, is the Oswald efficiency factor, Wing aspect ratio, It's the quality of the drone. is the acceleration due to gravity, Represents the total time taken by the drone mission.
[0025] There are constraints:
[0026] in, and Represent the initial position and final position of the UAV respectively, and Indicates the maximum speed and maximum acceleration, , Indicates the initial speed and the final speed.
[0027] Preferably, the calculation formula for constructing the task completion model T in step S4 is:
[0028]
[0029] in, represents the subtask set of the k-th task, represents the i-th subtask of the k-th task, Assign a collection of drones to perform subtasks, is the space-time constraint coefficient, Remaining time of the task, Indicates the maximum speed, is the initial energy, Represents the energy consumption of the task.
[0030] Preferably, the construction process of the optimization model in step S5 includes: S51. The state variables of the integrated collaboration, path, resource and task completion models are:
[0031] There is a multi-objective value function as an optimization model :
[0032]
[0033]
[0034] in; For task completion, is the learning rate, is the activation function, is the task change rate, For resource consumption, is the resource consumption benchmark value, Hyperbolic tangent function, is the position variance, is the path smoothness, For minimum collaborative efficiency, For safety, is the maximum energy distance, Urgency factor, Key indicators, The weight coefficient of task completion, represents the weight coefficient of energy consumption, The weight coefficient representing the smoothness of the path.
[0035] S25. Solve using improved multi-objective particle swarm optimization:
[0036]
[0037]
[0038]
[0039]
[0040] in; Represents particles No. Weidi The update speed after iterations, Represents particles No. Dimensional speed The update speed after iterations, Represents particles No. Vidi The updated position after iterations, Represents particles No. Vidi The updated position after iterations, represents the dynamic inertia weight, , represents the adaptive learning factor, , Represent the minimum and maximum values of cognitive learning factors, respectively. , represent the minimum and maximum values of the social learning factor, respectively. , is a random number, represents the optimal position of an individual, represents the global optimal position, , Indicates the upper and lower limits of the inertia weight, Indicates the current iteration number, Indicates the maximum number of iterations.
[0041] Preferably, the specific operation process of the model training in step S6 includes: S61. Construct multimodal fusion formula:
[0042] in Indicates the large model of the base, is the cross-model attention mechanism, Encoding historical decisions; S62. There is an attention mechanism:
[0043] S63. Construction strategy - value double-head output:
[0044] in, , is the policy header parameter, For the value header parameter, is the discount factor, Expresses expectations about future states; S64. Calculate the loss function L
[0045] in, It's expectation. Forecast returns for the model, To optimize comprehensive benefits, is the KL divergence, refers to the probability distribution of the current strategy, is the probability distribution of the expert strategy, is a random variable The variance of is the distribution coefficient, is the hidden coefficient; S65. Implementing Gradient Collaborative Updates:
[0046] in, is the gradient projection, is the collaborative learning rate, Indicates time The synergy matrix.
[0047] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0048] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0049] As can be seen from the above technical solutions, the present invention provides a large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems. Compared with the existing technology, the present invention has the following advantages: 1. By analyzing the collaborative mechanisms under different tasks, the present invention can establish a basic multi-agent collaboration framework, facilitate further design of dynamic adjustment mechanisms based on task changes and agent states, establish a path planning model that can be updated in real time, and significantly improve the path planning flexibility and adaptability of unmanned swarms in complex environments.
[0050] 2. The present invention constructs a resource consumption model and a task completion model, quantifies resource utilization efficiency and optimizes task execution effects based on experimental / simulation data, and can form a complete optimization system by combining comprehensive path planning, resource consumption and task completion models. It also uses an improved multi-objective particle swarm algorithm to achieve multi-dimensional collaborative optimization, while ensuring efficient collaboration and taking into account resource utilization and task completion rate, thereby comprehensively improving the execution efficiency of unmanned cluster tasks.
[0051] 3. This invention uses multimodal fusion large-scale model training, combines task data with agent data, and innovatively designs a strategy-value dual-head output architecture to promote the autonomous evolution of unmanned cluster collaborative intelligence, realize intelligent and autonomous collaborative task execution, and provide innovative solutions with environmental adaptability and resource optimization capabilities for complex scenarios.
[0052] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become easy to understand through the following description. Of course, it is not necessary to achieve all of the above-mentioned advantages simultaneously in order to implement any product of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings: Figure 1 It is a schematic diagram of the overall operation flow of the present invention. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0055] In the embodiment, see Figure 1 .
[0056] like Figure 1 The large-scale autonomous learning method for generating a swarm intelligence strategy for an unmanned system proposed in an embodiment of the present invention includes the following steps: Step 1: Classify tasks into m categories, such as emergency tasks, strike tasks, and detection tasks. Build a relationship model between any n agents and analyze their coordination when performing different tasks. For each task, study the performance relationship between multiple agents and tasks, including resource consumption rates and task completion rates. Finally, construct a coordination model to describe the performance, resource consumption, and path relationships of any n agents when they collaborate to complete a task.
[0057] Construct the drone attribute vector:
[0058] in, Indicates the The comprehensive capabilities of a drone, Indicates the The maximum flight speed of a drone, Indicates the Maximum payload weight of a drone, No. Maximum energy reserve; For m types of tasks, define the task feature matrix:
[0059] in, Indicates a task The urgency of Indicates a task The accuracy requirements, Task energy consumption; Constructing a collaborative relationship between drones, the collaborative efficiency tensor is expressed as:
[0060] in , indicating the drone and the The adaptation probability between the UAVs, Indicates a task The collaboration weight matrix, Represents the Sigmoid activation function; Construct the task fitness function between the drone and the task:
[0061] in, Indicates the drone and the The degree of suitability for a task, Indicates the The task feature vector matrix, represents the attenuation coefficient; Constructing the relationship between the completed tasks and the number of drones, the group collaborative efficiency function is:
[0062] in, Represents a subset of drones and tasks performed The synergistic relationship represents the collaboration gain coefficient, represents the number of drones in the subset, represents the scale penalty coefficient; Dynamic Programming Value Function
[0063] in, Indicates At this moment, the status is The maximum expected return under represents the path cost, Represents resource cost, Represents the expected value of the future state, which is the expected benefit after considering the uncertainty of state transition; Constraint formula:
[0064] in, Indicates the minimum threshold of synergistic effectiveness, Indicates the reserve amount of key resources; By dividing tasks into m categories at the initial stage of task planning and building a relationship model between any n agents, we can deeply analyze the coordination mechanism of agents under different tasks, thereby clarifying the basic framework of multi-agent collaboration and finally providing support for the generation of collaborative strategies in complex task scenarios. Step 2: Based on the collaborative model, establish a path planning model for the agent when performing the task. By selecting or designing an appropriate path planning algorithm, and combining task changes and agent status, design a dynamic adjustment mechanism to update the path planning in real time to adapt to complex task environments; The UAV path trajectory is parameterized and its formula is expressed as:
[0065] in, Indicates drone At the moment The three-dimensional position of , , Indicates The three-dimensional coordinates of the drone at any moment, represents the orbital parameter vector, No. basis functions; Establish a path cost function, whose expression is:
[0066] in, , , Represent path smoothing weight, resource distance weight, resource sensitivity weight, , Indicates drone , location; By building a path planning model based on the collaborative model and designing a dynamic adjustment mechanism that combines task changes with agent status, the path planning can be updated in real time to adapt to complex and changing task environments, thereby improving the flexibility and environmental adaptability of the unmanned swarm path planning. Step 3: Based on the collaborative model, analyze the resource consumption of the agent when performing the task. Collect resource consumption data through experiments or simulations, use statistical methods or machine learning algorithms to build a resource consumption model, and design an optimization algorithm to optimize resource utilization efficiency and reduce resource waste. Energy consumption model:
[0067]
[0068]
[0069] in, represents the energy consumption in completing the task, , represents the correlation coefficient of the UAV, Indicates the speed of the drone, is the initial velocity, is the acceleration of the drone, is the air density, is the zero-lift drag coefficient, is the wing area, is the Oswald efficiency factor, Wing aspect ratio, is the weight of the drone, is the acceleration due to gravity, Represents the total time of the UAV mission;
[0070] in, and Represent the initial position and final position of the UAV respectively, and Indicates the maximum speed and maximum acceleration, , Indicates the initial speed and the final speed; By analyzing resource consumption based on collaborative models, building resource consumption models using experimental or simulation data, and designing optimization algorithms, we can quantify resource consumption and optimize resource utilization efficiency, thereby reducing resource waste and improving resource utilization efficiency in unmanned swarms. Step 4: Based on the collaborative model, analyze the agent's task completion status during task execution. Collect task completion data through experiments or simulations, use statistical methods or machine learning algorithms to build a task completion model, and design an optimization algorithm to optimize the task completion rate and improve task execution efficiency.
[0071]
[0072] in, represents the subtask set of the k-th task, represents the i-th subtask of the k-th task, Assign a collection of drones to perform subtasks, is the space-time constraint coefficient, Remaining time of the task, Indicates the maximum speed, is the initial energy, represents the energy consumption of the task; Step 5: Integrate the path model, resource consumption model, and task completion model to form a complete optimization model. This model is solved using an improved multi-objective particle swarm optimization algorithm to ensure that the unmanned swarm can collaborate efficiently while performing tasks, while also balancing resource consumption and task completion rates.
[0073] Integrate state variables of collaboration, path, resource and task completion models:
[0074] Multi-objective value function:
[0075]
[0076]
[0077] in; For task completion, is the learning rate, is the activation function, is the task change rate, For resource consumption, is the resource consumption benchmark value, Hyperbolic tangent function, is the position variance, is the path smoothness, For minimum collaborative efficiency, For safety, is the maximum energy distance, Urgency factor, Key indicators, The weight coefficient of task completion, represents the weight coefficient of energy consumption, The weight coefficient representing the smoothness of the path.
[0078] Solve using improved multi-objective particle swarm optimization:
[0079]
[0080]
[0081]
[0082]
[0083] in; Represents particles No. Weidi The update speed after iterations, Represents particles No. Dimensional speed The update speed after iterations, Represents particles No. Vidi The updated position after iterations, Represents particles No. Vidi The updated position after iterations, represents the dynamic inertia weight, , represents the adaptive learning factor, , Represent the minimum and maximum values of cognitive learning factors, respectively. , represent the minimum and maximum values of the social learning factor, respectively. , is a random number, represents the optimal position of an individual, represents the global optimal position, , Indicates the upper and lower limits of the inertia weight, Indicates the current iteration number, Indicates the maximum number of iterations.
[0084] At this time, a complete optimization model is formed by integrating the path model, resource consumption model and task completion model, and the improved multi-objective particle swarm optimization algorithm is used to solve it. This can achieve multi-objective collaborative optimization, thereby balancing resource consumption and task completion rate while achieving efficient collaboration, and improving the overall task execution efficiency of the unmanned swarm. Step 6: Based on the above models and algorithms, train a large model to achieve autonomous evolution of collaborative swarm intelligence in unmanned swarms. By preparing task data, agent data, resource consumption data, and task completion data, conduct model training and evaluation, and ultimately achieve intelligent and autonomous collaborative task execution in unmanned swarms. Perform multimodal fusion on input data:
[0085] in Indicates the large model of the base, is the cross-model attention mechanism, Encoding historical decisions;
[0086] Construction strategy - value dual-head output:
[0087] in, , is the policy header parameter, For the value header parameter, is the discount factor, Expresses expectations about future states;
[0088] in, It's expectation. Forecast returns for the model, To optimize comprehensive benefits, is the KL divergence, refers to the probability distribution of the current strategy, is the probability distribution of the expert strategy, is a random variable The variance of is the distribution coefficient, is the hidden coefficient; Gradient collaborative update:
[0089] in; is the gradient projection, is the collaborative learning rate, Indicates time Synergy effectiveness matrix; In summary, by using task data, agent data, and other training large models, and conducting multimodal fusion and strategy-value dual-output design, it is possible to achieve the autonomous evolution of unmanned cluster collaborative intelligence, thereby promoting intelligent and autonomous collaborative task execution, and finally providing innovative solutions for the application of unmanned systems in complex scenarios.
[0090] Based on the collaborative model, a path planning model, a resource consumption model, and a task completion model are established respectively, which are then combined into an optimization model and solved by a multi-objective particle swarm optimization algorithm. Finally, a large-scale multimodal fusion model is trained and evaluated based on the optimization model, which significantly improves the path planning flexibility and adaptability of unmanned swarms in complex environments, optimizes resource utilization efficiency, improves task completion rate, promotes the autonomous evolution and autonomous task execution of collaborative group intelligence of unmanned swarms, and solves the problems of low efficiency, slow response, and poor adaptability of traditional collaborative planning methods in scenarios such as logistics, emergency rescue, and industrial manufacturing.
[0091] In a specific embodiment, the present application has the following evaluation table description for evaluating the effect:
[0092] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0093] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0094] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the large-model autonomous learning methods for generating swarm intelligence strategies for unmanned systems in the above-mentioned embodiments.
[0095] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the above method.
[0096] The embodiment of the present application further provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. Memory for storing computer programs; The processor is used to implement the above-mentioned large-model autonomous learning method for generating swarm intelligence strategies for unmanned systems when executing the program stored in the memory.
[0097] The communication bus mentioned in the above electronic device can be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0098] The communication interface is used for communication between the above electronic device and other devices.
[0099] The memory may include a random access memory, or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0100] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor, etc.; it can also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component.
[0101] It should also be noted that electronic devices also include terminal devices, which can also be called terminals, user equipment, mobile stations, mobile terminals, etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablet computers, computers with wireless transceiver functions, virtual reality terminal devices, augmented reality terminal devices, wireless terminals in industrial control, wireless terminals in unmanned driving, wireless terminals in remote surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technology and specific device form used by the terminal devices.
[0102] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0104] In addition, it should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components in a certain specific posture. If the specific posture changes, the directional indications will also change accordingly.
[0105] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or schemes in which A and B are satisfied at the same time. In addition, in the embodiments of the present invention, "multiple" refers to more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in this field to implement. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
Claims
1. A large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems, characterized by: include: S1. Classify tasks into multiple categories and build a relationship model between any number of agents. Build a coordination model based on the coordination relationship between agents when performing different tasks, as well as the effectiveness relationship between agents and different tasks. S2. Based on the collaborative model, establish a path planning model for the agent when performing tasks; S3. Based on the collaborative model, count the resource consumption of the agents when performing tasks and build a resource consumption model; S4. Based on the collaborative model, count the task completion status of the intelligent agent when performing the task and build a task completion model; S5. Combine the path planning model, resource consumption model, and task completion model to construct an optimization model, and solve the model using a multi-objective particle swarm optimization algorithm; S6. Based on optimized model training and evaluation, realize the autonomous evolution and autonomous collaborative task execution of unmanned swarm intelligence.
2. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 1, characterized in that: The process of constructing the relationship model between agents in step S1 includes: S11. The intelligent agent uses a drone. For n intelligent agents, construct the attribute vector: Among them, A i represents the comprehensive ability of the i-th agent, represents the maximum flight speed of the i-th agent, represents the maximum load weight of the i-th agent, The i-th maximum energy reserve, represents the sensor efficiency of the i-th agent; S12. For m types of tasks, define the task feature matrix: in, represents the urgency of task j, represents the accuracy requirement of task j, Energy consumption of task j, j∈(1,2,…,m); S13. Construct a collaborative relationship between agents. The collaborative effectiveness tensor is expressed as: Among them, E collab [i,j,k] represents the adaptation probability of the i-th agent and the j-th agent to perform task k, W k represents the collaborative weight matrix of task k, σ() represents the Sigmoid activation function; S14. Construct the task adaptation function between the agent and the task, which is: Among them, M(u i ,t k ) represents the degree of adaptation between the i-th agent and the k-th task, represents the kth task feature vector matrix, θ represents the attenuation coefficient; S15. Construct the relationship between the completed task and the number of agents. The group collaborative efficiency function is: Among them, Q(S,t k ) represents the agent subset S and the task t k The collaborative effectiveness relationship is, λ represents the collaborative gain coefficient, |S| represents the number of agents in the subset, and μ represents the scale penalty coefficient.
3. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 2, characterized in that: In step S1, the coordination model is constructed by executing a dynamic programming value function. The calculation formula of the dynamic programming value function is: Among them, V(s t ) means that at time t, the state is s t The maximum expected return under J p represents the path cost, J r represents resource cost, γ1 and γ2 represent the weights of path cost and resource cost respectively, E[V(s t+1 )] represents the expected value of the future state.
4. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 2, characterized in that: The process of establishing the path planning model in step S2 includes: The path trajectory of the agent is parameterized and the path cost function is established, which is expressed as follows: Among them, ω1, ω2, ω3 represent path smoothing weight, resource distance weight, and resource sensitivity weight respectively, and p i , p j Represents the agent u i ,u j location.
5. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 4, characterized in that: The mathematical expression of the resource consumption model in step S3 is: Where E(p(t)) represents the energy consumption in completing the task p(t), c1 and c2 represent the correlation coefficients of the UAV, v(t) represents the speed of the UAV, v(0) is the initial speed, a(t) is the acceleration of the UAV, and ρ is the air density. is the zero-lift drag coefficient, S is the wing area, e0 is the Oswald efficiency factor, A R Wing aspect ratio, m is the mass of the drone, g is the acceleration due to gravity, and N represents the total duration of the drone mission.
6. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 4, characterized in that: The calculation formula for constructing the task completion model U in step S4 is: among themU k ={τ k1 ,t k2 ,...,t kn } Among them, U k represents the subtask set of the k-th task, τ ki represents the i-th subtask of the k-th task, S ki The set of drones assigned to perform subtasks, λ is the spatiotemporal constraint coefficient, t remain Remaining time of the task, v jmax Indicates the maximum speed, is the initial energy, Represents the energy consumption of the task.
7. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 4, characterized in that: The construction process of the optimization model in step S5 includes: S51. Integrate the state variables of the coordination, path, resource and task completion models, and there is a multi-objective value function as the optimization model J(s t ): Among them; Q k is the task completion degree, η α is the learning rate, σ() is the activation function, is the task change rate, For resource consumption, is the resource consumption benchmark value, tanh() is the hyperbolic tangent function, Var() is the position variance, is the path smoothness, E min is the minimum collaborative effectiveness, d safe For safety, is the maximum energy distance, Q emergency Urgency index, I critical (k) Key indicator, α k (t) represents the weight coefficient of task completion, β i (t) represents the weight coefficient of energy consumption, γ i (t) represents the weight coefficient of path smoothness. S25. Solve using improved multi-objective particle swarm optimization: in; represents the updated velocity of particle i’s j-th dimension after the t+1th iteration, represents the updated velocity of particle i after the tth iteration of the j-th dimension velocity, represents the updated position of particle i after the j-th dimension t+1 iteration, represents the updated position of particle i after the tth iteration of the jth dimension, ω (t) represents the dynamic inertia weight, represents the adaptive learning factor, c 1,min , c 1,max Represent the minimum and maximum values of cognitive learning factors, c 2,min , c 2,max Represent the minimum and maximum values of the social learning factor, r1 and r2 are random numbers, and p ij represents the optimal position of an individual, g j represents the global optimal position, ω max ,ω min Indicates the upper and lower limits of the inertia weight, q indicates the current number of iterations, Q max Indicates the maximum number of iterations.
8. The large-scale autonomous learning method for generating swarm intelligence strategies for unmanned systems according to claim 1, characterized in that: The specific operation process of the model training in step S6 includes: S61. Construct multimodal fusion formula: Among them, GLM() represents the base model, CA() is the cross-model attention mechanism, Encoding historical decisions; S62. There is an attention mechanism: S63. Construction strategy - value double-head output: Among them, W π , b π is the strategy header parameter, W v is the value head parameter, γ is the discount factor, E[J(S τ )] expresses expectations about future states; S64. Calculate the loss function L: L=E[J pred -J opt ] 2 +β·KL(φ||φ expert )+ρ·Var(H) Among them, E[] is the expectation, J pred For the model to predict returns, J opt To comprehensively optimize the benefits, KL() refers to the KL divergence, φ refers to the probability distribution of the current strategy, and φ expert It refers to the probability distribution of the expert strategy, Var(H) is the variance of the random variable H, β is the distribution coefficient, and ρ is the hidden coefficient; S65. Implementing Gradient Collaborative Updates: Among them, Proj() is the gradient projection, μ is the collaborative learning rate, E t Represents the synergy effectiveness matrix at time t.
Citation Information
Patent Citations
Unmanned aerial vehicle cluster task allocation method based on crowd intelligence excitation coalition game
CN115963724A
Quantitative fault-tolerant adaptive bilateral tracking control method for unmanned aerial vehicle
CN116184818A
Unmanned aerial vehicle cluster intelligent confrontation strategy generation method based on group collaboration
CN116859989A
Digital-analog hybrid unmanned cluster brain-like crowd-sourcing collaborative navigation method
CN118643858A
Energy control method and system of wireless sensor network
CN119277421A
Cited By
Unmanned aerial vehicle cluster energy efficiency optimization control method and system based on input-output analysis
CN121477940A
Unmanned aerial vehicle cluster energy efficiency optimization control method and system based on input-output analysis
CN121477940B