Multi-unmanned-system decision-making method and device based on adaptive cooperation and dynamic optimization, and medium
By building an adaptive collaboration network and dynamic optimization model, the collaboration and decision-making problems of unmanned systems in complex environments are solved, information sharing and resource optimization are realized, and the decision-making efficiency and cooperation effect of the system are improved.
Patent Information
- Application Number
- CN202510475542.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
Existing unmanned systems collaborative decision-making methods are inflexible and decision-making efficiency when dealing with complex dynamic environments, and cannot meet the needs of fast response and efficient collaboration.
Build a multi-unmanned system decision-making method based on adaptive collaboration and dynamic optimization. By initializing the dynamic collaboration network, weight accumulation and trust adjustment are carried out, and information sharing and resource optimization are achieved by combining Bayesian inference and game theory models.
Achieve efficient collaboration and information sharing of multiple unmanned systems in complex environments, ensure optimal decision-making, and improve the system's cooperation effect under limited resources.
Smart Images

Figure CN120406492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned system cooperation, and more specifically, to a decision-making method, device, and medium for multi-unmanned systems based on adaptive cooperation and dynamic optimization. Background Art
[0002] With the wide application of unmanned systems in fields such as military, disaster relief, and monitoring, these systems face uncertainties and resource constraints in complex dynamic environments when performing tasks. Traditional decision-making models usually rely on static rules or simple logical judgments, making them inadequate in dealing with highly dynamic and complex environmental changes. Especially in the scenario of multi-unmanned system cooperation, how to achieve adaptive cooperation and information sharing to improve the overall efficiency has become an urgent problem to be solved. Existing unmanned system cooperation decision-making methods have obvious deficiencies in flexibility and decision-making efficiency, and cannot meet the requirements of rapid response and efficient cooperation in modern applications.
[0003] Therefore, how to provide a decision-making method for multi-unmanned systems based on adaptive cooperation and dynamic optimization is an urgent problem for those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention provides a decision-making method, device, and medium for multi-unmanned systems based on adaptive cooperation and dynamic optimization, which realizes efficient cooperation and information sharing among multi-unmanned systems and ensures optimal decision-making in complex dynamic environments.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A decision-making method for multi-unmanned systems based on adaptive cooperation and dynamic optimization, comprising:
[0007] Initializing a dynamic cooperation network based on unmanned systems;
[0008] Adjusting the weight accumulation and weight trust of the dynamic cooperation network;
[0009] Calculating and weighted aggregating and updating the posterior probability based on Bayesian inference according to the updated weight trust value;
[0010] Establishing a dynamic cooperation game model based on the updated weighted aggregation information to achieve a globally optimal cooperation result.
[0011] Preferably, initializing a dynamic cooperation network based on unmanned systems includes:
[0012] Let the set of unmanned systems be V = {v1, v2,..., v n}, and define the state X of each unmanned system node v i i (t), the state of the node includes task progress, resource usage, and enemy activities, and n represents the number of unmanned systems;
[0013] Construct a graph G=(V, E,)), where E={e ij} is the edge set, and W is the weight set;
[0014] Set an initial weight w ij (t) for each edge.
[0015] Preferably, the weight accumulation adjustment formula is:
[0016]
[0017] where f(X i , X j ) is the dynamic connection factor based on the nodes, δ(e ij ) is the indicator function of the cooperation relationship, α and β are adjustment coefficients, Δ is an observation window time, and w ij (t + 1) is the weight at time t + 1 after time integration update.
[0018] Preferably, the weight trust adjustment includes:
[0019] For each unmanned system node v i , calculate the trust value T j of its neighbor node v ij at time t:
[0020] T ij (t + 1)=γT ij (t)+(1 - γ)R ij (t),
[0021] where T ij (t)∈[0, 1]. When T ij (t)=0, it means that the unmanned system node v i completely distrusts the neighbor node v j of the unmanned system at time t; when T ij (t)=1, it means that the unmanned system node v i completely trusts the neighbor node v j of the unmanned system at time t. T ij (t + 1) is the trust value of the neighbor node v j of the unmanned system at time t + 1, γ is the smoothing factor, and R ij (t) is the trust value at the current moment in the memoryless state;
[0022] Perform weight trust adjustment based on the trust value:
[0023]
[0024] Among them, is the updated weight trust value.
[0025] Preferably, based on the updated weight trust value, posterior probability calculation and weighted aggregation update based on Bayesian inference are performed, including:
[0026] Obtain the current observation information, and calculate the posterior probability of the unmanned system node based on the Bayesian update formula:
[0027]
[0028] Among them, O i is the current observation information, P(X i ) is the prior probability, P(O i |X i ) represents the system's perception ability of the environment, and P(X i |O i ) represents the accuracy of the node state estimation under the current observation;
[0029] Perform weighted aggregation update based on the posterior probability and the updated weight trust value:
[0030]
[0031] Among them, X i represents the state of the unmanned system node v i . Aggregate the states of all i ∈ N(j) to obtain the updated state of the unmanned system node v j , represents the updated weighted aggregation information of the unmanned system node v j , and N(j) is the neighborhood of the unmanned system node v j .
[0032] Preferably, based on the updated weighted aggregation information, a dynamic cooperative game model is established, including:
[0033] Calculate the strategy of the unmanned system node v j according to the updated weighted aggregation information ω is a model parameter, and s j (ω) is the action set of the unmanned system node v j in response to the observation information;
[0034] Assume that the inter-system utility function is u j (s1, s2,..., s N ), where s1, s2,..., s N is the strategy set of other systems;
[0035] The collaborative utility and competitive utility between systems are represented by the following function:
[0036] u j (s1, s2, …, s N ) = C j (s1, s2, …, s N ) - P j (s1, s2, …, s N )
[0037] where C j (s1, s2, …, s N ) is the benefit brought by collaboration, and P j (s1, s2, …, s N ) is the penalty due to competition;
[0038] Update the strategy rules according to the collaborative utility and competitive utility functions between systems:
[0039]
[0040] where is the strategy that maximizes the utility function.
[0041] Preferably, it further includes:
[0042] Define a loss function according to the strategy s j of the unmanned system node v j and the strategy that maximizes the utility function Perform backpropagation on the model parameter ω where η is the learning rate, ω: is the updated LSTM parameter, represents the gradient operator.
[0043] Preferably, the dynamic connection factor is determined based on resource requirements and task importance.
[0044] A computer device, comprising: a memory and a processor, where a computer program that can run on the processor is stored in the memory, and when the processor executes the computer program, a multi-unmanned system decision-making method based on adaptive collaboration and dynamic optimization is implemented.
[0045] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, a multi-unmanned system decision-making method based on adaptive collaboration and dynamic optimization is implemented.
[0046] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a decision-making method, device, and medium for a multi-unmanned system based on adaptive collaboration and dynamic optimization, having the following effects:
[0047] (1) Bayesian inference network model under complex structures: This model constructs a dynamic collaboration network for a multi-unmanned system, where nodes exchange information and update it in real time. The states of the nodes include information such as task progress, resource usage, and enemy activities, and the information dissemination mechanism is based on the Bayesian update formula. The network topology is dynamically optimized according to task requirements and environmental changes to ensure smooth information transmission and achieve optimal collaboration.
[0048] (2) Network structure: The connections between multi-unmanned systems are constructed through a complex network model, and the weight w ij between nodes is dynamically adjusted to optimize information sharing and resource allocation.
[0049] (3) Information dissemination and aggregation: Information dissemination is updated through Bayesian inference, and information aggregation uses the weighted average method. The weight values are regulated by a trust mechanism to ensure the reliability of information sources.
[0050] (4) Dynamic collaboration game model: Each system collaborates and competes in task allocation and resource sharing according to the utility function in game theory. By adjusting strategies, the system gradually optimizes the cooperation effect to ensure a globally optimal decision.
[0051] (5) Strategy update: The system adjusts its own strategy by solving the maximum value of the utility function. Cooperation rewards and competition penalty terms are introduced into the utility function to ensure that the system reasonably balances cooperation and competition under limited resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.
[0053] Figure 1 It is a flowchart of a decision-making method for a multi-unmanned system based on adaptive collaboration and dynamic optimization provided by the present invention.
[0054] Figure 2 It is a schematic diagram of network weight update provided by the present invention.
[0055] Figure 3 It is a schematic diagram of parameter update of the dynamic collaboration game model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] The embodiment of the present invention discloses a decision-making method for a multi-unmanned system based on adaptive cooperation and dynamic optimization, as Figure 1 shown, including:
[0058] Initializing a dynamic cooperation network based on the unmanned system;
[0059] Performing weight accumulation adjustment and weight trust adjustment on the dynamic cooperation network;
[0060] Performing posterior probability calculation and weighted aggregation update based on Bayesian inference according to the updated weight trust value;
[0061] Establishing a dynamic cooperation game model based on the updated weighted aggregation information to achieve a globally optimal cooperation result.
[0062] Next, a specific description will be given for each step of the present invention.
[0063] (1) Bayesian inference network model under complex structure
[0064] By constructing a complex dynamic cooperation network, multiple unmanned systems are connected into a dynamic cooperation system. Each unmanned system serves as a node in the network, and the nodes are connected by weighted connections to form the network. The Bayesian inference method is used for information propagation to update the environmental state, task progress, enemy behavior, etc. in real time.
[0065] Specifically, (11) initializing a dynamic cooperation network based on the unmanned system includes:
[0066] Let the set of unmanned systems be V = {v1, v2,..., v n}, and define the state X i (t) of each unmanned system node v i , which includes information such as task progress, resource usage, and enemy activities, and n represents the number of unmanned systems.
[0067] The connection topology structure between unmanned system nodes adopts a dynamic complex network, denoted as G = (V, E, W), where E = {e ij} is the edge set, which is a qualitative description of the cooperation relationship between nodes, and W is the weight set.
[0068] Set an initial weight w for each edge ij (t), and the weight w ij (t) of the edge is dynamically adjusted according to factors such as task requirements, resource consumption, and enemy activities.
[0069] (12) Perform weight accumulation adjustment and weight trust adjustment on the dynamic collaboration network, as Figure 2 shown;
[0070] Design a dynamic evolution mechanism for the network topology. Let the unmanned system node v i and its neighbor node v j The connection strength w ij (t) changes over time, and the weight accumulation adjustment is performed through the following formula:
[0071]
[0072] where f(X i , X j ) is a dynamic connection factor calculated based on factors such as resource requirements and task importance between nodes, δ(e ij ) is an indicator function of the collaboration relationship, and α and β are adjustment coefficients representing factors such as task requirements and resource matching degree. Δ is an observation window time, and generally Δ < 1.
[0073] To improve information reliability, each unmanned system node v i needs to evaluate the trust value T j of its neighbor node v ij (t) at time t.
[0074] T ij (t + 1) = γT ij (t) + (1 - γ)R ij (t),
[0075] where T ij (t) ∈ [0, 1]. When T ij (t) = 0, it means that the unmanned system node v i completely distrusts the neighbor node v j of the unmanned system at time t; when T ij (t) = 1, it means that the unmanned system node v[[ID=E]] i completely trusts the neighbor node v j of the unmanned system at time t. T ij (t + 1) is similar. γ is a smoothing factor; R ij (t) is calculated based on historical collaboration results, and its meaning is the trust value at the current moment in the memoryless state:
[0076]
[0077] Weight trust adjustment based on trust value
[0078]
[0079] Among them, is the updated weight trust value.
[0080] (13) Calculate and update the posterior probability based on Bayesian inference according to the updated weight trust value, including:
[0081] Obtain the current observation information and calculate the posterior probability through the Bayesian update formula:
[0082]
[0083] Among them, O i is the current observation information; P(X i ) is the prior probability; P(O i |X i ) represents the system's perception ability of the environment, and P(X i |O i ) represents the accuracy of the node state estimation under the current observation.
[0084] The information of multiple nodes is weighted and aggregated and updated based on the posterior probability and the updated weight trust value:
[0085]
[0086] Among them, X i represents the state of the unmanned system node v i . Aggregate the states of all i ∈ N(j) to obtain the updated state of the unmanned system node v j , represents the updated weighted aggregation information of the unmanned system node v j . N(j) is the neighborhood of the unmanned system node v j . Through the probability P(X i |O i ) and the updated weight to fuse the latest information such as task progress, resource usage, and enemy activities.
[0087] (2) Dynamic cooperative game model:
[0088] In the process of task allocation and resource sharing, the system coordinates the cooperation and competition of different unmanned systems through a game theory model. Each unmanned system adjusts its strategy according to its own state and environmental changes to optimize the global cooperation benefit. This model ensures that the unmanned systems can reasonably allocate tasks and share resources under limited resources, and finally achieve the globally optimal cooperation result.
[0089] Specifically, a dynamic collaborative game model is established based on the updated weighted aggregation information, as Figure 3 shown, including:
[0090] The strategy of each unmanned system node v j is ω is the original parameter of the model. s j (ω) is the action set made by the current unmanned system node v j for the observed value, including adjusting the zoom, following the target, displacement control, etc.
[0091] Assume its utility function is u j (s1, s2, …, s N ), where s1, s2, …, s N is the strategy set of other systems.
[0092] The collaborative utility and competitive utility between systems are represented by the following function:
[0093] u j (s1, s2, …, s N ) = C j (s1, s2, …, s N ) - P j (s1, s2, …, s N )
[0094] where C j (s1, s2, …, s N ) is the benefit brought by collaboration, and P j (s1, s2, …, s N ) is the penalty due to competition. The system updates its strategy according to the above utility function in each round of decision-making.
[0095] The strategy update rule of the system is:
[0096]
[0097] where, is the strategy that maximizes the utility function.
[0098] Among them, the model training process is:
[0099] According to v j and s j define the loss function perform backpropagation on the model parameter ω where η is the learning rate, ω: is the updated LSTM parameter, represents the gradient operator.
[0100] This embodiment provides a computer device, including: a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, a multi-unmanned system decision-making method based on adaptive cooperation and dynamic optimization is implemented.
[0101] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a multi-unmanned system decision-making method based on adaptive cooperation and dynamic optimization is implemented.
[0102] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media that can store program codes, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0103] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0104] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A decision-making method for multi-unmanned systems based on adaptive collaboration and dynamic optimization, characterized in that Including: Initializing a dynamic collaboration network based on an unmanned system; Performing weight accumulation adjustment and weight trust adjustment on the dynamic collaboration network; Calculating the posterior probability based on Bayesian inference and performing weighted aggregation update according to the updated weight trust value; Establishing a dynamic collaboration game model based on the updated weighted aggregation information to achieve a globally optimal collaboration result.
2. The decision-making method for a multi-unmanned system based on adaptive cooperation and dynamic optimization according to claim 1, characterized in that Initializing a dynamic collaboration network based on an unmanned system, including: Let the set of unmanned systems be \(V = \{v_1, v_2, \ldots, v\) n \}\), and define the state \(X\) i \((t)\) of each unmanned system node \(v\) i . The state of the node includes task progress, resource usage, and enemy activities. \(n\) represents the number of unmanned systems; Construct a graph \(G=(V, E, W)\), where \(E = \{e ij \}\) is the edge set and \(W\) is the weight set; Set an initial weight w for each edge ij (t).
3. A decision-making method for a multi-unmanned system based on adaptive collaboration and dynamic optimization according to claim 1, characterized in that, The weight accumulation adjustment formula is: Among them, f(X i , X j ) is based on the dynamic connection factor between nodes, δ(e ij ) is the indicator function of the cooperation relationship, α and β are adjustment coefficients, Δ is an observation window time, and w ij (t + 1) is the weight at time t + 1 after time integration update.
4. A decision-making method for a multi-unmanned system based on adaptive collaboration and dynamic optimization according to claim 3, characterized in that, The weight trust adjustment includes: For each unmanned system node v i , calculate the trust value T j of its neighbor node v ij at time t: T ij (t + 1)=γT ij (t)+(1 - γ)R ij (t), Among them, T ij (t)∈[0,1], when T ij When (t) = 0, it means the unmanned system node v i At time t, the neighbor node v of the unmanned system j Complete distrust; when T ij When (t) = 1, it means the unmanned system node v i At time t, the neighbor node v of the unmanned system j Total trust, T ij (t+1) is the neighbor node v of the unmanned system j The trust value at time t+1, γ is the smoothing factor, R ij (t) is the trust value at the current moment in the memoryless state; Performing weight trust adjustment based on the trust value: Among them, is the updated weight trust value.
5. A decision-making method for a multi-unmanned system based on adaptive cooperation and dynamic optimization according to claim 4, characterized in that, Calculating the posterior probability based on Bayesian inference and performing weighted aggregation update according to the updated weight trust value, including: Obtaining the current observation information and calculating the posterior probability of the unmanned system node based on the Bayesian update formula: Among them, O i is the current observation information, P(X i ) is the prior probability, P(O i |X i ) represents the system's perception ability of the environment, and P(X i |O i ) represents the accuracy of the node state estimation under the current observation; Performing weighted aggregation update based on the posterior probability and the updated weight trust value: Among them, X i represents the state of the unmanned system node v i The states of all i ∈ N(j) are aggregated to obtain the updated state of the unmanned system node v j The updated state, represents the updated weighted aggregation information of the unmanned system node v j N(j) is the neighborhood of the unmanned system node v j The neighborhood.
6. The decision-making method for a multi-unmanned system based on adaptive collaboration and dynamic optimization according to claim 1, characterized in that Establishing a dynamic collaboration game model based on the updated weighted aggregation information, including: Calculate the strategy of the unmanned system node v based on the updated weighted aggregation information j of the strategy ω is a model parameter, and s j (ω) is the set of actions taken by the unmanned system node v j for the observed information; Suppose the utility function between systems is u j (s1, s2, …, s N ), where s1, s2, …, s N is the set of strategies of other systems; The collaboration utility and competition utility between systems are represented by the following function: u j (s1,s2,…,s N )=C j (s1,s2,…,s N )-P j (s1,s2,…,s N ) Among them, C j (s1, s2, …, s N ) is the benefit brought by collaboration, and P j (s1, s2, …, s N ) is the penalty due to competition; Updating the strategy rule according to the collaboration utility and competition utility functions between systems; Among them, is the strategy that maximizes the utility function.
7. A decision-making method for a multi-unmanned system based on adaptive collaboration and dynamic optimization according to claim 6, characterized in that, Also including: According to the strategy s of the unmanned system node v j and the strategy to maximize the utility function j Define the loss function Perform backpropagation on the model parameter ω where η is the learning rate, ω: is the updated LSTM parameter, represents the gradient operator. 8. A decision-making method for a multi-unmanned system based on adaptive cooperation and dynamic optimization according to claim 3, characterized in that The dynamic connection factor is determined based on resource requirements and task importance.
9. A computer device, characterized in that, Including: A memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.