Unmanned cluster cooperative combat digital twin deduction optimization method under complex meteorological conditions
By constructing a nonlinear meteorological condition model and an improved particle swarm optimization algorithm, combined with a dynamic Bayesian network, the problem of non-optimal task allocation under complex meteorological conditions in traditional military simulations was solved, and the accuracy and robustness of task allocation were improved.
Patent Information
- Application Number
- CN202510688815.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Traditional military simulations are difficult to accurately reflect the continuous evolution of environmental parameters and the nonlinear attenuation dynamic characteristics of equipment performance under complex meteorological conditions, resulting in non-optimal task allocation and inaccurate simulation results.
A digital twin deduction and optimization method for unmanned cluster collaborative combat under complex weather conditions is adopted. By constructing a nonlinear weather condition model, a dynamic attenuation factor and an improved particle swarm optimization algorithm, combined with a dynamic Bayesian network, accurate deduction and optimization of the task allocation strategy are achieved.
It improves the robustness of task allocation and the accuracy of deduction decisions, and can achieve efficient task allocation and system optimization under complex weather conditions.
Smart Images

Figure CN120597701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of meteorological services, quality, safety and environmental inspection and testing services, as well as digital twin and combat simulation technology. Specifically, it relates to a digital twin simulation optimization method for unmanned cluster collaborative combat under complex meteorological conditions. Background Art
[0002] Digital twin technology virtually maps real-time physical entities, synchronizes data with sensors, and constructs synchronized virtual and real models to aid decision-making and optimization. In military simulations, this technology can dynamically map meteorological parameters to equipment performance, enabling in-depth combat simulations.
[0003] Traditional military simulations rely on historical data and empirical rules, and have three main limitations: First, in terms of environmental factors, they ignore the multi-dimensional impact of complex meteorological conditions on equipment effectiveness; second, in model construction, it is difficult to analyze the relationship between cluster characteristics and mission requirements, resulting in suboptimal task allocation; third, in terms of accuracy, the use of discrete state models ignores the dynamic changes in the environment and the attenuation of equipment performance, resulting in simulation results that are difficult to truly reflect the continuous evolution of environmental parameters and the nonlinear attenuation dynamic characteristics of equipment performance. Summary of the Invention
[0004] The present invention provides a digital twin deduction and optimization method for unmanned cluster collaborative combat under complex weather conditions to improve the robustness of task allocation and the accuracy of deduction decisions, which can solve at least one of the above technical problems.
[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0006] The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions includes the following steps:
[0007] S1. Collect meteorological parameters and performance parameters of unmanned swarms from multiple sources to build a digital twin model of unmanned swarm collaborative operations under complex meteorological conditions.
[0008] S2, introducing cross-coupling terms into the Lorenz system, using the least squares method to dynamically optimize the parameters obtained in S1, and constructing a nonlinear meteorological condition model;
[0009] S3. Design a dynamic attenuation factor based on the nonlinear meteorological condition model in S2 to quantify the nonlinear attenuation effect of meteorological conditions on the performance parameters of the intelligent agents in the unmanned swarm. Combined with the geometric mean, variance constraint, and collaborative gain mechanism, a swarm collaborative capability model is constructed.
[0010] S4. Improve the particle swarm optimization algorithm, design a fitness function based on the cluster collaboration capability model in S3 and the actual task requirements, and dynamically adjust the inertia weight, adaptive learning factor, and global guidance factor to obtain a highly matching task allocation strategy;
[0011] S5. Construct a multi-dimensional coupled dynamic Bayesian network of weather, capability, mission, and strategy, combining forward prediction, posterior update, and online learning mechanisms to accurately deduce the task allocation strategy in S4.
[0012] S6. Based on the deduction results of S5, a multi-dimensional evaluation system including task effectiveness, resource consumption and system robustness is constructed, and the task allocation strategy in S4 is continuously improved through iterative optimization.
[0013] Furthermore, the multi-source meteorological parameters in S1 include at least temperature, humidity and visibility, and the performance parameters of the unmanned cluster include at least perception, communication and decision-making capability data of each intelligent agent in the cluster.
[0014] Furthermore, the meteorological conditions defined in S2 include:
[0015] T(t),H(t),V(t)
[0016] Where T is temperature, H is humidity, and V is visibility;
[0017] The nonlinear differential equation structure of the Lorenz system reflects the nonlinear dynamic characteristics of the real meteorological system. Cross-coupling terms are introduced into the nonlinear differential equation to expand the nonlinear interaction between different meteorological conditions, and the following nonlinear meteorological condition model is obtained:
[0018]
[0019] in:
[0020] a T 、a H and a V are all the original parameters of the Lorenz system, namely the Prandtl constant, the Rayleigh constant and a constant related to space;
[0021] b TV 、b TH and b TV are the coupling coefficients under different meteorological conditions, reflecting the nonlinear interaction between meteorological elements.
[0022] Furthermore, the three capabilities of the kth unmanned cluster are defined in S3 as follows: decision-making capability D k (t), communication capability C k (t) and perception ability Gk (t), there are N k Agents, each of which has the above three capabilities, defined as: and
[0023] Introducing dynamic attenuation factor α G (t), α C (t) and α D (t), respectively describe the attenuation of perception capability, communication capability and decision-making capability under the influence of meteorological conditions, and the expressions of the dynamic attenuation factors of perception capability, communication capability and decision-making capability are obtained as follows:
[0024]
[0025] in:
[0026] ρ T Represents the temperature sensitivity coefficient, which is determined by the material properties of the optical device;
[0027] ΔT(t) represents the absolute deviation between the current temperature and the sensor calibration temperature;
[0028] V th is the minimum visibility threshold, which is determined by the specific optical system;
[0029] ρ H and χ H are the specific attenuation coefficient and power law exponent, respectively;
[0030] T D Indicates the critical room temperature threshold that triggers dynamic frequency reduction;
[0031] Indicates the maximum room temperature threshold allowed by the processor;
[0032] ρ D Indicates the processor performance attenuation coefficient;
[0033] is the minimum performance guarantee threshold;
[0034] β is used to control the steepness of the curve. The better the heat dissipation design, the flatter the performance degradation curve.
[0035] Furthermore, in S3, the communication capability C of the unmanned cluster k (t), perception ability G k (t) and decision-making ability D k The modeling of (t) is as follows:
[0036] ①, Communication capability C k Modeling of (t):
[0037] While introducing the geometric mean mechanism to reflect the fact that the communication link efficiency is limited by the weakest node, the gain brought by cluster collaboration to the communication capability is also considered. Based on this, C is given. k The expression of (t):
[0038]
[0039] in:
[0040] Represents the geometric mean of the communication capabilities of each agent within the cluster, where N k Represents the number of agents in the cluster;
[0041] is the influencing term of communication stability, where is the variance of the communication capabilities of all agents. The larger the variance, the worse the communication stability.
[0042] Ω is the degree of influence of experimentally determined variance on synergy;
[0043] represents the gain of cluster collaboration on communication capability, where β C is the communication capability cooperation coefficient determined by the cluster hardware, M C is the number of communication collaborations, N k is the number of nodes, N k Expansion will increase the complexity of the unmanned cluster communication network, causing the communication capability to show logarithmic decay;
[0044] ②, Perception G k Modeling of (t):
[0045] In addition to considering the gain of cluster collaboration on perception capability, we also introduce the gain of cluster communication capability on perception capability, and thus give G k The expression of (t):
[0046]
[0047] in:
[0048] is the sum of the communication capabilities of agent i’s neighboring agents at time t;
[0049] ∑C total (t) is the sum of the communication capabilities of all agents at time t;
[0050] λ G The gain coefficient of cluster communication capability to perception capability determined in the experiment;
[0051] The overall representation represents the gain of the communication network on sensing capabilities;
[0052] β G ·M G represents the gain of cluster collaboration on perception capability, where β G is the perception capability cooperation coefficient, and its value is determined by the characteristics of the cluster hardware. G The number of communication collaborations;
[0053] ③. Decision-making ability D k Modeling of (t):
[0054] In addition to using the geometric mean mechanism to balance the decision weights of each agent, it is also necessary to build a fault-tolerant constraint mechanism to avoid global decision failure due to local communication failures, while taking into account the gains brought by cluster collaboration to decision-making capabilities. Based on this, D k The expression of (t):
[0055]
[0056] in:
[0057] Represents the geometric mean of the decision-making capabilities of each agent within the cluster;
[0058] It is the minimum value of the product of the communication capabilities of all agents, reflecting that the decision-making process is limited by the weakest communication link;
[0059] λ D is an experimentally determined coefficient used to adjust the degree of influence of communication capability on collaborative decision-making;
[0060] is the gain of cluster collaboration on decision-making ability, where β D is the decision-making ability cooperation coefficient, the value of which is determined by the characteristics of the cluster hardware, M D The number of communication collaborations;
[0061] In order to reduce the time consumed by excessive group negotiation, the hyperbolic tangent function is introduced to simulate the marginal diminishing effect. The nonlinear characteristics of the saturation region of the hyperbolic tangent function can accurately describe the benefit relationship between the number of negotiations and decision-making efficiency.
[0062] Furthermore, in S4, let the unmanned cluster set K = {k1, k2, ..., |K|}, divide the combat task J into several subtasks j, and use the particle swarm optimization algorithm to assign combat tasks to different clusters k:
[0063] S4.1. Encode the particles using a three-dimensional array X i =[x ijk] to define the position of the particle, where i represents the index of the particle, j represents the index of the task, k represents the index of the unmanned cluster, and x ijk ∈[0,1], and satisfy This holds true for all i and j;
[0064] S4.2. Construct the fitness function of the particle swarm optimization algorithm, which is used to guide the particle swarm optimization algorithm to search for the optimal solution. Let W J (t) is a set of weights of different cluster capability requirements for different tasks, defined as follows:
[0065]
[0066] In this collection: represents the demand weight of task j on the perception capability of the unmanned cluster at time t, represents the demand weight of task j for the unmanned cluster communication capability at time t, represents the demand weight of task j on the decision-making ability of the unmanned cluster at time t, and j traverses all tasks in the task set J;
[0067] Based on this, the fitness function is given as:
[0068]
[0069] in:
[0070] For the task requirement-capability matching item, the core is to sort the capability requirements from high to low weights, and prioritize matching high-weighted clusters with better corresponding capabilities;
[0071] is the time decay term, where It is the time decay coefficient, which is used to adjust the contribution of task completion speed to the fitness function. Its value depends on the actual problem scenario.
[0072] S4.3. Construct the velocity update equation of the particle swarm optimization algorithm. The parameters considered in this equation include at least: inertia weight w v (t), individual learning factor c1(t), social learning factor c2(t) and global learning guidance factor c3(t);
[0073] ①, inertia weight w v (t) is used for the exploration and development of the balancing algorithm, and the initial value of the inertia weight is set to The inertia weight w can be obtained v (t) is expressed as:
[0074]
[0075] Among them, l1, l2 and l3 represent the influence coefficients of temperature, humidity and visibility on inertia weight respectively;
[0076] ② Individual learning factor It is used to represent the ability of particle i to learn from its own historical optimal position. The initial value of the individual learning factor is Individual learning factor The expression is:
[0077]
[0078] The calculation of the matching degree in this formula further includes the following steps:
[0079] S4.3.1. Normalize the capabilities of each cluster to make all capabilities comparable at the same scale. max 、C max 、D max It is the maximum capacity obtained by traversing each cluster;
[0080] S4.3.2. Increase the sensitivity of cluster capabilities to demand weights through exponential functions, thereby increasing the sensitivity of individual learning factors to changes;
[0081] S4.3.3, after summing, divide by the total number of tasks |J| to ensure The adjustment range should be kept within a reasonable range under different task scales;
[0082] S4.3.4. Matching influence coefficient determined by experiment Control matching the adjustment range;
[0083] ③ The social learning factor c2(t) is used to represent the ability of particles to learn to the global optimal position. The initial value of the individual learning factor is The expression of the social learning factor c2(t) is:
[0084]
[0085] in:
[0086] is a Sigmoid function used to map changes in meteorological conditions to adjustments in social learning factors;
[0087] is the minimum value of the social learning factor, which is used to ensure that particles have the ability to learn from the global optimum under any circumstances;
[0088] It is the maximum value of the social learning factor, which is used to avoid excessive learning intensity that may lead to instability of the particle swarm optimization algorithm;
[0089] is the experimentally determined meteorological condition change influence coefficient, which is used to control the impact of meteorological condition changes on c2;
[0090] ④ The global guidance factor c3(t) is used to respond to sudden environmental changes and changes in task capability requirements. The initial value of the global guidance factor is The expression of the global learning guidance factor c3(t) is:
[0091]
[0092] in, is the experimentally measured adjustment coefficient, which indicates the degree of influence of the gradient term on the global guidance factor;
[0093] By solving the L2 norm of the gradient of each condition, the severity of the change of each condition over time can be characterized. In an emergency, the increase in the gradient value will increase the global guidance factor c3(t), thereby guiding the particle swarm to pay more attention to the global optimal solution and preventing the particle swarm optimization algorithm from falling into the local optimal solution.
[0094] S4.4. Combining the above parameters, we can get the speed update equation of the particle swarm optimization algorithm:
[0095]
[0096] Among them, p ibest is the optimal position found by particle i in its own history, r1 is a random number uniformly distributed in the interval [0,1], g best is the optimal position found by the entire particle swarm in history, g global is a pre-set global guide position, which is determined based on prior knowledge of the problem or other information. r2 and r3 are both random numbers uniformly distributed in the interval;
[0097] Perform multiple rounds of iterations on particle i as follows:
[0098]
[0099] Assume that the position of the particle in round t is Calculate the corresponding fitness value like Update No processing is done;
[0100] S4.5. After all particles have been iterated, compare all The fitness value of As g best , at this time g bestThat is the optimal task allocation solution, the g best Set to X opt =[x jk ].
[0101] Furthermore, in said S5, a dynamic Bayesian multidimensional coupled state space S is defined which is composed of meteorological state, capability state, task state and task allocation strategy. t :
[0102]
[0103] Among them, n j (t) indicates the task status:
[0104]
[0105] Based on this, each meteorological parameter and cluster capability in the state space are discretized respectively;
[0106] Let θ t is the evolution coefficient matrix, used for discretization processing:
[0107] θ t =[β T (t),β H (t),β V (t),β G (t),β D (t),β C (t)]
[0108] For θ t Adopt online mechanism to update: continuously update θ t , reflecting the latest data distribution in real time and improving the adaptability of the entire deduction process to dynamic environments. The online update rule usually adopts gradient descent:
[0109]
[0110] Dynamically adjust the parameter θ through the gradient ascent method t To maximize the state transition probability:
[0111] logP(S t+1 ∣S t ,θ t )
[0112] in:
[0113] η t is the adaptive learning rate, which can be expressed as:
[0114]
[0115] When the historical gradient is large, reduce the learning rate η t , to avoid oscillation, when the historical gradient is small, maintain or increase the learning rate η t , accelerated convergence;
[0116] η0 is the initial learning rate, which is set to a fixed value;
[0117] is the Bayesian gradient at the historical moment τ;
[0118]
[0119] cumulative fluctuations;
[0120] Assume O t Observation variables in a dynamic Bayesian network:
[0121]
[0122] Observed variables are various data obtained from real scenes, including those measured by meteorological sensors. Current capabilities obtained from unmanned cluster self-test and task execution
[0123] Furthermore, the S5 further includes:
[0124] S5.1. Forward deduction of the dynamic Bayesian network. The goal is to calculate the forward direction based on the current state S. t and task allocation scheme X opt , predict the state distribution P(S t+1 ∣S t ,X opt ), the specific steps are as follows:
[0125] S5.1.1, assume state S t include:
[0126] Weather conditions:
[0127] Ability status:
[0128] Task Status:
[0129] Each state component is conditionally independent, and the joint probability is decomposed into:
[0130]
[0131] Based on this, the transition probabilities of each meteorological state and cluster capability state are calculated respectively;
[0132] S5.1.2. Use the sigmoid function to project the task execution status at time t into the range of (0,1) and obtain the transition probability of the task state:
[0133]
[0134] S5.2, the dynamic Bayesian network is updated a posteriori, the purpose is to know all the historical observation data O 1:t+1 and the current task allocation plan X opt Under the condition of t+1 The optimal estimate P(S t+1 |O 1:t+1 ,X opt ), assume that the transfer of observed variables follows:
[0135] Meteorological observation
[0136] Ability Self-Test
[0137] Status Report
[0138] Get the joint observation likelihood:
[0139]
[0140] The formula for posterior update is:
[0141]
[0142] in:
[0143] P(S t+1 ∣O 1:t ,X opt ) can be expressed as:
[0144]
[0145] In the above formula, P(S t ∣O 1:t ) is the posterior probability, and the solution formula is:
[0146]
[0147] P(O t+1 |O 1:t ,X opt ) is a normalization constant, expressed as:
[0148]
[0149] Furthermore, in said S6, at t endAfter the time deduction is completed, the task allocation plan X opt Evaluate the deduction results and construct the evaluation vector V:
[0150] V=[E,U,R] T
[0151] ① In the vector V, E represents the task execution efficiency, which is expressed as:
[0152]
[0153] Among them, τ is the time decay coefficient, which punishes the delay in task completion time. The longer the time, the lower the efficiency contribution. jend is the actual completion time of task j, t start is the deduction starting time, E[n j ] is the expected value of task j completed, E[n j The calculation expression of ] is:
[0154]
[0155] Assume that the performance loss of executing task j is ΔE j , the expression is:
[0156]
[0157] For task j with high performance loss, reallocate cluster k according to capability matching:
[0158]
[0159] In this formula, η E,j is the personalized learning rate for task j, which can be expressed as: E,j =η E ·ΔE j , η E The basic learning rate, the value of which needs to be determined through multiple experiments based on experience;
[0160] ② In vector V, U represents resource consumption, expressed as:
[0161]
[0162] in, They represent the resource consumption of cluster k in performing perception, communication, and decision-making tasks respectively;
[0163] Assume that the energy consumption of cluster k is For high-energy-consuming clusters, reduce task load and balance resource usage:
[0164]
[0165] In this formula, ξ C The energy consumption penalty coefficient is used to control the intensity of the penalty, and its value is determined by the actual battlefield environment;
[0166] The exponential term maps energy consumption to the interval (0,1). When it approaches 1, it means that the cluster energy consumption is high. Item approaches Significantly reduced;
[0167] ③ In the vector V, R represents the system robustness, which is used to measure the system's ability to maintain task effectiveness when the environment suddenly changes and the task requirements change. Let T DBN is the total time deduction steps of dynamic Bayes, then the expression of R is:
[0168]
[0169] By averaging the cumulative changes of environmental mutations and task requirements over the entire period, we can avoid the accidental impact of a single mutation on the evaluation results.
[0170] Through normalization design, the robustness R is mapped to the interval [0,1], and the larger the value, the more stable the system;
[0171] Let the robustness of cluster k be R k , the expression is:
[0172]
[0173] Where J′ is the set of tasks j being executed by cluster k at time t;
[0174] For highly robust clusters, tasks should be assigned first:
[0175]
[0176] In this formula, is the robust linear enhancement term, ξ R is the enhancement coefficient, and its value is determined by the actual battlefield environment;
[0177] The higher the robustness of cluster k, the larger the linear enhancement term, Significant increase.
[0178] Furthermore, in said S6, and Form a multi-objective joint optimization formula:
[0179]
[0180] Among them, q E ,qU ,q R are the weight coefficients of each target, and their specific values must be determined in combination with the actual mission requirements in the battlefield environment;
[0181] The adjusted X′ opt =[x′ jk ] is input into the dynamic Bayesian network to obtain a new evaluation vector V′:
[0182] V′=[E′,U′,R′] T
[0183] Assume that the optimization thresholds of each target are E th 、U th and R th , if the following convergence conditions are met:
[0184] E′≥E tj ,U′≤U th ,R′≥R th
[0185] Then stop the optimization and get the optimal task allocation solution X best =X′ opt ;
[0186] Otherwise, repeat S5 and S6 until the above convergence condition is met.
[0187] The beneficial effects of the present invention are embodied in:
[0188] 1. A nonlinear meteorological evolution model is constructed based on the Lorenz system, integrating dynamic attenuation factors with digital twin technology to improve the accuracy of unmanned swarm collaborative combat simulations under complex weather conditions. At the same time, cross-coupling terms are used to enhance the nonlinear effects between meteorological variables, accurately modeling the dynamic evolution of temperature, humidity, and visibility, and breaking through the limitations of traditional linear models.
[0189] 2. Design a dynamic attenuation factor to quantify the nonlinear impact of weather on cluster capabilities, integrate geometric mean, variance constraint and collaborative gain mechanism, and combine the hyperbolic tangent function to simulate the weather attenuation saturation effect to solve the defect of traditional methods that ignore dynamic weather interference.
[0190] 3. An improved particle swarm optimization algorithm is proposed. Through dynamic inertia weight and adaptive learning factor, the task allocation strategy is adjusted in real time in combination with meteorological data to generate a highly adaptable solution.
[0191] 4. Design a multi-source coupled dynamic Bayesian network of meteorological environment, cluster capability, combat mission, and task allocation strategy, build a two-way deduction closed loop of forward prediction and posterior update, and realize online strategy optimization under environmental mutation conditions through an adaptive learning mechanism.
[0192] 5. Establish a joint evaluation system of task effectiveness, resource consumption, and robustness to achieve accurate evaluation of task allocation plans, while sustainably improving task allocation strategies through iterative optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0193] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0194] Figure 1 It is a schematic diagram of the overall flow of the method according to an embodiment of the present invention.
[0195] Figure 2 3 is a schematic diagram comparing the improved particle swarm optimization algorithm and the standard particle swarm optimization algorithm according to an embodiment of the present invention.
[0196] Figure 3 3. It is a schematic diagram comparing the improved dynamic Bayesian network according to an embodiment of the present invention and the standard dynamic Bayesian network.
[0197] Figure 4 It is a structural block diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0198] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0199] It should be noted that the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or solutions in which both A and B are satisfied. In addition, "multiple" refers to more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0200] See also Figure 1 , an embodiment of the present invention provides a digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions, comprising the following steps:
[0201] S1. Collect meteorological parameters and performance parameters of unmanned swarms from multiple sources to build a digital twin model of unmanned swarm collaborative operations under complex meteorological conditions.
[0202] S2, introducing cross-coupling terms into the Lorenz system, using the least squares method to dynamically optimize the parameters obtained in S1, and constructing a nonlinear meteorological condition model;
[0203] S3. Design a dynamic attenuation factor based on the nonlinear meteorological condition model in S2 to quantify the nonlinear attenuation effect of meteorological conditions on the performance parameters of the intelligent agents in the unmanned swarm. Combined with the geometric mean, variance constraint, and collaborative gain mechanism, a swarm collaborative capability model is constructed.
[0204] S4. Improve the particle swarm optimization algorithm, design a fitness function based on the cluster collaboration capability model in S3 and the actual task requirements, and dynamically adjust the inertia weight, adaptive learning factor, and global guidance factor to obtain a highly matching task allocation strategy;
[0205] S5. Construct a multi-dimensional coupled dynamic Bayesian network of weather, capability, mission, and strategy, combining forward prediction, posterior update, and online learning mechanisms to accurately deduce the task allocation strategy in S4.
[0206] S6. Based on the deduction results of S5, a multi-dimensional evaluation system including task effectiveness, resource consumption and system robustness is constructed, and the task allocation strategy in S4 is continuously improved through iterative optimization.
[0207] In this embodiment, the multi-source meteorological parameters in S1 include at least temperature, humidity and visibility, and the performance parameters of the unmanned cluster include at least perception, communication and decision-making capability data of each intelligent agent in the cluster.
[0208] In this embodiment, the meteorological conditions defined in S2 include:
[0209] T(t),H(t),V(t)
[0210] Where T is temperature, H is humidity, and V is visibility;
[0211] The Lorenz system in chaos theory is used to construct a nonlinear meteorological condition model. Compared with traditional models or simplified models, the nonlinear differential equation structure of the Lorenz system can effectively reflect the nonlinear dynamic characteristics of the real meteorological system.
[0212] However, the interactions among the three variables of temperature, humidity, and visibility are more complex than those in the original Lorenz model. Therefore, a cross-coupling term is introduced into the equation to expand the nonlinear interactions between different meteorological conditions, resulting in the following nonlinear meteorological condition model:
[0213]
[0214] In this equation, a T 、a H and a V are all the original parameters of the Lorenz system, namely the Prandtl constant, the Rayleigh constant and a constant related to space;
[0215] Now, with the help of the least squares optimization strategy, we re-do the solution based on the system variables, with a T As an example, when the humidity H is stable, the coupling term is ignored and the above model equation is simplified first:
[0216]
[0217] Then, the equation is discretized and linearly regressed:
[0218]
[0219] Arranged as a linear equation:
[0220] ΔT i =a T (H i -T i )Δt+ε i
[0221] Finally, the least squares solution is obtained:
[0222]
[0223] In this equation, b TV 、b TH and b TV are the coupling coefficients under different meteorological conditions, reflecting the nonlinear interaction between meteorological elements, and b TV As an example, first simplify the above model equation:
[0224]
[0225] Next, construct the objective function:
[0226]
[0227] Finally, take the derivative, Find the optimal solution:
[0228]
[0229] In this embodiment, the three capabilities of the kth unmanned cluster defined in S3 are: decision-making capability D k (t), communication capability C k (t) and perception ability G k(t), there are N k Agents, each of which has the above three capabilities, defined as: and
[0230] Introducing dynamic attenuation factor α G (t), α C (t) and α D (t), respectively describes the attenuation of perception capability, communication capability and decision-making capability under the influence of meteorological conditions:
[0231] The main meteorological conditions that affect perception capabilities are temperature and visibility. The former causes the sensitivity of optical devices to decay exponentially, while the latter limits the measurement accuracy to a threshold function.
[0232] The main meteorological condition affecting communication capabilities is humidity;
[0233] The key meteorological condition that affects decision-making capabilities is temperature. In high-temperature environments, heat is transferred through the device casing to the internal chipset, causing the processor to throttle, thereby reducing decision accuracy and prolonging decision-making time.
[0234] Based on this, the expressions of the dynamic attenuation factors of perception ability, communication ability and decision-making ability are obtained as follows:
[0235]
[0236] in:
[0237] ρ T Represents the temperature sensitivity coefficient, which is determined by the material properties of the optical device;
[0238] ΔT(t) represents the absolute deviation between the current temperature and the sensor calibration temperature;
[0239] V th is the minimum visibility threshold, which is determined by the specific optical system;
[0240] ρ H and χ H are the specific attenuation coefficient and power law exponent, respectively. Their values can be found in the ITU-R Rain Attenuation Model Recommendation;
[0241] T D Indicates the critical room temperature threshold that triggers dynamic frequency reduction;
[0242] Indicates the maximum room temperature threshold allowed by the processor;
[0243] ρ D Indicates the processor performance attenuation coefficient;
[0244] is the minimum performance guarantee threshold;
[0245] β is used to control the steepness of the curve. The better the heat dissipation design, the flatter the performance degradation curve.
[0246] The above parameters are measured with different processor models at different room temperature conditions.
[0247] In this embodiment, in S3, the communication capability C of the unmanned cluster is k (t), perception ability G k (t) and decision-making ability D k The modeling of (t) is as follows:
[0248] ①, Communication capability C k Modeling of (t):
[0249] While introducing the geometric mean mechanism to reflect the fact that the communication link efficiency is limited by the weakest node, the gain brought by cluster collaboration to the communication capability is also considered. Based on this, C is given. k The expression of (t):
[0250]
[0251] in:
[0252] Represents the geometric mean of the communication capabilities of each agent within the cluster, where N k Represents the number of agents in the cluster;
[0253] is the influencing term of communication stability, where is the variance of the communication capabilities of all agents. The larger the variance, the worse the communication stability.
[0254] Ω is the degree of influence of experimentally determined variance on synergy;
[0255] represents the gain of cluster collaboration on communication capability, where β C is the communication capability cooperation coefficient determined by the cluster hardware, M C is the number of communication collaborations, N k is the number of nodes, N k Expansion will increase the complexity of the unmanned cluster communication network, causing the communication capability to show logarithmic decay;
[0256] ②, Perception G k Modeling of (t):
[0257] In addition to considering the gain of cluster collaboration on perception capability, we also introduce the gain of cluster communication capability on perception capability, and thus give G kThe expression of (t):
[0258]
[0259] in:
[0260] is the sum of the communication capabilities of agent i’s neighboring agents at time t;
[0261] ∑C total (t) is the sum of the communication capabilities of all agents at time t;
[0262] λ G The gain coefficient of cluster communication capability to perception capability determined in the experiment;
[0263] The overall representation represents the gain of the communication network on sensing capabilities;
[0264] β G ·M G represents the gain of cluster collaboration on perception capability, where β G is the perception capability cooperation coefficient, and its value is determined by the characteristics of the cluster hardware. G The number of communication collaborations;
[0265] ③. Decision-making ability D k Modeling of (t):
[0266] In addition to using the geometric mean mechanism to balance the decision weights of each agent, it is also necessary to build a fault-tolerant constraint mechanism to avoid global decision failure due to local communication failures, while taking into account the gains brought by cluster collaboration to decision-making capabilities. Based on this, D k The expression of (t):
[0267]
[0268] in:
[0269] Represents the geometric mean of the decision-making capabilities of each agent within the cluster;
[0270] It is the minimum value of the product of the communication capabilities of all agents, reflecting that the decision-making process is limited by the weakest communication link;
[0271] λ D is an experimentally determined coefficient used to adjust the degree of influence of communication capability on collaborative decision-making;
[0272] is the gain of cluster collaboration on decision-making ability, where β D is the decision-making ability cooperation coefficient, the value of which is determined by the characteristics of the cluster hardware, MD The number of communication collaborations;
[0273] In order to reduce the time consumed by excessive group negotiation, the hyperbolic tangent function is introduced to simulate the marginal diminishing effect. The nonlinear characteristics of the saturation region of the hyperbolic tangent function can accurately describe the benefit relationship between the number of negotiations and decision-making efficiency.
[0274] In this embodiment, in S4, let the unmanned cluster set K = {k1, k2, ..., |K|}, divide the combat task J into several subtasks j, and use the particle swarm optimization algorithm to assign combat tasks to different clusters k:
[0275] S4.1. Encode the particles using a three-dimensional array X i =[x ijk ] to define the position of the particle. Compared with the one-dimensional vector representation, the three-dimensional array can clearly represent the probability of each task being assigned to each unmanned cluster, where i represents the index of the particle, j represents the index of the task, k represents the index of the unmanned cluster, and x represents the index of the task. ijk ∈[0,1], and satisfy This holds true for all i and j;
[0276] S4.2. Construct the fitness function of the particle swarm optimization algorithm, which is used to guide the particle swarm optimization algorithm to search for the optimal solution. Let W J (t) is a set of weights of different cluster capability requirements for different tasks, defined as follows:
[0277]
[0278] In this collection: represents the demand weight of task j on the perception capability of the unmanned cluster at time t, represents the demand weight of task j for the unmanned cluster communication capability at time t, represents the demand weight of task j on the decision-making ability of the unmanned cluster at time t, and j traverses all tasks in the task set J;
[0279] Compared with the static fitness function of the general particle swarm optimization algorithm, this fitness function can dynamically adjust the weights and calculation methods of each factor according to real-time meteorological data and task progress to ensure the timeliness and accuracy of the evaluation. Based on this, the fitness function is given as:
[0280]
[0281] in:
[0282] For the task requirement-capability matching item, the core is to sort the capability requirements from high to low weights, and prioritize matching high-weighted clusters with better corresponding capabilities;
[0283] is the time decay term, where It is the time decay coefficient, which is used to adjust the contribution of task completion speed to the fitness function. Its value depends on the actual problem scenario.
[0284] S4.3. Construct the velocity update equation of the particle swarm optimization algorithm. The parameters considered in this equation include at least: inertia weight w v (t), individual learning factor c1(t), social learning factor c2(t) and global learning guidance factor c3(t). Compared with the fixed parameters used in general particle swarm optimization algorithms, the dynamic parameter adjustment strategy adopted can adaptively adjust the speed update equation in different search stages and complex environments;
[0285] ①, inertia weight w v (t) Used for the exploration and development of balancing algorithms. Combined with the current scenario, in high temperature, high humidity, and low visibility, the inertia weight should be reduced to speed up convergence, avoid local optimality, and quickly adapt to environmental changes. The initial value of the inertia weight is set to The inertia weight w can be obtained v (t) is expressed as:
[0286]
[0287] Among them, l1, l2 and l3 represent the influence coefficients of temperature, humidity and visibility on inertia weight respectively. These coefficients need to be adjusted through experiments or according to the characteristics of specific problems;
[0288] ② Individual learning factor It is used to represent the ability of particle i to learn from its own historical optimal position. Combined with the current scenario, when the cluster's perception, communication or decision-making capabilities match the needs of the current task, individual learning should be strengthened and encouraged to use its own advantages. The initial value of the individual learning factor is set to Individual learning factor The expression is:
[0289]
[0290] The calculation of the matching degree in this formula further includes the following steps:
[0291] S4.3.1. Normalize the capabilities of each cluster to make all capabilities comparable at the same scale. max 、C max 、D max It is the maximum capacity obtained by traversing each cluster;
[0292] S4.3.2. Increase the sensitivity of cluster capabilities to demand weights through exponential functions, thereby increasing the sensitivity of individual learning factors to changes;
[0293] S4.3.3, after summing, divide by the total number of tasks |J| to ensure The adjustment range should be kept within a reasonable range under different task scales;
[0294] S4.3.4. Matching influence coefficient determined by experiment Control matching the adjustment range;
[0295] ③. The social learning factor c2(t) is used to represent the ability of particles to learn to the global optimal position. Combined with the current scenario, under extreme weather conditions (such as high temperature or low visibility), individuals may behave unstable. In this case, more reliance should be placed on group experience. The initial value of the individual learning factor is set to The expression of the social learning factor c2(t) is:
[0296]
[0297] in:
[0298] The Sigmoid function is used to map changes in meteorological conditions to adjustments in social learning factors. The Sigmoid function prevents c2 from fluctuating dramatically due to small changes in environmental factors and better reflects the complex nonlinear relationship between environmental factors and c2.
[0299] is the minimum value of the social learning factor, which is used to ensure that particles have the ability to learn from the global optimum under any circumstances;
[0300] It is the maximum value of the social learning factor, which is used to avoid excessive learning intensity that may lead to instability of the particle swarm optimization algorithm;
[0301] is the experimentally determined meteorological condition change influence coefficient, which is used to control the impact of meteorological condition changes on c2;
[0302] ④ The global guidance factor c3(t) is used to respond to sudden environmental changes and changes in task capability requirements. The initial value of the global guidance factor is The expression of the global learning guidance factor c3(t) is:
[0303]
[0304] in, is the experimentally measured adjustment coefficient, which indicates the degree of influence of the gradient term on the global guidance factor;
[0305] By solving the L2 norm of the gradient of each condition, the severity of the change of each condition over time can be characterized. In an emergency, the increase in the gradient value will increase the global guidance factor c3(t), thereby guiding the particle swarm to pay more attention to the global optimal solution and preventing the particle swarm optimization algorithm from falling into the local optimal solution.
[0306] S4.4. Combining the above parameters, we can get the speed update equation of the particle swarm optimization algorithm:
[0307]
[0308] in, is the optimal position found by particle i in its own history, r1 is a random number uniformly distributed in the interval [0,1]. The introduction of random numbers is to increase the randomness and diversity of the search, g best is the optimal position found by the entire particle swarm in history, g global is a pre-set global guide position, which is determined based on prior knowledge of the problem or other information. r2 and r3 are both random numbers uniformly distributed in the interval;
[0309] Perform multiple rounds of iterations on particle i as follows:
[0310]
[0311] Assume that the position of the particle in round t is Calculate the corresponding fitness value like Update like No processing is done;
[0312] S4.5. After all particles have been iterated, compare all The fitness value of As g best , at this time g best That is the optimal task allocation solution, the g best Set to X opt =[x jk ].
[0313] See also Figure 2 ,Comparing the improved particle swarm optimization algorithm proposed in the present invention with the standard particle swarm optimization algorithm, it can be seen that the improved particle swarm optimization algorithm proposed in the present invention performs better in terms of fitness value.
[0314] In this embodiment, the S5 defines a dynamic Bayesian multidimensional coupled state space S consisting of meteorological state, capability state, task state and task allocation strategy. t Compared with the single-type state modeling of ordinary dynamic Bayesian networks, it can achieve accurate simulation of complex task allocation scenarios. Its expression is:
[0315]
[0316] Among them, n j (t) indicates the task status:
[0317]
[0318] Discretize each function in the state space. First, take temperature T(t) as an example to discretize the meteorological conditions:
[0319] T(t+1)=T(t)+β T (t)(T env -T(t))Δt+∈ T
[0320] in:
[0321] β T (t) is the temperature evolution coefficient under meteorological conditions with known initial values;
[0322] T env is the ambient equilibrium temperature, and the current temperature T(t) increases at a rate of β T (t)To T env Approximation, simulating the natural thermal equilibrium process;
[0323] is the Gaussian noise term of temperature;
[0324] Similarly, the humidity H(t) and visibility V(t) are also calculated using a similar method;
[0325] Secondly, the perception ability G k (t) is used as an example to discretize the cluster capability:
[0326]
[0327] in:
[0328] β G (t) is the perception ability evolution coefficient with a known initial value;
[0329] is the Gaussian noise term of perception ability;
[0330] Similarly, for the communication capability C k (t) and decision-making ability Dk (t) A similar method is also used;
[0331] Let θ t is the evolution coefficient matrix:
[0332] θ t =[β T (t),β H (t),β V (t),β G (t),β D (t),β C (t)]
[0333] For θ t Adopt online mechanism update: In general dynamic Bayesian networks, the parameters θ are usually t Fixed, its value depends on offline training data, which is difficult to cope with real-time changing environments. In this application, the online mechanism continuously updates θ t , reflecting the latest data distribution in real time and improving the adaptability of the entire deduction process to dynamic environments. The online update rule usually adopts gradient descent:
[0334]
[0335] Dynamically adjust the parameter θ through the gradient ascent method t To maximize the state transition probability:
[0336] logP(S t+1 ∣S t ,θ t )
[0337] in:
[0338] η t is the adaptive learning rate, which can be expressed as:
[0339]
[0340] When the historical gradient is large, reduce the learning rate η t , to avoid oscillation, when the historical gradient is small, maintain or increase the learning rate η t , accelerated convergence;
[0341] η0 is the initial learning rate, which is set to a fixed value;
[0342] is the Bayesian gradient at the historical moment τ;
[0343] is the average of the sum of squares of historical gradients, used to quantify the cumulative fluctuation of parameter updates;
[0344] Assume O tObservation variables in a dynamic Bayesian network:
[0345]
[0346] Observed variables are various data obtained from real scenes, including those measured by meteorological sensors. Current capabilities obtained from unmanned cluster self-test and task execution
[0347] In this embodiment, the S5 further includes:
[0348] S5.1. Forward deduction of the dynamic Bayesian network. The goal is to calculate the forward direction based on the current state S. t and task allocation scheme X opt , predict the state distribution P(S t+1 ∣S t ,X oPt ), the specific steps are as follows:
[0349] S5.1.1, assume state S t include:
[0350] Weather conditions:
[0351] Ability status:
[0352] Task Status:
[0353] Each state component is conditionally independent, and the joint probability is decomposed into:
[0354]
[0355] Taking temperature T as an example, we calculate the transition probability of the meteorological state. In the dynamic Bayesian network, it is usually assumed that the state transition follows the normal distribution:
[0356]
[0357] Where, the mean μ T =T(t)+β T (t)(T env -T(t)), variance It needs to be measured separately through experiments or historical data to reflect the random disturbance intensity under different meteorological conditions;
[0358] Based on perception ability G k As an example, calculate the transition probability of the cluster capability state:
[0359]
[0360] In the formula, the mean variance It also needs to be determined independently through experiments or historical data;
[0361] S5.1.2. Use the sigmoid function to project the task execution status at time t into the range of (0,1) and calculate the transition probability of the task state:
[0362]
[0363] S5.2, the dynamic Bayesian network is updated a posteriori, the purpose is to know all the historical observation data O 1:t+1 and the current task allocation plan X opt Under the condition of t+1 The optimal estimate P(S t+1 |O 1:t+1 ,X opt ), assume that the transfer of observed variables follows:
[0364] Meteorological observation
[0365] Ability Self-Test
[0366] Status Report
[0367] Get the joint observation likelihood:
[0368]
[0369] The formula for posterior update is:
[0370]
[0371] in:
[0372] P(S t+1 ∣O 1:t ,X opt ) can be expressed as:
[0373]
[0374] In the above formula, P(S t ∣O 1:t ) is the posterior probability, and the solution formula is:
[0375]
[0376] P(O t+1 |O 1:t ,X opt ) is a normalization constant, expressed as:
[0377]
[0378] In this embodiment, in S6, end After the time deduction is completed, the task allocation plan X opt Evaluate the deduction results and construct the evaluation vector V:
[0379] V=[E,U,R] T
[0380] ① In the vector V, E represents the task execution efficiency, which is expressed as:
[0381]
[0382] Among them, τ is the time decay coefficient, which punishes the delay in task completion time. The longer the time, the lower the efficiency contribution. jend is the actual completion time of task j, t start is the deduction starting time, E[n j ] is the expected value of task j completed, E[n j The calculation expression of ] is:
[0383]
[0384] Assume that the performance loss of executing task j is ΔE j , the expression is:
[0385]
[0386] For task j with high performance loss, reallocate cluster k according to capability matching:
[0387]
[0388] In this formula, η E,j is the personalized learning rate for task j, which can be expressed as: E,j =η E ·ΔE j , η E The basic learning rate, the value of which needs to be determined through multiple experiments based on experience;
[0389] ② In vector V, U represents resource consumption, expressed as:
[0390]
[0391] in, They represent the resource consumption of cluster k for performing perception, communication, and decision-making tasks respectively. Taking perception as an example, it can be expressed as:
[0392]
[0393] In this formula:
[0394] ∈ S The energy consumption coefficient per unit capacity is determined by the cluster hardware;
[0395] E[G k (t+1)] represents the expected value of perception ability, and its mathematical expression is:
[0396]
[0397] In this formula, Represents the discretized value of the perception ability of cluster k;
[0398] Similarly, the process of solving the energy consumption of communication and decision-making tasks is similar;
[0399] Assume that the energy consumption of cluster k is For high-energy-consuming clusters, reduce task load and balance resource usage:
[0400]
[0401] In this formula, ξ C The energy consumption penalty coefficient is used to control the intensity of the penalty, and its value is determined by the actual battlefield environment;
[0402] The exponential term maps energy consumption to the interval (0,1). When it approaches 1, it means that the cluster energy consumption is high. Item approaches Significantly reduced;
[0403] ③ In the vector V, R represents the system robustness, which is used to measure the system's ability to maintain task effectiveness when the environment suddenly changes and the task requirements change. Let T DBN is the total time deduction steps of dynamic Bayes, then the expression of R is:
[0404]
[0405] By averaging the cumulative changes of environmental mutations and task requirements over the entire period, we can avoid the accidental impact of a single mutation on the evaluation results.
[0406] Through normalization design, the robustness R is mapped to the interval [0,1], and the larger the value, the more stable the system;
[0407] Let the robustness of cluster k be R k , the expression is:
[0408]
[0409] Where J′ is the set of tasks j being executed by cluster k at time t;
[0410] For highly robust clusters, tasks should be assigned first:
[0411]
[0412] In this formula, is the robust linear enhancement term, ξ R is the enhancement coefficient, and its value is determined by the actual battlefield environment;
[0413] The higher the robustness of cluster k, the larger the linear enhancement term, Significant increase.
[0414] In this embodiment, in S6, and Form a multi-objective joint optimization formula:
[0415]
[0416] Among them, q E ,q U ,q R are the weight coefficients of each target, and their specific values must be determined in combination with the actual mission requirements in the battlefield environment;
[0417] The adjusted X′ opt =[x′ jk ] is input into the dynamic Bayesian network to obtain a new evaluation vector V′:
[0418] V′=[E′,U′,R′] T
[0419] Assume that the optimization thresholds of each target are E th 、U th and R th , if the following convergence conditions are met:
[0420] E′≥E tj ,U′≤U th ,R′≥R th
[0421] Then stop the optimization and get the optimal task allocation solution X best =X′ opt ;
[0422] Otherwise, repeat S5 and S6 until the above convergence condition is met.
[0423] See also Figure 3Comparing the improved dynamic Bayesian network proposed in the present invention with the standard dynamic Bayesian network, it can be seen that the improved dynamic Bayesian network proposed in the present invention has better performance in each component of the evaluation vector.
[0424] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the digital twin deduction and optimization method for unmanned cluster collaborative combat under complex weather conditions as described above.
[0425] See also Figure 4 An embodiment of the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the digital twin deduction and optimization method for unmanned cluster collaborative combat under complex weather conditions as described above.
[0426] An embodiment of the present invention also provides a computer program product containing instructions, which, when run on a computer, enables the computer to execute the steps of the digital twin deduction and optimization method for unmanned cluster collaborative combat under complex weather conditions as described above.
[0427] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention. The explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the digital twin deduction and optimization method for unmanned cluster collaborative combat under complex weather conditions mentioned above.
[0428] It should be noted that those skilled in the art will understand that all or part of the steps implemented in the embodiments of the present invention can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using hardware, it can be implemented in whole or in part in the form of purchased standard parts or modified parts. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).
[0429] In summary, in order to solve the shortcomings of traditional deduction methods in meteorological dynamic modeling, cluster collaborative capability quantification and real-time strategy optimization, the present invention provides a digital twin deduction and optimization method for unmanned cluster collaborative combat under complex meteorological conditions. In this method: first, a battlefield digital twin model is constructed based on meteorological parameters and unmanned cluster performance parameters, and a nonlinear evolution model of meteorological conditions is established by introducing cross-coupling terms in the Lorenz system; secondly, a dynamic attenuation factor is designed to quantify the nonlinear attenuation effect of temperature, humidity and visibility on the perception, communication and decision-making capabilities of the unmanned cluster, and a cluster collaborative capability model is constructed in combination with geometric mean and variance constraints; then, by dynamically adjusting the inertia weight, learning factor and global guidance factor, a deep improvement of the particle swarm optimization algorithm is achieved, and then a task allocation strategy based on the algorithm is obtained; again, a multidimensional coupled dynamic Bayesian network is used to construct a meteorological-capability-task-strategy state space, and forward prediction, posterior update and online parameter optimization mechanism are combined to achieve accurate deduction of the task allocation strategy; finally, a multidimensional evaluation system is established based on the deduction results, and the solution is continuously improved through iterative optimization, and finally an optimal task allocation strategy that adapts to the battlefield situation is formed. The present invention significantly improves the robustness and deduction accuracy of task allocation in complex environments, and provides highly adaptable decision support for unmanned swarm collaborative combat.
[0430] It should be understood that the examples and implementation methods described herein are for illustrative purposes only and are not intended to limit the present invention. Those skilled in the art may make various modifications or changes based on them. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions, characterized by: The following steps are involved: S1. Collect meteorological parameters and performance parameters of unmanned swarms from multiple sources to build a digital twin model of unmanned swarm collaborative operations under complex meteorological conditions. S2, introducing cross-coupling terms into the Lorenz system, using the least squares method to dynamically optimize the parameters obtained in S1, and constructing a nonlinear meteorological condition model; S3. Design a dynamic attenuation factor based on the nonlinear meteorological condition model in S2 to quantify the nonlinear attenuation effect of meteorological conditions on the performance parameters of the intelligent agents in the unmanned swarm. Combined with the geometric mean, variance constraint, and collaborative gain mechanism, a swarm collaborative capability model is constructed. S4. Improve the particle swarm optimization algorithm, design a fitness function based on the cluster collaboration capability model in S3 and the actual task requirements, and dynamically adjust the inertia weight, adaptive learning factor, and global guidance factor to obtain a highly matching task allocation strategy; S5. Construct a multi-dimensional coupled dynamic Bayesian network of weather, capability, mission, and strategy, combining forward prediction, posterior update, and online learning mechanisms to accurately deduce the task allocation strategy in S4. S6. Based on the deduction results of S5, a multi-dimensional evaluation system including task effectiveness, resource consumption and system robustness is constructed, and the task allocation strategy in S4 is continuously improved through iterative optimization.
2. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 1 is characterized in that: The multi-source meteorological parameters in S1 include at least temperature, humidity and visibility, and the performance parameters of the unmanned cluster include at least the perception, communication and decision-making capability data of each intelligent agent in the cluster.
3. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 1 is characterized in that: The meteorological conditions defined in S2 include: T(t),H(t),V(t) Where T is temperature, H is humidity, and V is visibility; The nonlinear differential equation structure of the Lorenz system reflects the nonlinear dynamic characteristics of the real meteorological system. Cross-coupling terms are introduced into the nonlinear differential equation to expand the nonlinear interaction between different meteorological conditions, and the following nonlinear meteorological condition model is obtained: in: a T 、a H and a V are all the original parameters of the Lorenz system, namely the Prandtl constant, the Rayleigh constant and a constant related to space; b TV 、b TH and b TV are the coupling coefficients under different meteorological conditions, reflecting the nonlinear interaction between meteorological elements.
4. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 1 is characterized in that: In S3, the three capabilities of the kth unmanned cluster are defined as follows: decision-making capability D k (t), communication capability C k (t) and perception ability G k (t), there are N k Agents, each of which has the above three capabilities, defined as: and Introducing dynamic attenuation factor α G (t), α C (t) and α D (t), respectively describe the attenuation of perception capability, communication capability and decision-making capability under the influence of meteorological conditions, and the expressions of the dynamic attenuation factors of perception capability, communication capability and decision-making capability are obtained as follows: in: ρ T Represents the temperature sensitivity coefficient, which is determined by the material properties of the optical device; ΔT(t) represents the absolute deviation between the current temperature and the sensor calibration temperature; V th is the minimum visibility threshold, which is determined by the specific optical system; ρ H and χ H are the specific attenuation coefficient and power law exponent, respectively; T D Indicates the critical room temperature threshold that triggers dynamic frequency reduction; Indicates the maximum room temperature threshold allowed by the processor; ρ D Indicates the processor performance attenuation coefficient; is the minimum performance guarantee threshold; β is used to control the steepness of the curve. The better the heat dissipation design, the flatter the performance degradation curve.
5. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 4 is characterized in that: In S3, the communication capability C of the unmanned cluster k (t), perception ability G k (t) and decision-making ability D k The modeling of (t) is as follows: ①, Communication capability C k Modeling of (t): While introducing the geometric mean mechanism to reflect the fact that the communication link efficiency is limited by the weakest node, the gain brought by cluster collaboration to the communication capability is also considered. Based on this, C is given. k The expression of (t): in: Represents the geometric mean of the communication capabilities of each agent within the cluster, where N k Represents the number of agents in the cluster; is the influencing term of communication stability, where is the variance of the communication capabilities of all agents. The larger the variance, the worse the communication stability. Ω is the degree of influence of experimentally determined variance on synergy; represents the gain of cluster collaboration on communication capability, where β C is the communication capability cooperation coefficient determined by the cluster hardware, M C is the number of communication collaborations, N k is the number of nodes, N k Expansion will increase the complexity of the unmanned cluster communication network, causing the communication capability to show logarithmic decay; ②, Perception G k Modeling of (t): In addition to considering the gain of cluster collaboration on perception capability, we also introduce the gain of cluster communication capability on perception capability, and thus give G k The expression of (t): in: is the sum of the communication capabilities of agent i’s neighboring agents at time t; ∑C total (t) is the sum of the communication capabilities of all agents at time t; λ G The gain coefficient of cluster communication capability to perception capability determined in the experiment; The overall representation represents the gain of the communication network on sensing capabilities; β G ·M G represents the gain of cluster collaboration on perception capability, where β G is the perception capability cooperation coefficient, and its value is determined by the characteristics of the cluster hardware. G The number of communication collaborations; ③. Decision-making ability D k Modeling of (t): In addition to using the geometric mean mechanism to balance the decision weights of each agent, it is also necessary to build a fault-tolerant constraint mechanism to avoid global decision failure due to local communication failures, while taking into account the gains brought by cluster collaboration to decision-making capabilities. Based on this, D k The expression of (t): in: Represents the geometric mean of the decision-making capabilities of each agent within the cluster; It is the minimum value of the product of the communication capabilities of all agents, reflecting that the decision-making process is limited by the weakest communication link; λ D is an experimentally determined coefficient used to adjust the degree of influence of communication capability on collaborative decision-making; is the gain of cluster collaboration on decision-making ability, where β D is the decision-making ability cooperation coefficient, the value of which is determined by the characteristics of the cluster hardware, M D The number of communication collaborations; In order to reduce the time consumed by excessive group negotiation, the hyperbolic tangent function is introduced to simulate the marginal diminishing effect. The nonlinear characteristics of the saturation region of the hyperbolic tangent function can accurately describe the benefit relationship between the number of negotiations and decision-making efficiency.
6. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 1 is characterized in that: In S4, let the unmanned cluster set K = {k1, k2, ..., |K|}, divide the combat task J into several subtasks j, and use the particle swarm optimization algorithm to assign combat tasks to different clusters k: S4.
1. Encode the particles using a three-dimensional array X i =[x ijk ] to define the position of the particle, where i represents the index of the particle, j represents the index of the task, k represents the index of the unmanned cluster, and x ijk ∈[0,1], and satisfy This holds true for all i and j; S4.
2. Construct the fitness function of the particle swarm optimization algorithm, which is used to guide the particle swarm optimization algorithm to search for the optimal solution. Let W J (t) is a set of weights of different cluster capability requirements for different tasks, defined as follows: In this collection: represents the demand weight of task j on the perception capability of the unmanned cluster at time t, represents the demand weight of task j for the unmanned cluster communication capability at time t, represents the demand weight of task j on the decision-making ability of the unmanned cluster at time t, and j traverses all tasks in the task set J; Based on this, the fitness function is given as: in: For the task requirement-capability matching item, the core is to sort the capability requirements from high to low weights, and prioritize matching high-weighted clusters with better corresponding capabilities; is the time decay term, where It is the time decay coefficient, which is used to adjust the contribution of task completion speed to the fitness function. Its value depends on the actual problem scenario. S4.
3. Construct the velocity update equation of the particle swarm optimization algorithm. The parameters considered in this equation include at least: inertia weight w v (t), individual learning factor c1(t), social learning factor c2(t) and global learning guidance factor c3(t); ①, inertia weight w v (t) is used for the exploration and development of the balancing algorithm, and the initial value of the inertia weight is set to The inertia weight w can be obtained v (t) is expressed as: Among them, l1, l2 and l3 represent the influence coefficients of temperature, humidity and visibility on inertia weight respectively; ② Individual learning factor It is used to represent the ability of particle i to learn from its own historical optimal position. The initial value of the individual learning factor is Individual learning factor The expression is: The calculation of the matching degree in this formula further includes the following steps: S4.3.
1. Normalize the capabilities of each cluster to make all capabilities comparable at the same scale. max 、C max 、D max It is the maximum capacity obtained by traversing each cluster; S4.3.
2. Increase the sensitivity of cluster capabilities to demand weights through exponential functions, thereby increasing the sensitivity of individual learning factors to changes; S4.3.3, after summing, divide by the total number of tasks |J| to ensure The adjustment range should be kept within a reasonable range under different task scales; S4.3.
4. Matching influence coefficient determined by experiment Control matching the adjustment range; ③、The social learning factor c2(t) is used to represent the ability of particles to learn to the global optimal position. The initial value of the individual learning factor is The expression of the social learning factor c2(t) is: in: is a Sigmoid function used to map changes in meteorological conditions to adjustments in social learning factors; is the minimum value of the social learning factor, which is used to ensure that particles have the ability to learn from the global optimum under any circumstances; It is the maximum value of the social learning factor, which is used to avoid excessive learning intensity that may lead to instability of the particle swarm optimization algorithm; is the experimentally determined meteorological condition change influence coefficient, which is used to control the impact of meteorological condition changes on c2; ④ The global guidance factor c3(t) is used to respond to sudden environmental changes and changes in task capability requirements. The initial value of the global guidance factor is The expression of the global learning guidance factor c3(t) is: in, is the experimentally measured adjustment coefficient, which indicates the degree of influence of the gradient term on the global guidance factor; By solving the L2 norm of the gradient of each condition, the severity of the change of each condition over time can be characterized. In an emergency, the increase in the gradient value will increase the global guidance factor c3(t), thereby guiding the particle swarm to pay more attention to the global optimal solution and preventing the particle swarm optimization algorithm from falling into the local optimal solution. S4.
4. Combining the above parameters, we can get the speed update equation of the particle swarm optimization algorithm: in, is the optimal position found by particle i in its own history, r1 is a random number uniformly distributed in the interval [0,1], g best is the optimal position found by the entire particle swarm in history, g global is a pre-set global guide position, which is determined based on prior knowledge of the problem or other information. r2 and r3 are both random numbers uniformly distributed in the interval; Perform multiple rounds of iterations on particle i as follows: Assume that the position of the particle in round t is Calculate the corresponding fitness value like Update like No processing is done; S4.
5. After all particles have been iterated, compare all The fitness value of As g best , at this time g best That is the optimal task allocation solution, the g best Set to X opt =[x jk ].
7. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 1 is characterized in that: In the S5, a dynamic Bayesian multidimensional coupled state space S is defined, which is composed of meteorological state, capability state, task state and task allocation strategy. t : Among them, n j (t) indicates the task status: Based on this, each meteorological parameter and cluster capability in the state space are discretized respectively; Let θ t is the evolution coefficient matrix, used for discretization processing: i t =[β T (t),β H (t),β V (t),β G (t),β D (t),β C (t)] For θ t Adopt online mechanism to update: continuously update θ t , reflecting the latest data distribution in real time and improving the adaptability of the entire deduction process to dynamic environments. The online update rule usually adopts gradient descent: Dynamically adjust the parameter θ through the gradient ascent method t To maximize the state transition probability: logP(S t+1 ∣S t ,i t ) in: η t is the adaptive learning rate, which can be expressed as: When the historical gradient is large, reduce the learning rate η t , to avoid oscillation, when the historical gradient is small, maintain or increase the learning rate η t , accelerated convergence; η0 is the initial learning rate, which is set to a fixed value; is the Bayesian gradient at the historical moment τ; is the average of the sum of squares of historical gradients, used to quantify the cumulative fluctuation of parameter updates; Assume O t Observation variables in a dynamic Bayesian network: Observed variables are various data obtained from real scenes, including those measured by meteorological sensors. Current capabilities obtained from unmanned cluster self-test and task execution 8. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 7 is characterized in that: Said S5 further comprises: S5.
1. Forward deduction of the dynamic Bayesian network. The goal is to calculate the forward direction based on the current state S. t and task allocation scheme X opt , predict the state distribution P(S t+1 ∣S t ,X opt ), the specific steps are as follows: S5.1.1, assume state S t include: Weather conditions: Ability status: Task Status: Each state component is conditionally independent, and the joint probability is decomposed into: Based on this, the transition probabilities of each meteorological state and cluster capability state are calculated respectively; S5.1.
2. Use the sigmoid function to project the task execution status at time t into the range of (0,1) and obtain the transition probability of the task state: S5.2, the dynamic Bayesian network is updated a posteriori, the purpose is to know all the historical observation data O 1:t+1 and the current task allocation plan X opt Under the condition of t+1 The optimal estimate P(S t+1 |O 1:t+1 ,X opt ), assume that the transfer of observed variables follows: Meteorological observation Ability Self-Test Status Report Get the joint observation likelihood: The formula for posterior update is: in: P(S t+1 ∣O 1:t ,X opt ) can be expressed as: In the above formula, P(S1|O 1:t ) is the posterior probability, and the solution formula is: P(O t+1 |O 1:t ,X ppt ) is a normalization constant, expressed as:
9. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 1 is characterized in that: In S6, at t end After the time deduction is completed, the task allocation plan X opt Evaluate the deduction results and construct the evaluation vector V: V=[E,U,R] T ① In the vector V, E represents the task execution efficiency, which is expressed as: Among them, τ is the time decay coefficient, which punishes the delay in task completion time. The longer the time, the lower the efficiency contribution. jend is the actual completion time of task j, t start is the deduction starting time, E[n j ] is the expected value of task j completed, E[n j The calculation expression of ] is: Assume that the performance loss of executing task j is ΔE j , the expression is: For task j with high performance loss, reallocate cluster k according to capability matching: In this formula, η E,j is the personalized learning rate for task j, which can be expressed as: E,j =η E ·ΔE j , η E The basic learning rate, the value of which needs to be determined through multiple experiments based on experience; ② In vector V, U represents resource consumption, expressed as: in, They represent the resource consumption of cluster k in performing perception, communication, and decision-making tasks respectively; Assume that the energy consumption of cluster k is For high-energy-consuming clusters, reduce task load and balance resource usage: In this formula, ξ C The energy consumption penalty coefficient is used to control the intensity of the penalty, and its value is determined by the actual battlefield environment; The exponential term maps energy consumption to the interval (0,1). When it approaches 1, it means that the cluster energy consumption is high. Item approaches Significantly reduced; ③ In the vector V, R represents the system robustness, which is used to measure the system's ability to maintain task effectiveness when the environment suddenly changes and the task requirements change. Let T DBN is the total time deduction steps of dynamic Bayes, then the expression of R is: By averaging the cumulative changes of environmental mutations and task requirements over the entire period, we can avoid the accidental impact of a single mutation on the evaluation results. Through normalization design, the robustness R is mapped to the interval [0,1], and the larger the value, the more stable the system; Let the robustness of cluster k be R k , the expression is: Among them, J ′ is the set of tasks j being executed by cluster k at time t; For highly robust clusters, tasks should be assigned first: In this formula, is the robust linear enhancement term, ξ R is the enhancement coefficient, and its value is determined by the actual battlefield environment; The higher the robustness of cluster k, the larger the linear enhancement term, Significant increase.
10. The digital twin deduction and optimization method for unmanned swarm collaborative combat under complex weather conditions as claimed in claim 9, characterized in that: In the S6, the and Form a multi-objective joint optimization formula: Among them, q E ,q U ,q R are the weight coefficients of each target, and their specific values must be determined in combination with the actual mission requirements in the battlefield environment; The adjusted X′ opt =[x′ jk ] is input into the dynamic Bayesian network to obtain a new evaluation vector V′: V′=[E′,U′,R′] T Assume that the optimization thresholds of each target are E th 、U th and R th , if the following convergence conditions are met: E′≥E th ,U′≤U th ,R′≥R th Then stop the optimization and get the optimal task allocation solution X best =X′ opt ; Otherwise, repeat S5 and S6 until the above convergence condition is met.
Citation Information
Patent Citations
Method for evaluating weapon combat effectiveness in meteorological environment based on DGWO-SVM
CN114186482A
Simulation deduction system based on digital twinning
CN117150757A
Unmanned cluster dynamic collaborative optimization method oriented to multi-task requirements
CN118819188A
Photovoltaic power station power prediction method and system based on BP neural network
CN119852981A
Intelligent inspection path planning method and system based on reinforcement learning
CN119990496A
Cited By
Air defense combat cooperation method and system based on digital twinning
CN120876191A
Digital Twin-Based Air Defense Operational Coordination Method and System
CN120876191B
A cooperative game method and system for cross-domain heterogeneous unmanned cluster
CN122528684A