Model certification and verification method and system of dual-order T-norm-Choquet-OWA resource aggregator oriented to multi-unmanned aerial vehicle cooperation
By predicting future resource trajectories using the LSTM-EMA model and combining the T-norm and Joquit-OWA operator, the problem of low resource allocation efficiency in UAV collaboration is solved, achieving efficient and stable resource management in dynamic environments and improving task completion rate and energy efficiency.
Patent Information
- Application Number
- CN202511189163.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-19
AI Technical Summary
Existing drone collaborative technologies lack effective prediction of future resource trends, resulting in poor performance in highly dynamic flight environments. Traditional methods cannot balance mission performance and system robustness, especially in resource allocation efficiency in computationally intensive or communication-sensitive tasks.
The Long Short-Term Memory Network-Exponential Moving Average (LSTM-EMA) model is used to predict the resource trajectory in the next 3 seconds. Combined with the T-norm and Joquet-OWA operator, a two-stage aggregator design is used to first protect the bottleneck resources and then perform elastic compensation, realizing a collaborative strategy of "solving the bottleneck first and then restoring efficiency".
It significantly improves resource utilization efficiency and execution stability in multi-UAV collaborative tasks, reduces latency and packet loss rate, enhances system security and flexibility, and adapts to short-term resource fluctuations and extreme bottleneck conditions.
Smart Images

Figure CN121166348A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of unmanned aerial vehicle cooperation, and particularly relates to a model proving and verifying method and system of a double-stage T-norm-Choquet-OWA resource aggregator for multi-unmanned aerial vehicle cooperation. BACKGROUND
[0002] In recent years, the rapid development of unmanned aerial vehicle (UAV) technology has made multi-unmanned aerial vehicle cooperative systems play an increasingly key role in related fields such as target tracking, environmental monitoring, and emergency rescue. In practical applications, multiple unmanned aerial vehicles must build a stable and efficient cooperative network to meet the needs of complex tasks while adapting to various restrictions brought by dynamic environments. When performing computationally intensive or communication-sensitive tasks, the unmanned aerial vehicle group has particularly strict real-time cooperation requirements on three key resources: airborne battery energy, wireless link bandwidth, and airborne computing capacity. Therefore, balancing task performance and system robustness under resource constraints has become a core scientific challenge in the field of unmanned aerial vehicle cooperation.
[0003] Among current multi-unmanned aerial vehicle resource scheduling methods, two typical schemes dominate: a rigid bottleneck protection scheme (such as the minimum operator) and a simple linear weighted aggregator (such as the Weighted Sum Model, WSM). The former strictly follows the short board effect, using the weakest resource dimension to determine the feasibility of the overall task. Although this method can guarantee reliability under extreme conditions, it cannot utilize the excess capacity of non-bottleneck resources, thereby severely limiting efficiency. In contrast, the linear weighted model has a certain flexibility by combining resource dimensions with fixed weights, but lacks adaptive adjustment capability when conditions change rapidly; this often leads to resource overload or waste, thereby reducing overall efficiency and stability.
[0004] Therefore, a more balanced method is needed that can both quickly respond to emerging bottlenecks and simultaneously utilize the flexible potential of non-bottleneck resources to optimize task performance and system robustness.
[0005] To address these challenges, fuzzy set theory and aggregation operators have recently been widely applied in resource allocation and task scheduling. In particular, two-stage fuzzy aggregators have shown significant advantages: the first stage employs strict T-norms for bottleneck protection, and the second stage uses nonlinear fusion operators, such as Choquet integral and Ordered Weighted Averaging (OWA), to dynamically balance bottleneck and non-bottleneck capacities. However, in the field of UAV cooperation, the application of these methods is still relatively small, especially in the absence of effective prediction of future resource trends, which seriously hinders their performance in highly dynamic flight environments.
[0006] Through the above analysis, the problems and defects of the prior art are:
[0007] In the field of UAV cooperation, the application of these methods is still relatively small, especially in the absence of effective prediction of future resource trends, which seriously hinders their performance in highly dynamic flight environments. SUMMARY
[0008] In view of the problems existing in the prior art, the present application provides a model proof and verification method for a two-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation.
[0009] The present application is implemented as follows: a model proof and verification method for a two-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation includes:
[0010] Step 1: use a long short-term memory network-exponential moving average (LSTM-EMA) model to predict the resource trajectory for the next 3 seconds;
[0011] Step 2: the first stage uses T-norm (minimum value) to determine the bottleneck resource, and the second stage uses the Choquet-OWA method driven by the adaptive interaction metric φ to perform elastic compensation according to the instantaneous power usage, realizing the cooperative strategy of "first solving the bottleneck, then restoring efficiency".
[0012] Further, the long short-term memory network-exponential moving average (LSTM-EMA) model is used to predict the resource trajectory for the next 3 seconds:
[0013] (1) Resource usage multi-dimensional fuzzy set design;
[0014] (2) Construction of comprehensive fuzzy benefit function.
[0015] Further, the resource usage multi-dimensional fuzzy set design:
[0016] 1) Fuzzy aggregation under extreme protection
[0017] Let and e raw respectively represent the CPU utilization UAV synergy, bandwidth utilization UAV synergy and residual energy UAV synergy measured at the UAV node; to make them dimensionless and comparable, each index is normalized to the unit interval [0, 1] before entering the fuzzy set model, obtaining:
[0018]
[0019] where, e min , e max is the observed value or nominal bound for each resource (typically 0 and 100%); after normalization, all three indices satisfy F {·} ∈ [0, 1], where a value close to 1 indicates a high utilization of CPU and bandwidth (for CPU and bandwidth) or an energy sufficiency (for energy);
[0020] This min-max normalization rescales each raw index into a standard dimensionless form, enabling a fair aggregation of heterogeneous resources in the multi-dimensional fuzzy set framework; Fcand Fbtend to 1 when CPU or bandwidth is heavily used, while Fe tends to 1 when residual energy is sufficient;
[0021] 2) Future-aware membership functions
[0022] CPU prediction-enhanced membership function (alarm contraction): one step penalty on "current load + upcoming predicted load" to avoid resource decision errors due to upcoming peaks;
[0023]
[0024] Here, c is the current CPU utilization; Fc is the k-seconds-ahead predicted utilization; their sum measures the upcoming total pressure on CPU; λc∈ [0, 1] controls the weight of the prediction - the more sensitive the task to computation, the larger λc; θcis the utilization inflection point; κccontrols the steepness of the S-shaped curve; if Ifc+ λ c F c <<θ c , the exponential term tends to 0, indicating that the computational resource is sufficient; otherwise, it rapidly decreases;
[0025] Bandwidth membership function (alarm contraction): for high real-time communication services, a small θband a large κb can be set to achieve "early braking";
[0026]
[0027] Here, b is the current bandwidth utilization; Fb is the predicted future congestion situation; λb is the weight; if b + λ b F b Very small (bandwidth sufficient); once the inflection point θb is reached, it decays rapidly;
[0028] Energy membership function (power response correction):
[0029]
[0030] Here, is the original trigonometric function, and the offset term realizes the "future power drop" penalty; er is the current remaining energy ratio; Fe is the predicted remaining energy ratio after the task is completed; the offset term λ e (1-F e ) deducts the expected energy consumption in advance, and λe reflects the sensitivity of the task to energy; after the offset, if the predicted remaining energy drops sharply, the function input becomes small, and the energy protection triggers faster;
[0031] Instantaneous power consumption rate membership (linear):
[0032] μ p (p) = 1-p, p ∈ [0, 1] (5)
[0033] Where p is the ratio of the current power load to the reference power; it decreases linearly: the higher the power → the lower the membership, simply and intuitively quantifying the impact of "power consumption rate" on task feasibility;
[0034] 3) Two-stage extreme protection + elastic coupling;
[0035] 4) Two-stage T-norm - Jouquet resource aggregation.
[0036] Further, the two-stage extreme protection + elastic coupling:
[0037] First-stage extreme protection (core three dimensions): preserve the "bucket effect", the weakest of the three main resources determines the overall feasibility; if any membership equals 0 → μ pre = 0, the task is immediately rejected or migrated;
[0038]
[0039] Second-stage elastic coupling:
[0040] φ = η + (1-η)F e , φ ∈ (0, 1] (7)
[0041] Here, η is the task elasticity; for high-performance tasks, η→1, indicating more emphasis on performance; Fe is the predicted residual energy after the task is completed; as the energy sufficiency increases, Fe→1; φ combines the two factors: if the task has high performance requirements and energy sufficiency increases; if the task is energy-saving or energy is low decreases; φ adjusts the strength of the punishment on the power rate.
[0042] Further, the two-stage T-norm-Jouquette resource aggregation:
[0043] To capture both bottleneck protection and cross-dimension complementarity, the present application adopts a two-stage aggregation structure; in the lower stage, the T-norm (minimum value) is used to extract the bottleneck value from the three main resources—CPU, bandwidth, and energy;
[0044]
[0045] In the higher stage, the binary Jouquette-OWA operator with interaction measure ν is used to model the complementarity between resource sufficiency and power rate in μ pre , μ p , and is defined as follows:
[0046] f1=μ pre , f2=μ p , f (1) ≥f (2) (9)
[0047] Sort the two membership degrees in descending order, set S1={1,2}, S2={2}; then the Jouquette aggregation is:
[0048]
[0049] Here, the interaction measure is defined as:
[0050] ν({μ pre})=φν({μ p})=1-φ,ν({μ pre ,μ p})=1 (11)
[0051] and φ=η+(1-η)Fe∈(0,1], consistent with the original model; substituting the above formula gives the closed form:
[0052] μ resource =φμ pre +(1-φ)max{μ pre ,μ p} (12)
[0053] Here, when μpre ≤μ p (when power rate is better than bottleneck value), μ resource = φμ pre + (1-φ)μ p reflects the "power compensation" effect; it reverts to μresource=μprewhen μpre>μp to ensure bottleneck protection priority; parameter φ is still modulated by task elasticity η and predicted residual energy Fe to achieve scenario adaptive behavior;
[0054] The overall mechanism is as follows:
[0055] Future perception: formula Injecting predicted values into the membership function can pre-punish upcoming resource conflicts or energy drops;
[0056] Strict bottleneck protection: formula μpreuses the minimum operator to ensure that any major resource shortage triggers protection immediately;
[0057] Flexible rate suppression: the formula of φ and μresource incorporates instantaneous power consumption, smoothly fusing power mean and power penalty terms to achieve elasticity-power coupling adjustment;
[0058] Final membership output: membership μresource combines with μtrust and μdelay to form the benefit R, driving the adaptive evolution of subsequent fuzzy game strategies.
[0059] Further, the construction of the comprehensive fuzzy benefit function:
[0060] (1) First, sort the three fuzzy values to get the descending triplet μ(1)≥μ(2)≥μ(3);
[0061] (μ (1) , μ (2) , μ (3) ):=sort_desc(μ trust , μ delay , μ resource ) (13)
[0062] (2) Potential game modeling
[0063] Let OWA weight w={w1,w2,w3}∈Δ 2 represent the action of the "central dispatcher", and let the joint strategy π of the UAV / edge node represent the action of the subordinate participants;
[0064] Define the potential benefit function Φ(w,π) as the potential function of the game;
[0065]
[0066] where w prior is the prior knowledge of the governance layer, and λ > 0 is the regularization coefficient;
[0067] The Lemma A.1 in the Appendix proves that the weight updating game and the UAV / edge node strategy game share the potential function Φ, thus forming a single-person potential game;
[0068] (3) the central weight is updated by using the regret learning algorithm to perform weighted regret update;
[0069]
[0070] where J is the learning rate, and μ(τ, k) represents the kth largest membership degree value after sorting at step τ; specifically, μ(τ, k) is the hth value after sorting the three membership degrees in descending order at time τ; according to the regret learning theory, the external regret of the sequence wt is RT / T→0; combined with the potential game characteristics, it is obtained that:
[0071] Theorem 1 For any ∈ > 0, there is a step number T(∈) = O(1 / ∈ 2 ) such that the average weight forms an ∈-Nash equilibrium;
[0072] In the scheduling formula of, the interaction between UAVs can be modeled as a finite potential game, where each UAV is a participant, and its strategy corresponds to the relative contribution to selecting different fuzzy aggregation components; if there is a scalar function Φ: S→R such that for any participant i, strategies si, s'i∈Si and s -i ∈S -i ,
[0073] u i (s′ i , s -i )-u i (s i , s -i )=Φ(s′ i , s -i )-Φ(s i , s -i ) (16)
[0074] The game G = <N, {Si}, {ui}> is called a potential game: this property ensures that any unilateral improvement in the utility of any UAV is consistent with an increase in the global potential function, which means that local optimization is coordinated with global network performance improvement; in the case of, the regret learning weight update in formula (14) is based on this potential game structure, so that the learning dynamics can converge to an ∈-Nash equilibrium that maximizes the potential function; this directly links the local adaptability adjustment of the aggregation weight to the maximization of the overall system efficiency;
[0075] (4) OWA aggregation: the integrated return at time t is calculated as follows:
[0076]
[0077] where, controlling the degree of emphasis on the best indicator, controlling the degree of punishment on the worst indicator; regret learning automatically increases in congestion or low energy scenarios (risk aversion); when resources are abundant, turn to average or bias towards maximum (pursue efficiency).
[0078] Another object of the present application is to provide a model proof and verification system for a two-stage T-norm-Choquet-OWA resource aggregator oriented to multi-UAV cooperation, comprising:
[0079] a prediction module for predicting the resource trajectory for the next 3 seconds using a long short-term memory network-exponential moving average (LSTM-EMA) model;
[0080] a compensation module for determining the bottleneck resource in the first stage using T-norm (taking the minimum value), and in the second stage using a Choquet-OWA method driven by an adaptive interaction measure φ to perform flexible compensation according to the instantaneous power usage, realizing the cooperative strategy of "first solving the bottleneck, then restoring efficiency".
[0081] Another object of the present application is to provide a computer device comprising a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to make the processor execute the steps of the model proof and verification method for a two-stage T-norm-Choquet-OWA resource aggregator oriented to multi-UAV cooperation.
[0082] Another object of the present application is to provide a computer-readable storage medium storing a computer program, the computer program being executed by a processor to make the processor execute the steps of the model proof and verification method for a two-stage T-norm-Choquet-OWA resource aggregator oriented to multi-UAV cooperation.
[0083] Another object of the present application is to provide an information data processing terminal for implementing the model proof and verification system for a two-stage T-norm-Choquet-OWA resource aggregator oriented to multi-UAV cooperation.
[0084] In combination with the above technical solutions and solved technical problems, the technical solution to be protected by the present application has the following advantages and positive effects:
[0085] The application proposes a two-stage T-norm-Joiner-OWA resource aggregator based on prediction, the core advantage of which is that it can capture the dynamic resources of the UAV group in computing, communication and energy at the same time, and realize low delay and high reliability scheduling through the prediction mechanism and the two-stage aggregation strategy. Compared with the traditional static method, this scheme shows stronger adaptability and robustness under short-term resource fluctuations and extreme bottleneck conditions, effectively improving the resource utilization efficiency and execution stability in multi-UAV cooperative tasks.
[0086] In terms of specific mechanisms, the application realizes the forward-looking prediction of the future three-second resource trajectory through the long short-term memory network (LSTM) combined with the exponential moving average (EMA) model. This design breaks through the limitations of traditional schemes that rely only on current state decision-making, and can identify the upcoming computing load peaks, communication congestion and energy decline trends in advance, thereby actively avoiding and reserving resources at the scheduling level, significantly reducing the risk of task failure due to lag response.
[0087] In terms of aggregation strategy, the application adopts a two-stage structure, the first stage uses the minimum value operation of T-norm to strictly guarantee the "bottleneck effect", ensuring that any single resource shortage can be identified and trigger the protection mechanism in time; the second stage introduces the Joiner-OWA operator driven by the adaptive interaction metric φ, which elastically couples the power consumption rate and resource sufficiency, thereby achieving a balance between performance priority and energy efficiency priority. This design embodies the core idea of "protecting the bottleneck first, then restoring efficiency", ensuring the safety and flexibility of the system.
[0088] The application not only has rigorous theoretical support in method design. Through strict proof of the monotonicity, boundary correctness, bottleneck priority and Lyapunov stability of the aggregator, the explainability, robustness and convergence of the model in the running process are ensured. This theoretical basis provides a reliable guarantee for its industrial application in complex flight environments and large-scale UAV clusters.
[0089] In terms of performance verification, the application conducts system evaluation through joint simulation experiments involving 360 UAVs. The results show that this method controls the average round-trip time (RTT) to within 55 milliseconds, reducing the delay by 5%, 10%, 15% and 20% respectively compared with the minimum value scheduler, DRL-PPO algorithm, single-layer OWA method and weighted sum (WSM) method. At the same time, the jitter is controlled within 11 milliseconds, the packet loss rate is less than 0.03%, and the remaining power is increased by about 12%. These data fully prove that this method has low delay, high stability and excellent energy efficiency performance in large-scale UAV groups.
[0090] The two-stage T-norm-Choquet-OWA resource aggregator based on prediction proposed in the application has significant advantages in theory and experiment. Through the combination of future perception and fuzzy aggregation, it successfully solves the bottleneck resource protection and cross-dimension compensation problem of UAV swarm in complex task scenarios, realizes the safety, efficiency and robustness of task execution, and provides a reliable technical path and industrial value for the application of UAV cluster in high load, long endurance and complex communication environment. BRIEF DESCRIPTION OF DRAWINGS
[0091] Figure 1 is the model proof and verification method flowchart of the two-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation provided by the embodiment of the application.
[0092] Figure 2 is the model proof and verification system structure block diagram of the two-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation provided by the embodiment of the application.
[0093] Figure 3 is the overall data flow and functional module diagram of the aggregator provided by the embodiment of the application.
[0094] Figure 4 is the resource membership degree distribution diagram of four resource aggregation strategies in three task scenarios provided by the embodiment of the application.
[0095] Figure 5 is the available bandwidth (left) and packet loss rate (right) diagram of the UAV cluster provided by the embodiment of the application.
[0096] Figure 6 is the round trip time (left) and signal strength (right) distribution diagram provided by the embodiment of the application.
[0097] Figure 7 is the link jitter node distribution diagram in the UAV scenario provided by the embodiment of the application.
[0098] Figure 8 is the round trip time comparison diagram of five scheduling strategies provided by the embodiment of the application.
[0099] Figure 9 is the performance diagram (mean ± standard deviation of 5 random seed tests) under different UAV scales provided by the embodiment of the application.
[0100] Figure 10 is the ablation experiment result comparison diagram provided by the embodiment of the application. DETAILED DESCRIPTION
[0101] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0102] The core technical problem to be solved by the present application comes from the contradiction between real-time aggregation of heterogeneous resources and prediction of failure risk in a multi-unmanned aerial vehicle cooperative environment. In current industrial applications, the unmanned aerial vehicle performs tasks often involving three key resources: computing (CPU), communication (bandwidth), and energy. However, traditional resource management relies on static thresholds or average scheduling, which cannot model future short-term resource fluctuations, extreme bottleneck effects, and power responses synchronously, resulting in delays, packet loss, or even task failure when sudden loads or energy drops occur. The present application introduces a long short-term memory network combined with an exponential moving average method at the prediction layer to estimate resource trajectories 3 seconds in advance, so that the decision is no longer a "lagging response", but a "forward-looking avoidance". This is a substantial improvement over the single real-time monitoring mode in the prior art.
[0103] In terms of working principle, the present application first establishes a multi-dimensional fuzzy set, which normalizes the CPU utilization, bandwidth occupancy, and remaining energy and introduces them into the fuzzy membership degree framework. This solves the industrial pain point that heterogeneous resources are difficult to compare directly. Further, by using a future perception membership degree function, the prediction results are injected into the membership degree curve. For example, the CPU and bandwidth trigger penalties in advance under the upcoming peak load, and the energy function deducts the predicted consumption in advance to achieve "prediction contraction". This mechanism significantly reduces the problem of incorrect resource allocation due to insufficient prediction, making the scheduling system more suitable for the dynamic environment of real flight tasks.
[0104] Secondly, the T-norm minimum operation is introduced in the first stage aggregation, which strictly preserves the "bottleneck effect", that is, any extreme shortage of a single resource will directly become a system rejection or migration signal. This extreme protection mechanism can avoid the collapse of a single resource leading to the failure of the entire task in industrial applications, ensuring a minimum level of safety. Compared with the existing average allocation or weighted sum method, it is more in line with the actual needs of emergency scenarios.
[0105] In the second stage, the present application uses the Choquet-OWA operator combined with an adaptive interaction measure φ to model the power consumption rate and resource sufficiency, forming an elastic adjustment. The core principle is that when the task performance requirement is high and the energy is sufficient, the φ value increases, allowing the system to give priority to performance; conversely, when the task prioritizes energy saving or energy is tight, the φ value decreases, strengthening the penalty on power. This coupling method is significantly superior to the fixed weight model in industrial applications, enabling dynamic adaptation across task categories, making the scheduling strategy flexible and migratory.
[0106] At the income modeling level, the application introduces a potential game framework, maps the fuzzy output of resource aggregation and the weight decision of the central scheduler to a unified potential function Φ, and realizes weight update through regret learning. This forms a dynamic game balance between the central governance layer and the individual UAVs, avoiding the problems of "global rigidity" in traditional centralized scheduling or "local greed" in distributed game. The working principle ensures that the system tends to be globally stable and efficient in the long-term evolution, which is extremely groundbreaking in the industrial application of multi-UAV cluster.
[0107] The method solves three long-standing technical problems in industrial applications: first, the advance response to short-term future resource fluctuations enables the UAV cluster to have forward-looking scheduling capabilities; second, a strict bottleneck protection mechanism reduces the risk of task failure; and third, an elastic power coupling mechanism achieves dynamic balance between performance and energy efficiency under different task preferences. Combined with potential game and learning optimization, the resource aggregation model is finally verified and adaptively evolved. These mechanisms together form a complete link from prediction, aggregation to income feedback, significantly improving the task completion rate, energy efficiency, and robustness of the system in the industrial application of multi-UAV cooperative tasks.
[0108] As shown in Figure 1 The model proof and verification method of the double-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation provided by the embodiment of the application includes the following steps:
[0109] S101, using a long short-term memory network-exponential moving average (LSTM-EMA) model to predict the resource trajectory for the next 3 seconds;
[0110] S102, the first stage uses T-norm (minimum value) to determine the bottleneck resource, and the second stage uses the Choquet-OWA method driven by the adaptive interaction metric φ to perform elastic compensation according to the instantaneous power usage, realizing the cooperative strategy of "first solving the bottleneck, then restoring the efficiency".
[0111] The embodiment of the application provides a long short-term memory network-exponential moving average (LSTM-EMA) model for predicting the resource trajectory for the next 3 seconds:
[0112] (1) Resource use multi-dimensional fuzzy set design;
[0113] (2) Construction of a comprehensive fuzzy income function.
[0114] The resource use multi-dimensional fuzzy set design provided by the embodiment of the application is:
[0115] 1) Fuzzy aggregation under extreme protection
[0116] Let and eraw respectively represent the CPU utilization drone synergy, bandwidth utilization drone synergy and residual energy drone synergy measured at the drone node; to make them dimensionless and comparable, each index is normalized to the unit interval [0, 1] before entering the fuzzy set model, obtaining:
[0117]
[0118] where, e min , e max is the observed value or the nominal bound of each resource (typically 0 and 100%); after normalization, all three indices satisfy F {·} ∈ [0, 1], where a value close to 1 indicates a high utilization of CPU and bandwidth (for CPU and bandwidth) or an energy sufficiency (for energy);
[0119] This min-max normalization rescales each raw index into a standard dimensionless form, enabling a fair aggregation of heterogeneous resources in the multi-dimensional fuzzy set framework; Fcand Fbtend to 1 when CPU or bandwidth is heavily used, while Fe tends to 1 when residual energy is sufficient;
[0120] 2) Future-aware membership functions
[0121] CPU prediction-enhanced membership function (alarm contraction): one step penalty on "current load + upcoming predicted load" to avoid resource decision errors due to upcoming peaks;
[0122]
[0123] Here, c is the current CPU utilization; Fcis the k-seconds-ahead predicted utilization; their sum measures the upcoming total pressure on CPU; λc∈ [0, 1] controls the weight of the prediction - the more sensitive the task to computation, the larger λc; θcis the utilization inflection point; κccontrols the steepness of the S-shaped curve; if Ifc+ λ c F c <<θ c , the exponential term tends to 0, indicating that the computational resource is sufficient; otherwise, it rapidly decreases;
[0124] Bandwidth membership function (alarm contraction): for high real-time communication services, a small θband a large κbcan be set to achieve an "early braking";
[0125]
[0126] Here, b is the current bandwidth utilization; F b is the predicted future congestion; λ bis the weight; if b + λ b F b very small (bandwidth sufficient); once the inflection point θ b , it decays rapidly;
[0127] Energy membership function (power response correction):
[0128]
[0129] Here, is the original trigonometric function, and the offset term realizes the "future power drop" penalty; er is the current remaining energy ratio; Fe is the predicted remaining energy ratio after the task is completed; the offset term λ e (1-F e ) deducts the expected energy consumption in advance, and λe reflects the sensitivity of the task to energy; after the offset, if the predicted remaining energy sharply decreases, the function input becomes smaller, and the energy protection triggers faster;
[0130] Instantaneous power consumption membership (linear):
[0131] μ p (p) = 1-p, p∈[0,1] (5)
[0132] Where p is the ratio of the current power load to the reference power; it decreases linearly: the higher the power → the lower the membership, which simply and intuitively quantifies the influence of "power consumption" on the feasibility of the task;
[0133] 3) Two-stage extreme protection + elastic coupling;
[0134] 4) Two-stage T-norm - Jouquet resource aggregation.
[0135] The two-stage extreme protection + elastic coupling provided by the embodiments of the present application:
[0136] First-stage extreme protection (core three dimensions): retain the "bucket effect", the weakest one of the three main resources determines the overall feasibility; if any membership is equal to 0 → μ pre = 0, the task is immediately rejected or migrated;
[0137]
[0138] Second-stage elastic coupling:
[0139] φ = η + (1-η)F e , φ∈(0,1] (7)
[0140] Here, η is the task elasticity; for high performance tasks, η→1, indicating more emphasis on performance; Fe is the predicted residual energy after the task is completed; as the energy sufficiency increases, Fe→1; φ combines the two factors: if the task has high performance requirements and energy sufficiency increases; if the task is energy-saving or energy is low decreases; φ adjusts the strength of the penalty on the power rate.
[0141] The two-stage T-norm-Joiner resource aggregation provided by the embodiment of the application:
[0142] In order to capture the bottleneck protection and cross-dimension complementarity at the same time, the application adopts a two-stage aggregation structure; in the lower stage, the T-norm (minimum value) is used to extract the bottleneck value from the three main resources: CPU, bandwidth and energy;
[0143]
[0144] In the higher stage, the binary Joiner-OWA operator with interaction measure ν is used to model the complementarity between resource sufficiency and power rate in μ pre , μ p , and is defined as follows:
[0145] f1=μ pre ,f2=μ p , f (1) ≥f (2) (9)
[0146] Sort the two membership degrees in descending order, set S1={1,2}, S2={2}; then the Joiner aggregation is:
[0147]
[0148] Here, the interaction measure is defined as:
[0149] ν({μ pre})=φ,ν({μ p})=1-φ,ν({μ pre ,μ p})=1 (11)
[0150] And φ=η+(1-η)Fe∈(0,1], consistent with the original model; substituting the above formula gives the closed form:
[0151] μ resource =φμ pre +(1-φ)max{μ pre ,μ p} (12)
[0152] Here, when μpre ≤μ p (when power consumption rate is better than bottleneck value), μ resource = φμ pre + (1-φ)μ p reflects the "power compensation" effect; when μpre> μp, it reverts to μresource= μpre to ensure bottleneck protection priority; parameter φ is still modulated by task elasticity η and predicted residual energy Fe to achieve scenario adaptive behavior;
[0153] The overall mechanism is as follows:
[0154] Future perception: formula Injecting the predicted value into the membership function can pre-punish the upcoming resource conflict or energy decline;
[0155] Strict bottleneck protection: formula μpre uses the minimum operator to ensure that any major resource shortage immediately triggers protection;
[0156] Flexible rate suppression: the formula of φ and μresource incorporates instantaneous power consumption, smoothly fuses power mean and power penalty term, and realizes elasticity-power coupling adjustment;
[0157] Final membership output: the membership μresource combines with μtrust and μdelay to form the benefit R, which drives the adaptive evolution of the subsequent fuzzy game strategy.
[0158] The construction of the comprehensive fuzzy benefit function provided by the embodiment of the application is as follows:
[0159] (1) First, sort the three fuzzy values to obtain a descending three-tuple μ(1)≥μ(2)≥μ(3);
[0160] (μ (1) ,μ (2) ,μ (3) ):=sort_desc(μ trust ,μ delay ,μ resource ) (13)
[0161] (2) Potential game modeling
[0162] Let OWA weight w={w1,w2,w3}∈Δ 2 represent the action of the "central dispatcher", and let the joint strategy π of the UAV / edge node represent the action of the subordinate participant;
[0163] Define the potential benefit function Φ(w,π) as the potential function of the game;
[0164]
[0165] where w prior is the prior knowledge of the governance layer, and λ > 0 is the regularization coefficient;
[0166] The Lemma A.1 in the Appendix proves that the weight updating game and the UAV / edge node strategy game share the potential function Φ, thus forming a single-person potential game;
[0167] (3) The central weight is updated by using the regret learning algorithm to perform weighted regret update;
[0168]
[0169] where J is the learning rate, and μ(τ, k) represents the kth largest membership value after ordering at step τ; specifically, μ(τ, k) is the hth value after descending ordering of the three membership degrees at time τ; according to the regret learning theory, the external regret of the sequence wt is RT / T→0; combined with the potential game property, it is obtained that
[0170] Theorem 1 For any ∈ > 0, there exists a step number T(∈) = O(1 / ∈ 2 ) such that the average weight forms an ∈-Nash equilibrium;
[0171] In the scheduling formula of, the interaction between UAVs can be modeled as a finite potential game, where each UAV is a participant, and its strategy corresponds to the relative contribution to the selection of different fuzzy aggregation components; if there exists a scalar function Φ: S→R such that for any participant i, strategies si, s'i∈ Si, and s -i ∈ S -i ,
[0172] u i (s′ i ,s -i )-u i (s i ,s -i )=Φ(s′ i ,s -i )-Φ(s i ,s -i ) (16)
[0173] The game G = <N, {Si}, {ui}> is called a potential game: this property ensures that any unilateral improvement of the utility of any UAV is consistent with the increase of the global potential function, which means that local optimization is coordinated with the improvement of the global network performance; in the case of, the regret learning weight update in formula (14) is based on this potential game structure, so that the learning dynamics can converge to an ∈-Nash equilibrium that maximizes the potential function; this directly links the local adaptability adjustment of the aggregation weight to the maximization of the overall system efficiency;
[0174] (4) OWA aggregation: the integrated return at time t is calculated as follows:
[0175]
[0176] where, control the degree of emphasis on the best indicator, control the degree of punishment for the worst indicator; regret learning automatically increases in congestion or low energy scenarios (risk aversion); when resources are abundant, turn to the average or favor the maximum (pursue efficiency).
[0177] As Figure 2 shown, the model proving and verifying system of the double-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation provided by the embodiment of the present application comprises:
[0178] a prediction module, configured to predict resource trajectories in the future 3 seconds by using a long short-term memory network-exponential moving average (LSTM-EMA) model;
[0179] a compensation module, configured to determine bottleneck resources by using T-norm (taking the minimum value) in the first stage, and to perform elastic compensation according to instantaneous power usage by using a Choquet-OWA method driven by an adaptive interaction metric φ in the second stage, so as to realize the cooperative strategy of “first solving the bottleneck, and then restoring efficiency”.
[0180] Another object of the present application is to provide a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the model proving and verifying method of the double-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation.
[0181] Another object of the present application is to provide a computer readable storage medium storing a computer program, and the computer program is executed by a processor to make the processor execute the steps of the model proving and verifying method of the double-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation.
[0182] Another object of the present application is to provide an information data processing terminal for realizing the model proving and verifying system of the double-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation.
[0183] The present application is embodied as follows:
[0184] 1. The main contributions of the present application are as follows:
[0185] Predictive two-stage T-norm-Jouquette-OWA aggregator: A resource aggregator is introduced to optimize the battery energy, bandwidth, and onboard computing capacity of a UAV swarm in real time. A predictive enhanced membership function can predict resource dynamics several seconds in advance, and a two-stage design first protects bottleneck resources using a T-norm and then fuses non-bottleneck capacities using a Jouquette-OWA. This ensures the safety and efficiency of task execution in resource-limited scenarios.
[0186] Rigorous theoretical foundation: Comprehensive theoretical analysis and parameter design are conducted, including monotonicity proof, boundary correctness proof, bottleneck priority proof, and Lyapunov stability proof. These results guarantee the explainability, robustness, and convergence of the aggregator in practical applications.
[0187] Simulation-based verification: Extensive simulation studies show that the aggregator outperforms traditional bottleneck protection methods based on minimum values and linear weighted WSM methods, with better aggregation performance and higher resource saving efficiency.
[0188] The rest of the invention is arranged as follows: Section 2 reviews related research work; Section 3 details the theoretical properties of the proposed aggregator model and algorithm; Section 4 provides proofs and analysis of the algorithm; Section 5 conducts simulation experiments and performance analysis; Section 6 discusses the work of the invention; finally, Section 7 summarizes the full text and looks forward to future research directions.
[0189] 2. Related work
[0190] 2.1 Overview of UAV resource scheduling and task allocation methods
[0191] Unmanned aerial vehicles (UAVs) play an indispensable role in numerous modern applications, which has led to a demand for efficient resource scheduling and task allocation methods. In flying ad hoc networks (FanETS), a hierarchical framework clusters user UAVs, enabling tasks to be performed locally or offloaded to mobile edge computing (MEC) UAVs, and minimizes energy consumption through an iterative optimization algorithm [5]. Mobile crowd sensing (MCS) with UAV assistance employs a multi-task allocation scheme and deep reinforcement learning to expand data collection range, optimize flight paths, and reduce energy consumption [6]. In the context of emergency response, a swarm-level scheduling method applies a particle swarm optimization (PSO) algorithm to balance and minimize total flight distance, outperforming traditional methods [7]. In hybrid mobile edge computing networks, thermal-aware scheduling addresses central processing unit (CPU) cooling limitations by jointly optimizing user access, task scheduling, and UAV trajectories, thereby reducing task time while controlling CPU temperature [8]. Decentralized swarm scheduling methods, such as the consensus-based bundle algorithm (CBBA) and performance impact (PI) algorithm, integrate task-aware functionality, improving task allocation accuracy and reducing flight time compared to traditional methods [9]. A biologically inspired wolf pack strategy models UAV swarm behavior in complex environments to achieve dynamic task allocation, achieving high task completion rates and balanced workloads
[10] . FlexEdge utilizes genetic algorithms to convert multi-objective scheduling into a single-objective problem, optimizing task allocation and UAV positioning to reduce execution time and energy consumption
[11] . Finally, the maximum UAV trajectory and task allocation algorithm (MUTAA) provides real-time route planning and scheduling in delay-sensitive scenarios, significantly improving task completion rates
[12] . Overall, these methods highlight the challenges and innovations in the field of UAV resource scheduling and task allocation, particularly in improving energy efficiency, thermal management, and real-time decision-making capabilities.
[0192] 2.2 Fuzzy membership functions and T-norm / Ordered Weighted Averaging (OWA) aggregation
[0193] Fuzzy membership functions and T-norms / OWA aggregation are the foundation of fuzzy logic and decision making, especially in the context of Multi-Attribute Decision Making (MADM) and Multi-Criteria Decision Making (MCDM). Membership functions extend the membership relation of classical sets to more flexibly model uncertainty and fuzziness. For example, in opportunistic mobile networks, an optimal membership function based on asymmetric triangular fuzzy numbers improves routing metrics that outperform symmetric fuzzy numbers in terms of both transmission cost and delay
[13] . In rule-based classifiers, membership functions improve the explainability and reliability of decision support systems, balancing the succinctness of the knowledge base with the quality of generalization
[14] .
[0194] T-norms and OWA operators are essential for aggregating fuzzy information. T-norms like the Aczel-Alsina norm provide versatile means for combining fuzzy sets, benefiting MADM scenarios that require fusing uncertain inputs. Applications include interval-valued Pythagorean fuzzy sets and T-spherical fuzzy data, where the Aczel-Alsina aggregation reduces information loss and enhances the robustness of decisions [15, 16]. The use of generalized T-norms and T-copulas in advanced fuzzy aggregation operators further demonstrates their adaptability in complex decision environments
[17] . In the context of q-order orthomodular fuzzy environments, the integration of T-norms creates a flexible, robust framework for handling unknown weight information
[18] . Together, these techniques form the basis of sophisticated models capable of addressing the uncertainty and complexity of real-world decision problems.
[0195] 2.3 Choquet Integrals and Bottleneck Protection in UAV Networks
[0196] The combination of the Quetelet metric with bottleneck protection strategies has recently emerged as a promising approach to improve the efficiency and security of UAV networks. Quetelet integrals, as a powerful tool, can capture the interdependencies between performance criteria and optimize resource allocation and decision-making under conflicting objectives such as energy usage, bandwidth allocation, and computational load. As UAV systems become flexible platforms for wireless communication and edge computing, effective management of these resources is crucial to maintain optimal performance
[19] . Bottleneck protection techniques like Software Defined Network (SDN)-driven topology deception play a vital role in protecting critical UAV nodes. By generating virtual network layouts, these schemes can mislead adversaries and protect UAVs acting as communication relays in sensor-assisted deployments
[20] . However, UAVs still face severe limitations: limited onboard energy, narrow wireless channels, and limited processing power. To address these issues, researchers have explored the integration of mobile edge computing and blockchain, providing a secure framework for task offloading and resource coordination, which is a critical step in protecting privacy and reducing power consumption
[21] . However, the practicality of UAVs is hindered by their inherent Size-Weight-Power (SWaP) limitations
[22] .
[0197] Optimizing UAV placement and control to maximize quality of service while minimizing energy consumption remains a challenge. Recent algorithms that utilize Virtual Force Fields (VFF) and Optimal Transport Theory (OTT) for resource management have shown promising results, but also highlight the complexity of the problem
[23] . Overall, these advances indicate that further research is needed to overcome current limitations and fully unleash the potential of UAV networks in various applications.
[0198] 2.4 Comparison and limitations of two-stage and multi-stage aggregation frameworks
[0199] In complex environments, multi-UAV systems are now tasked with increasingly diverse and dynamic missions. To cope with this complexity, researchers have designed hierarchical frameworks that separate resource aggregation, path planning, and task scheduling into two or more stages. Each stage employs a specialized aggregation strategy, which simplifies computation and improves adaptability. However, these methods often create information silos between stages and can only achieve limited global optimality.
[0200] In practice, staged aggregation is common in task allocation and route optimization, such as coordinating data collection, energy management, and real-time communication. While such frameworks can improve resource utilization and alleviate real-time processing burden, researchers have found that stage coupling and error propagation often lead to suboptimal end-to-end performance. As shown in Table 1, current methods still have deficiencies in multi-objective trade-offs and overall system robustness when balancing information freshness, energy efficiency, and the ability to adapt to changing conditions.
[0201] Table 1. Comparison of existing two-stage / multi-stage resource aggregation frameworks
[0202]
[0203]
[0204] 2.5 Research based on auction and hybrid reinforcement learning-fuzzy methods
[0205] Recently, in the study of task scheduling in the context of UAV / edge computing, several studies have combined reinforcement learning with fuzzy logic or auction mechanisms. Specifically:
[0206] Li Dong et al.
[27] proposed a deep progressive reinforcement learning scheduling framework for intelligent reflector-assisted UAV–MEC systems. In this framework, a progressive scheduler and a tabu search algorithm jointly optimize the positioning, task offloading, and resource allocation of UAVs, demonstrating real-time scheduling capabilities and robustness against catastrophic forgetting (arXiv).
[0207] He et al.
[28] developed an edge computing framework for intelligent agricultural supply chains that combines an auction mechanism with a fuzzy optimizer. Through multi-stage auctions and fuzzy neural networks, the framework supports coordination and scheduling between stages of the agricultural supply chain, emphasizing the combination of market mechanisms and rule-based fuzzy optimization to enable real-time decision-making in complex agricultural scenarios (SpringerOpen).
[0208] Zander et al.
[29] studied the integration of reinforcement learning with TSK fuzzy systems, exploring the application of architectures such as actor-critic and DQN-ANFIS in standard reinforcement learning tasks, highlighting the potential of reinforcement learning-fuzzy systems in terms of interpretability and performance (ResearchGate).
[0209] Differences and innovations from the above research
[0210] Differences in fusion mechanisms: Unlike Li et al. who focus on the reinforcement learning structure evolution, He et al. who emphasize auction and fuzzy neural network-driven real-time decision making, and Zand et al. who focus on reinforcement learning-fuzzy control, the proposed two-stage scheduler explicitly integrates heterogeneous fuzzy aggregation operators - T-norm, Choquet, and ordered weighted average (OWA) operators - with a long short-term memory network-exponential moving average (LSTM-EMA) prediction layer with a 3-second look-ahead window. This design enables multi-dimensional fusion of scheduling evaluation metrics and prediction-based weight self-adaptation, while auction / reinforcement learning-based approaches mainly rely on post-mortem scheduling.
[0211] Differences in theory and interpretability: The proposed framework provides rigorous proofs on monotonicity, boundary correctness, and Lyapunov stability while maintaining O(1) computational complexity per node, ensuring interpretability and embedded deployability. In contrast, existing hybrid reinforcement learning-fuzzy and auction-based schedulers often lack such formal guarantees.
[0212] Predictive two-stage T-norm–Choquet–OWA resource aggregator
[0213] To meet the real-time coordination needs of UAV swarms in terms of battery energy, bandwidth, and on-board computing capacity, a predictive two-stage T-norm–Choquet–OWA resource aggregator is proposed. First, it defines a fuzzy set representing the multi-dimensional resource occupation and uses prediction-enhanced membership functions to predict and adjust upcoming loads. Next, a two-stage "strict protection + flexible integration" design is adopted to ensure robust performance:
[0214] First stage (T-norm / min): Accurately isolates bottleneck resources, preventing any "short board" from being overlooked.
[0215] Second stage (Choquet-OWA): Self-adapts the trade-off between identified bottleneck levels and instantaneous power consumption, achieving a smooth balance between system performance and endurance.
[0216] Figure 3 The aggregator is demonstrated in five stages of data flow and modules, from real-time monitoring and future prediction to the final membership output:
[0217] Data acquisition layer: Embedded sensors capture current loads, and edge prediction modules provide s-second look-ahead predictions.
[0218] Fuzzification layer: Four prediction-enhanced membership functions convert each resource dimension into fuzzy membership values.
[0219] Bottleneck protection layer (first stage): The minimum of these membership values is computed to determine the bottleneck membership μb. If any primary resource is below its threshold, protection is immediately triggered.
[0220] Elastic fusion layer (second stage): The coupling factor λ is computed from the task elasticity e and the predicted remaining energy Erem. Then, the μb is fused with the elastic membership μe using the Jocchite-OWA to generate the final membership μout.
[0221] Scheduler interface layer: The scheduler uses μout to evaluate task feasibility and determine priority order.
[0222] 3.1 Resource usage multi-dimensional fuzzy set design (μresource)
[0223] 3.1.1 Fuzzy aggregation under extreme protection
[0224] To ensure safe and stable execution of tasks in the blockchain-enabled UAV network fuzzy game model, an extreme protection strategy for multi-dimensional resource fuzzy aggregation is adopted. Future perception and elastic coupling enhancement are performed on this method: (1) Extreme protection isolates the bottleneck by selecting the minimum membership value in the three dimensions—central processing unit (CPU), network bandwidth, and battery capacity—as μresource. (2) If any dimension is below its threshold, the task is immediately rejected or the game strategy is adjusted to prevent excessive consumption and system instability ("short board" effect). Although this method is more conservative than average or weighted aggregation, it ensures high reliability, which is crucial for safety-sensitive tasks.
[0225] Table 2 shows the resource usage multi-dimensional fuzzy set indicators.
[0226]
[0227] As shown in Table 2, let and e raw represent the CPU utilization UAV coordination, bandwidth utilization UAV coordination, and remaining energy UAV coordination measured at the UAV node, respectively. To make them dimensionless and comparable, each indicator is normalized to the unit interval [0, 1] before entering the fuzzy set model, resulting in:
[0228]
[0229] where, e min ,e max is the observed value or nominal boundary for each resource (typically 0 and 100%). After normalization, all three indicators satisfy F {·}∈ [0, 1], where a value close to 1 indicates high utilization of CPU and bandwidth (for CPU and bandwidth) or energy sufficiency (for energy).
[0230] This min-max normalization rescales each raw indicator into a standard dimensionless form, enabling fair aggregation of heterogeneous resources in the multi-dimensional fuzzy set framework. Fcand Fbtend to 1 when CPU or bandwidth is heavily used, while Feis close to 1 when there is sufficient energy left.
[0231] 3.1.2 Future-aware membership functions
[0232] (1) CPU prediction-enhanced membership function (alert contraction): One step penalty on "current load + upcoming predicted load" to avoid resource decision errors due to upcoming peaks.
[0233]
[0234] Here, c is the current CPU utilization; Fcis the k-seconds-ahead predicted utilization; their sum measures the upcoming total pressure on CPU. λc∈ [0, 1] controls the weight of prediction - the more sensitive the task is to computation, the larger λc. θcis the utilization knee; κccontrols the steepness of the S-shaped curve. If Ifc+ λ c F c << θ c , the exponential term tends to 0, indicating that computation resources are sufficient; otherwise, it rapidly decreases.
[0235] (2) Bandwidth membership function (alert contraction): For high real-time communication services, a small θband a large κbcan be set to achieve "early braking".
[0236]
[0237] Here, b is the current bandwidth utilization; Fbis the predicted future congestion; λbis the weight. If b + λ b F b is very small (bandwidth is sufficient); once the knee θbis reached, it rapidly decays.
[0238] (3) Energy membership function (power response correction):
[0239]
[0240] Here, is the original trigonometric function, and the offset term implements a "future power drop" penalty. eris the current residual energy proportion; Feis the predicted residual energy proportion after the task is completed. The offset term λ e (1 - Fe ) Pre-deduct the expected consumed energy, λe reflects the sensitivity of the task to energy. After offset, if the predicted remaining energy drops sharply, the function input becomes smaller, the membership decreases, and the energy protection triggers faster.
[0241] (4) Instantaneous power consumption membership (linear):
[0242] μ p (p) = 1 - p, p ∈ [0, 1] (5)
[0243] Where p is the proportion of the current power load to the reference power. It decreases linearly: the higher the power → the lower the membership, simply and intuitively quantifying the impact of "power consumption" on the feasibility of the task.
[0244] 3.1.3 Two-stage extreme protection + elastic coupling
[0245] (1) First stage extreme protection (core three dimensions): retain the "bucket effect", the weakest of the three main resources determines the overall feasibility. If any membership is equal to 0 → μ pre = 0, the task is immediately rejected or migrated.
[0246]
[0247] (2) Second stage elastic coupling:
[0248] φ = η + (1 - η) F e , φ ∈ (0, 1] (7)
[0249] Here, η is the task elasticity; for high-performance tasks, η → 1, indicating more emphasis on performance. Fe is the predicted remaining energy after the task is completed; as the energy sufficiency increases, Fe → 1. φ combines these two factors: if the task has high performance requirements and energy is sufficient Increase; if the task is energy-saving or energy is low Decrease. φ adjusts the strength of the punishment on power consumption.
[0250] (3) Two-stage T-norm - Choquet resource aggregation
[0251] To capture both bottleneck protection and cross-dimensional complementarity, the invention adopts a two-stage aggregation structure. In the lower stage, the T-norm (minimum value) is used to extract the bottleneck value from the three main resources - CPU, bandwidth and energy.
[0252]
[0253] In the higher stage, the binary Choquet-OWA operator with interaction measure v is used to aggregate resource sufficiency and power consumption in μ pre , μp The complementarity between f1and f2is modeled, defined as follows:
[0254] f1= μ pre f2= μ p (1) ≥ f (2) (9)
[0255] Sort the two membership degrees in descending order, set S1 = {1, 2}, S2 = {2}. Then the Jaccard aggregation is:
[0256]
[0257] Here, the interaction metric is defined as:
[0258] v({μ pre}) = φ, v({μ p}) = 1 - φ, v({μ pre , μ p}) = 1 (11)
[0259] and φ = η + (1 - η) Fe ∈ (0, 1], consistent with the original model. Substituting the above formula gives a closed form:
[0260] μ resource = φμ pre + (1 - φ) max{μ pre , μ p} (12)
[0261] Here, when μ pre ≤ μ p (power consumption is better than the bottleneck value), μ resource = φμ pre + (1 - φ) μ p reflects the "power compensation" effect; when μpre> μp, it returns to μresource= μpreto ensure that bottleneck protection takes priority. The parameter φ is still modulated by the task elasticity η and the predicted residual energy Fe, achieving scenario adaptive behavior.
[0262] The overall mechanism is as follows:
[0263] Future awareness: formula Injecting the predicted value into the membership function can pre-punish the upcoming resource conflict or energy decline.
[0264] Strict bottleneck protection: formula μpreuses the minimum operator to ensure that any major resource shortage triggers protection immediately.
[0265] Flexible rate throttling: The formula of φ and μresource incorporates the instantaneous power consumption, which smoothly combines the power mean and power penalty term, to achieve the elasticity-power coupling adjustment.
[0266] Final membership output: The membership μresource combines with μtrust and μdelay to form the payoff R, which drives the adaptive evolution of the subsequent fuzzy game strategy.
[0267] 3.2 Construction of the comprehensive fuzzy payoff function
[0268] In this section, after obtaining the single resource evaluation value μresource, further consideration is given to large-scale UAV swarm simulation and highly dynamic scenarios. Based on the three fuzzy sets—trustworthiness μtrust, communication delay requirement μdelay, and resource evaluation value μresource— an ordering plus regret-based weight learning (OWA-RL) method is used to aggregate these indicators into a comprehensive payoff function R. The OWA-RL mechanism assigns weights w to the reinforcement learning (RL) agent for online output. This function is used to evaluate the overall payoff of inter-vehicle cooperation or competition strategies:
[0269] (1) First, sort the three fuzzy values to obtain the descending triplets μ(1) ≥ μ(2) ≥ μ(3).
[0270] (μ (1) , μ (2) , μ (3) ):=sort_desc(μ trust , μ delay , μ resource ) (13)
[0271] (2) Potential game modeling
[0272] Let OWA weights w = {w1, w2, w3} ∈ Δ 2 represent the action of the "central dispatcher", and let the joint strategy π of the UAVs / edge nodes represent the action of the subordinate participants.
[0273] Define the potential payoff function Φ(w, π) as the potential function of the game.
[0274]
[0275] where w prior is the prior knowledge of the governance layer, and λ > 0 is the regularization coefficient.
[0276] The Lemma A.1 in the Appendix proves that the weight updating game and the UAV / edge node strategy game share a potential function Φ, thus forming a single-person potential game.
[0277] (3) Regret learning algorithm is adopted to update the central weight with regret.
[0278]
[0279] where J is the learning rate, and μ(τ, k) represents the kth largest membership value after ordering at step τ. Specifically, μ(τ, k) is the hth value after ordering the three membership values in descending order at time τ. According to the regret learning theory, the external regret of the sequence wt is RT / T→0. Combining the potential game property, we have
[0280] Theorem 1 For any ∈ > 0, there exists a step number T(∈) = O(1 / ∈ 2 ) such that the average weight forms an ∈-Nash equilibrium.
[0281] In the scheduling formula of, the interaction among UAVs can be modeled as a finite potential game, where each UAV is a player, and its strategy corresponds to the relative contribution to the selection of different fuzzy aggregation components. If there exists a scalar function Φ: S→R such that for any player i, strategies si, s' i∈ Si and s -i ∈S -i ,
[0282] u i (s′ i ,s -i )-u i (s i ,s -i )=Φ(s′ i ,s -i )-Φ(s i ,s -i ) (16)
[0283] The game G = <N, {Si}, {ui}> is called a potential game: this property ensures that any unilateral improvement of the utility of any UAV is consistent with the increase of the global potential function, which means that local optimization is coordinated with the global network performance improvement. In the case of, the regret learning weight update in formula (14) is based on this potential game structure, so that the learning dynamics can converge to an ∈-Nash equilibrium that maximizes the potential function. This directly links the local adaptive adjustment of the aggregation weight to the maximization of the overall system efficiency.
[0284] (4) OWA aggregation: the comprehensive benefit at time t is calculated as follows:
[0285]
[0286] where, controlling the degree of emphasis on the best indicator, controlling the degree of penalty on the worst indicator. Regret learning automatically increases (risk-aversion) in congestion or low-energy scenarios; and shifts towards averaging or favoring the maximum (efficiency-seeking) when resources are abundant.
[0287] 3.3 Aggregator algorithm flow;
[0288]
[0289] 4. Algorithm proof and analysis
[0290] 4.1 Proof of monotonicity, boundary correctness, and bottleneck priority
[0291] These three properties establish the mathematical predictability of the aggregator (proof in Appendix A):
[0292] • Monotonicity ensures that measurement noise in the input does not induce counter-intuitive jumps in the output.
[0293] • Boundary correctness guarantees that the membership values always lie within the valid interval and behave consistently with extreme cases.
[0294] • Bottleneck priority assigns the most constrained resource the largest decision weight while reserving power compensation space, thus balancing safety and efficiency.
[0295] 4.2 Stability of conjunctive-disjunctive switching (Lyapunov stability)
[0296] The stability results are as follows (proof in Appendix B):
[0297] • The common Lyapunov function V(x) is non-increasing under both modes M1 and M2.
[0298] • The switching surface Σ is continuous, with no sliding mode or Zeno phenomena.
[0299] • According to the geometric convergence in (B-5), the system is globally asymptotically stable in the set M1∪Σ, i.e., eventually satisfies μpre≥ μp—the bottleneck protection mode.
[0300]
[0301] 4.3 Complexity and scalability
[0302] The conclusions are as follows (proof in Appendix C):
[0303] • Time complexity: O(1) per drone using exponential moving average (EMA); O(Lh2) using long short-term memory network (LSTM); linearly scales with the number of resource dimensions and drones (and is parallelizable).
[0304] • Space complexity: O(M) floating-point values are occupied; can be implemented in a streaming fashion on microcontroller units (MCUs) and field-programmable gate arrays (FPGAs).
[0305] • Communication and ordering: low overhead; O(N log N) complexity of centralized ordering is not a bottleneck.
[0306] • Scalable security: theoretical convergence rate is decoupled from parallelism, supporting swarms of thousands of drones.
[0307] 5. Simulation and performance analysis
[0308] To comprehensively verify the practical performance and advantages of the proposed prediction-enhanced two-stage T-norm–Choquet–ordered weighted average (T-norm–Choquet–OWA) aggregator, a high-fidelity joint simulation platform was built using the PX4 software-in-the-loop (SITL) flight controller simulator, the Robot Operating System (ROS2 Humble), and the network simulator (ns-3). PX4-SITL models the kinematics and dynamics of the drones, including flight path control, battery consumption characteristics, and onboard computing load characteristics. ROS2 Humble provides a distributed node communication environment that supports task instruction distribution, resource state monitoring, and real-time inference execution of the aggregator in the swarm. ns-3 provides accurate wireless link simulation for long-term evolution (LTE) and Wi-Fi networks to evaluate metrics such as bandwidth usage, latency, and packet loss rate during task execution.
[0309] To evaluate the performance of the aggregator in large-scale drone swarms, a target tracking and edge inference scenario was designed, in which 360 drones operate in a 10 km x 10 km area. The targets are randomly distributed and dynamically move to simulate emergency tracking tasks in the real world. Each drone must continuously acquire and track its designated target, use onboard edge inference to analyze captured video and sensor data in real time, and transmit the analysis results to the designated edge server through a wireless link.
[0310] Table 3. Simulation resource configuration and task configuration.
[0311]
[0312] Through the above scene setting and accurate simulation running, the method is analyzed and evaluated from resource aggregation strategy and network link quality (including link delay and packet loss rate) and the like, and further verifies that the prediction enhanced two-stage T-norm-choquet-ordered weighted average aggregator has obvious advantages in balancing system safety and efficiency.
[0313] As shown in Figure 4 The violin plot compares the resource membership distribution of the four resource aggregation strategies in the bandwidth-sensitive, computation-intensive, and energy-sensitive task scenarios. The black band represents the 25%-75% quartile range, the horizontal bar represents the full 1.5 times the quartile range, and the dot represents the median:
[0314] (1) Two-stage T-norm-choquet (Resource_choquet)
[0315] The median in the three scenarios is kept between 0.50-0.65, with tight convergence and short tail. This means that the first-stage bottleneck protection prevents extremely low membership values, and the second-stage elastic coupling improves the overall score using non-bottleneck resources. This confirms the "bottleneck priority + power compensation" mechanism described in formula 18.
[0316] (2) Single-layer ordered weighted average (Resource_owa)
[0317] The mean and quartile range are slightly higher than the choquet aggregation, but the tail is longer. The fixed weight can improve the average membership, but cannot prevent extreme bottlenecks, resulting in occasional low values - as predicted for "no prediction peak clipping".
[0318] (3) Min operator (Resource_min)
[0319] All scenarios show a clear left skew. In computation and energy-sensitive tasks, the median drops to 0.15-0.25, and the tail reaches 0.05. This indicates that the "too strict short board effect" severely reduces the feasibility score, sacrificing throughput.
[0320] (4) Arithmetic mean (Resource_mean)
[0321] The median is between 0.45-0.55, and the quartile range is wide, indicating that simple averaging cannot prevent bottlenecks or provide compensation complementarity. Its performance is between ordered weighted average and minimum, consistent with the evaluation of "rigid weight and task insensitivity" in section 2.4.
[0322] These violin plots clearly show that the proposed Joquet aggregator maintains the highest, most stable resource membership across all three scenarios—avoiding the over-converging of the min-operator, overcoming the volatility of the ordered weighted average / arithmetic average—thus validating the synergistic advantages of the prediction enhancement and two-stage design in optimizing resource bottlenecks and power rates.
[0323] As Figure 5 shown, for the 360-UAV swarm employing the prediction-enhanced two-stage aggregation scheduler, the available bandwidth (left) and packet loss rate (right) as a function of node index. These results directly validate the "prediction peaking-bottleneck prioritization" mechanism described in Chapter 4.
[0324] Bandwidth curve: The bandwidth of most nodes stabilizes in the 0.49-0.52 Mbps range. Only nodes 220-240 exhibit a short-lived peak (~0.57 Mbps) before quickly falling back to the baseline. This peak corresponds to a local surge in bandwidth demand triggered by highly concurrent tasks. As the first-stage T-norm has locked the bottleneck and pre-peak, the curve stabilizes again immediately, confirming the immediate protection of bottleneck resources described in Equation 15.
[0325] Packet loss rate curve: The mean remains around 0.025%, with very low variance. A short-lived jitter occurs in the same node segment, below 0.03%, before dropping. This behavior is consistent with the second-stage Joquet-ordered weighted average elastic compensation: after bandwidth peaking, the queue depth decreases, and the packet loss rate drops simultaneously, proving the effectiveness of the power / bandwidth complementary trade-off.
[0326] No tail phenomenon: Neither curve exhibits sustained peaks or oscillations, indicating that the Long Short-Term Memory-Exponential Moving Average (LSTM-EMA) prediction module successfully predicts bursty traffic within a 3-second window and suppresses its propagation. This is consistent with the Lyapunov stability analysis in Section 4.2—the system state quickly returns to the bottleneck stable set.
[0327] Therefore, these slight fluctuations and transient corrections in bandwidth and packet loss rate fully demonstrate that the proposed aggregator not only peaks on the Round-Trip-Time (RTT) but also maintains stable throughput and extremely low packet loss rates at the link layer.
[0328] As Figure 6As shown, under the predictive-enhanced two-stage aggregation scheduler, the round-trip time and signal strength distribution of 360 UAVs are strictly controlled. The round-trip time curve fluctuates narrowly between 53ms and 56ms, with only a brief peak (≈66ms) at nodes 220-240 before quickly returning to the baseline—demonstrating the immediate suppression of sudden bottlenecks by the first-stage T-norm and the elastic compensation by the second-stage Joquit-ordered weighted average. During the same period, the signal strength is concentrated in the range of -76dBm ± 1dB, indicating that the round-trip time variation is mainly caused by link load rather than physical attenuation. By using a 3-second prediction window to pre-clip the peak, the algorithm ensures millisecond-level latency stability for most nodes under small signal strength fluctuations. This figure supports the theoretical premise of this invention, namely, that bottleneck prioritization, power compensation, and forward prediction work together to guarantee a robust real-time control loop.
[0329] like Figure 7 As shown, the node-level link jitter distribution in the scenario with 360 drones is limited to 9.8-11.2ms, with no sustained peaks. Only brief pulses appear at nodes 220-240, which quickly disappear. This is due to:
[0330] A 3-second prediction window pre-cuts peak traffic to suppress queue depth oscillations;
[0331] The first phase, T-norm bottleneck protection, prevents low-membership tasks from monopolizing the link.
[0332] The second stage involves Joquet-ordered weighted average elastic compensation for instantaneous power consumption, balancing transmission intervals.
[0333] These results indicate that the proposed method not only reduces the average round-trip time but also significantly smooths out latency jitter, providing more stable real-time communication quality for drone swarms.
[0334] Figure 8 The mean round-trip time (RTT) curves of five scheduling strategies were compared as the drone swarm size increased from 0 to 360 drones. All methods exhibited sublinear growth, confirming that the marginal impact of queuing delay diminishes with increasing node count. However, the curves showed distinct stratification in both level and slope: the predictive-enhanced two-stage aggregator consistently provided the lowest RTT and the most gradual growth, reaching approximately 55 ms at 360 drones—5%, 10%, 15%, and 20% lower than the minimum, Deep Reinforcement Learning-Proximal Policy Optimization (DRL-PPO), Single-Layer Ordered Weighted Average, and Weighted Sum Model (WSM), respectively. This is consistent with the theoretical model: forward-looking peak shaving lowers the baseline, the first-stage T-norm ensures bottleneck safety, and the second-stage Joquit-Ordered Weighted Average provides elastic compensation to suppress the slope. In contrast, fixed-weight methods (Weighted Sum Model, Ordered Weighted Average) and non-predictive Deep Reinforcement Learning-Proximal Policy Optimization showed steeper growth due to resource mismatch.Figure 8 The trends in Figure 31 show that the method of maintains the lowest initial latency and the smallest growth rate in the large-scale scenario, highlighting its scalable real-time performance.
[0335] To evaluate the scalability and robustness of the proposed algorithm, tests were conducted under four drone swarm sizes (50, 180, 360, and 500 drones), with each configuration running independently five times using different random seeds. Figure 9 The results for round-trip time, jitter, and signal strength are shown, where each data point represents the mean ± standard deviation of repeated trials.
[0336] As shown in Figure 32, the average round-trip time gradually increases as the drone swarm size expands, due to the rising network load and routing complexity. However, the standard deviation remains consistently low (<39 ms), indicating that the latency performance remains stable and predictable even under heavy network conditions. Figure 9 The jitter amplitude is low across all drone swarm sizes, with a slight fluctuation observed at 500 drones. For time-sensitive drone cooperative tasks, this minor variation is within an acceptable range, confirming that the proposed scheme maintains reliable packet timing in large-scale scenarios.
[0337] The signal strength (in dBm) remains relatively stable across all drone swarm sizes, with minimal differences between different random seeds. This consistency demonstrates that the proposed topology control and adaptive link maintenance mechanisms can maintain link quality as node density increases.
[0338] Overall, these results confirm that the proposed algorithm has robust and scalable performance, maintaining low latency, low jitter, and stable signal quality under different drone swarm sizes and random conditions.
[0339] Finally, to ensure a fair and reproducible comparison with the baseline methods, the hyperparameter settings, training procedures, and evaluation environment are described in detail.
[0340] Table 4 lists the hyperparameters, tuning strategies, and training details for all baseline methods, as well as the uniform hardware and software environment to ensure a fair and reproducible comparison.
[0341] Table 4. Baseline configurations and experimental environment.
[0342]
[0343]
[0344] Table 5. Uniform experimental environment
[0345]
[0346] 1. Hyperparameter tuning:
[0347] • For the minima, single-layer ordered weighted average, and weighted sum model baselines, a grid search over the weight vector and operator parameters (l, p) was performed on the validation set to select the configuration that maximized the average resource score without overfitting to specific scenarios.
[0348] • For deep reinforcement learning-proximal policy optimization (DRL-PPO), the default policy network structure and learning rate schedule from the original paper were adopted, followed by a random search over 20 trials to tune the learning rate, clipping ratio, and entropy coefficient. The final configuration was selected based on convergence speed and average reward stability.
[0349] • All baselines adopted the same input normalization and preprocessing pipeline as the proposed method to ensure comparability.
[0350] 2. Training duration: Deep reinforcement learning-proximal policy optimization was trained for 2.5 x 105 episodes (approximately 6 hours of real time) until the moving average reward stabilized within ±1% for 20 consecutive periods. For non-learning baselines, parameter optimization consumed approximately 1.2 hours of CPU time.
[0351] 3. Evaluation environment:
[0352] • Hardware: All methods were executed on the same server equipped with an AMD Ryzen 9 5950X CPU @ 3.4 GHz, 128 GB of memory, and an NVIDIA RTX 3090 GPU (only for deep reinforcement learning-proximal policy optimization).
[0353] • Software: Ubuntu 22.04, Python 3.10, PyTorch 2.0.1, NS-3.37, and ROS2 Humble.
[0354] • Simulation parameters (drone swarm size, topology structure, link model, and channel conditions) were strictly consistent across runs. Each result is the average of 50 independent random seeds.
[0355] Figure 10The presented ablation study results clearly demonstrate the contribution of each module in the proposed two-stage fuzzy aggregation framework. When the prediction component is removed (No Prediction), the task completion rate decreases by about 6.8%, while the average latency increases significantly, indicating that the prediction scheduling is crucial for pre-empting congestion and maintaining time efficiency. Excluding the Jolting-Ordered Weighted Average aggregation stage (No Jolting-Ordered Weighted Average) leads to a more significant performance decrease - the task completion rate decreases by 9.5%, and the throughput decreases by 7.2%, reflecting the importance of nonlinear multi-factor fusion in balancing bottlenecks and remaining resources. In contrast, removing the T-norm stage (No T-norm) results in the most severe performance loss, with the task completion rate decreasing by 12.4% and the throughput decreasing by 10.1%, highlighting the necessity of initial bottleneck-oriented screening before higher-order aggregation. Overall, the complete model outperforms all simplified variants in all metrics, confirming that each module is indispensable and that their joint effect enables the highest robustness and efficiency in large-scale drone swarm scenarios.
[0356] 6. Discussion
[0357] 6.1 Limitations
[0358] During the experiment, the following limitations of this study were found, as shown in Table 6.
[0359]
[0360] 6.2 Complementarity with swarm reinforcement learning and auction games
[0361] In the future, more dimensions can be expanded based on the invention content, complementing methods such as swarm reinforcement learning / auction games. This mainly manifests in the following three aspects:
[0362] Swarm reinforcement learning: Reinforcement learning is good at global optimization with long-term vision, which can provide advanced task allocation priors for the aggregator; the aggregator ensures the safety of low-level resources, forming a two-layer collaborative architecture.
[0363] Auction game: In resource-scarce scenarios, auction-based pricing can suppress excessive requests; the aggregator can map the membership output to the upper limit of the bid, implementing a "trust-elastic" auction mechanism.
[0364] Mixed advantages: Reinforcement learning and auction-driven strategic exploration, while the two-stage aggregator enforces safety constraints. Their combination can achieve fast convergence and effective risk control.
[0365] 6.3 Scalability of the two-stage aggregator
[0366] Scalability is evaluated for N = {60, 120, 180, 240, 300, 360, 480, 600} drones (5 random seed trials per N) reporting mean ± std. dev. of round-trip time / jitter / signal strength. Figure 9 Non-linear least squares fit of round-trip time RTT(N) = a + b√N indicates that the proposed method has the smallest slope b among all schedulers, with narrow confidence intervals across random seeds, indicating robustness to random dynamic changes as the swarm size scales up.
[0367] From a theoretical perspective, each node's computation involves a fixed number of operations: three prediction-enhanced membership a T-norm μpre= min{·}, and a binary Borda-ordered weighted average based on a fixed capacity v of {μpre, μp}. Thus, each node's time and memory complexity is O(l); the entire swarm's runtime per control step is therefore O(N). In a decentralized setting with a limited neighborhood degree d (due to radio range limitations), each node only exchanges local statistics, resulting in O(l) messages per node and O(N) messages in total.
[0368] Finally, the queuing / interference analysis explains the empirical scaling law: due to spatial reuse, the effective contention radius grows sublinearly with N, causing the delay envelope to increase as √N. The first phase (T-norm) guarantees bottleneck priority, preventing compensatory bias that could amplify delay with N; the second phase (Borda-ordered weighted average) provides rank-sensitive compensation via φ, cutting peaks without violating bottleneck safety. This division explains the observed low slope b and tight variability relative to the min, weighted sum, and single-layer ordered weighted average models, and deep reinforcement learning - proximal policy optimization.
[0369] 6.4 Security Considerations
[0370] While this study focuses on scheduling and resource aggregation, security is an integral dimension in large-scale drone-iot systems. Drone nodes typically run embedded firmware that can contain exploitable vulnerabilities. Recent research on iot firmware vulnerability detection shows that combining static binary analysis, dynamic execution tracing, and symbolic execution can effectively expose buffer overflow, command injection, and authentication bypass flaws in resource-constrained devices
[30] . These methods are typically integrated with fuzzing frameworks to ensure that each drone's communication and control stack is free of critical flaws before deployment.
[0371] In addition to traditional static / dynamic analysis, emerging research leverages large language models (LLMs) to perform software security tasks
[31] . LLMs can assist in code review, vulnerability classification, and even automated patch synthesis by understanding natural language security recommendations and translating them into actionable code changes. For a drone swarm, such capabilities can enhance the firmware verification process, helping to detect known common vulnerability and exposures (CVEs) and zero-day vulnerabilities before field deployment.
[0372] In the operational scenario, the two-stage T-norm-Joinville-ordered weighted average scheduler can be integrated with these security mechanisms at two points: (i) pre-flight - ensuring that only verified firmware images participate in the drone swarm; (ii) in-flight - combining resource scheduling with real-time anomaly detection, so that nodes exhibiting abnormal traffic patterns or latency spikes (possibly indicating a breach) are downgraded or isolated. This joint design will enhance the reliability and security posture of the entire system.
[0373] The present invention aims to provide a solution to the problem of real-time aggregation of heterogeneous resources and prediction of failure risks in multi-drone cooperative tasks. Existing technologies typically rely on static threshold control or average allocation, making it difficult to balance short-term fluctuations, extreme bottlenecks, and instantaneous power responses of three key resources: computation, communication, and energy. This can easily lead to delays, packet loss, and even task interruption in high-load or energy drop situations. The present invention introduces a combination of long short-term memory networks and exponential moving average models in the prediction layer to predict the resource trajectory for the next three seconds, making scheduling decisions forward-looking and effectively improving the stability and security of task execution.
[0374] In the resource modeling section, the present invention establishes a multi-dimensional fuzzy set, normalizing CPU utilization, bandwidth occupancy, and remaining energy to a unified dimensionless interval, and introduces a future-aware membership function to inject prediction information. In this way, CPU and bandwidth will trigger shrinkage penalties in advance when peak loads are predicted, and the energy indicator will be corrected in combination with expected consumption, thereby achieving a "predicted shrinkage" effect. This mechanism solves the problem of non-comparable heterogeneous indicators and insufficient prediction in traditional methods, enabling the scheduling system to make more accurate resource allocation decisions in dynamic environments.
[0375] In the design of the aggregation strategy, the application adopts a two-stage architecture. The first stage retains the "bottleneck effect" based on T-norm minimum operation, ensuring that extreme shortage of any single resource will trigger task rejection or migration, fundamentally ensuring the bottom line safety. The second stage introduces the Choquet-OWA operator and combines the adaptive interaction metric φ to realize the elastic coupling of power consumption and resource adequacy. When the task is sensitive to performance and energy sufficient, the value of φ increases to highlight performance priority; on the contrary, in the energy-saving priority or energy shortage scene, the value of φ decreases to strengthen the power constraint, thereby realizing adaptive dynamic adjustment across task categories.
[0376] In the income feedback mechanism, the application maps the fuzzy aggregation output and the weight update of the central scheduler to the potential function Φ through potential game modeling, and adjusts the weights by using weighted regret learning. This framework not only avoids the problem of centralized scheduling rigidity, but also overcomes the defects of distributed game greed, thereby achieving a balance between central governance and individual UAVs. The operation mechanism ensures that the system gradually converges to a globally stable and efficient equilibrium state in the long-term evolution, providing reliable self-learning scheduling capabilities for UAV clusters.
[0377] The two-stage T-norm-Choquet-OWA resource aggregation model proposed by the application constructs a closed-loop link covering prediction, aggregation and income feedback, and can realize adaptive resource scheduling for multi-UAV cooperative tasks in complex environments. This method not only significantly improves the task completion rate and energy efficiency in industrial applications, but also enhances the robustness of the system, providing a solid technical guarantee for UAV clusters in high-load operations, complex communication and long-duration tasks.
[0378] The application aims to solve the technical problem of real-time aggregation of heterogeneous resources and prediction of failure risks in a multi-UAV cooperative environment. Existing technologies rely on static threshold control or average scheduling, which cannot model short-term fluctuations, extreme bottlenecks and instantaneous power responses of three key resources: computation, communication and energy. This leads to delays, packet loss and even task interruption in sudden load or energy drop scenarios. The application introduces a long short-term memory network combined with an exponential moving average model in the prediction layer to predict the future three-second resource trajectory, thereby changing the scheduling decision from "passive response" to "active avoidance", effectively improving the stability of task execution.
[0379] At the resource modeling level, the application establishes a multi-dimensional fuzzy set, normalizes CPU utilization, bandwidth occupancy and residual energy to a dimensionless interval, and injects prediction information through a future perception membership function to trigger shrinkage penalty in advance under the upcoming peak of CPU and bandwidth, and the energy function is modified in combination with expected consumption to realize the "predicted shrinkage" effect. This mechanism overcomes the resource allocation errors of traditional schemes under the incomparable and insufficient prediction of heterogeneous indicators, ensuring that the scheduling decision is closer to the dynamic actuality of the flight task.
[0380] In terms of aggregation policy, the application adopts a two-stage structure: the first stage strictly retains the "bottleneck effect" through T-norm minimum value operation, and extreme shortage of any key resource directly determines task rejection or migration, thereby ensuring the bottom line safety; the second stage introduces the Choquet-OWA operator combined with adaptive interaction measure φ to elastically couple power consumption and resource sufficiency, when task performance is prioritized and energy is sufficient, φ takes a large value to ensure performance; when task energy saving is prioritized or energy is insufficient, φ takes a small value to strengthen power constraints, forming adaptive dynamic adjustment across task categories.
[0381] In the benefit feedback mechanism, the application constructs a potential game framework, maps the fuzzy aggregation output and the central scheduler weight to a unified potential function Φ, and realizes weight update through weighted regret learning. This mechanism effectively balances the game relationship between the central governance layer and the individual UAVs, avoids centralized scheduling rigidity or distributed game greed, and ensures that the system gradually evolves to a globally stable and efficient equilibrium state in the long-term operation.
[0382] The two-stage T-norm-Choquet-OWA resource aggregation model proposed by the application forms a "prediction-aggregation-benefit" closed-loop link, which significantly improves the task completion rate, energy efficiency and system robustness in multi-UAV cooperative industrial applications. Through the combination of forward-looking prediction, bottleneck protection and elastic coupling, the application realizes adaptive resource scheduling for complex dynamic environments, providing reliable technical support for UAV clusters in high-load tasks, complex communication and long-duration applications.
[0383] It should be noted that embodiments of the present application can be realized by hardware, software, or a combination of software and hardware. The hardware portion can be realized by a special logic; the software portion can be stored in a memory and executed by a proper instruction execution system, such as a microprocessor or a specially designed hardware. A person of ordinary skill in the art can understand that the above-mentioned apparatus and method can be realized by computer executable instructions and / or included in processor control codes, such as a carrier medium, such as a magnetic disk, CD or DVD-ROM, a programmable memory, such as a read-only memory (firmware), or a data carrier, such as an optical or electronic signal carrier. The apparatus of the present application and its modules can be realized by a hardware circuit, such as a very large scale integrated circuit or a gate array, a semiconductor, such as a logic chip, a transistor, or a programmable hardware device, such as a field programmable gate array, a programmable logic device, or the like, by software executed by various types of processors, or by a combination of the above-mentioned hardware circuit and software, such as firmware.
[0384] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any modification, equivalent replacement, and improvement within the technical range disclosed by the present application, and within the spirit and principle of the present application, should be covered within the protection scope of the present application.
Claims
1. A model proof and verification method of a two-stage T-norm-Choquet-OWA resource aggregator for multi-UAV cooperation, characterized in that, The method comprises the following steps: The LSTM-EMA prediction process comprises: normalizing CPU utilization, bandwidth utilization, and residual energy, constructing a resource use multi-dimensional fuzzy set, and combining a comprehensive fuzzy benefit function to realize forward-looking prediction of resource trajectory. The method comprises the following steps:
2. The method of claim 1, wherein, The CPU prediction enhancement membership function is based on the weighted sum of current utilization and predicted utilization, and adjusts the task sensitivity to calculation, inflection point, and curve steepness through parameters λc, θc, and κc, to realize the upcoming peak penalty.
3. A multi-unmanned aerial vehicle (UAV) coordination oriented resource usage multi-dimensional fuzzy set modeling method, characterized in that, The method comprises the following steps: In the first stage, the T-norm is used to perform extreme protection on three-dimensional resources of CPU, bandwidth, and energy. When any membership is zero, it is directly determined that the task is infeasible. In the second stage, the φ factor of the task elasticity η and the predicted residual energy Fe jointly act on the power consumption rate membership to realize the dynamic balance of protection and recovery.
4. The method of claim 3, wherein, The method comprises the following steps:
5. A two-stage extreme protection and elastic coupling method for multi-UAV cooperation, characterized in that, In the first stage, the bottleneck membership μpre is extracted based on the minimum operator; In the second stage, the Choquet-OWA operator with interaction measure ν is used to aggregate the complementarity between resource sufficiency and power consumption rate, and the scene adaptive adjustment is realized through φ parameter modulation.
6. A two-stage T-norm-Choquet-OWA resource aggregation method for multi-UAV cooperation, characterized in that, The Choquet-OWA operator models the complementary relationship of the two dimensions through descending order sorting of membership and interaction measure ν, and performs power compensation between μpre and μp. When the power consumption rate is better than the bottleneck, the compensation value is adopted, otherwise the bottleneck protection priority is maintained. The method comprises the following steps: The three fuzzy values are sorted in descending order, the potential benefit function Φ(w, π) is defined by combining the OWA weight of the central scheduler and the joint strategy of the UAV node, and the regret learning algorithm is used to dynamically update the central weight, so that the resource scheduling converges to the ∈-Nash equilibrium.
7. The method of claim 6, wherein, The regret learning algorithm increases the weight of the worst indicator in the congestion or low energy scene, and increases the weight of the best indicator in the resource sufficient scene, to realize the adaptive balance between risk aversion and efficiency pursuit.
8. A multi-UAV cooperation-oriented comprehensive fuzzy benefit function construction method, characterized in that, The method comprises the following steps: A prediction module is configured to predict resource trajectory in the next 3 seconds based on an LSTM-EMA model; 9. The method of claim 8, wherein, A compensation module is configured to perform two-stage T-norm and Choquet-OWA aggregation to realize joint decision of bottleneck protection and power compensation; and through a fuzzy benefit function and a potential game mechanism, the convergence and stability of the system under the cooperation of the UAV group are guaranteed.
10. A model proof and verification system for multi-UAV cooperative oriented two-stage T-norm-Choquet-OWA resource aggregator, characterized in that,