Task scheduling optimization method based on double-agent strategy in computing network environment

By employing the Dual Agent Policy Optimization (DAPO) algorithm in an edge computing environment, combined with communication and computing models, and dynamically adjusting privacy protection, the problems of conservative policy updates, poor adaptability, and insufficient privacy protection in existing technologies are solved, achieving efficient and stable task offloading and privacy protection.

CN121037441APending Publication Date: 2025-11-28NORTHEASTERN UNIV CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511217935.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing task offloading methods exhibit overly conservative policy updates, poor adaptability, insufficient balance between exploration and exploitation in complex and dynamic edge computing environments, low efficiency in parameter sharing in multi-agent systems, suboptimal resource allocation among heterogeneous devices, and fixed privacy protection measures that make them vulnerable to inference attacks, making it difficult to dynamically adjust according to computing needs and network conditions.

Method used

A task scheduling optimization method based on dual agent strategy is adopted. By combining the dual agent strategy optimization algorithm (DAPO) with communication and computing models, the privacy protection level is dynamically adjusted, the privacy entropy is quantified to optimize task offloading, and the dual agent collaboration is used to accelerate convergence and explore the policy space, thereby enhancing the system stability and adaptability.

Benefits of technology

It achieves more efficient task offloading in dynamic edge computing environments, dynamically adjusts privacy protection, improves system stability and adaptability, reduces latency and energy consumption, and enhances resistance to inference attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037441A_ABST
    Figure CN121037441A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud task unloading and reinforcement learning, and particularly provides a task scheduling optimization method based on a double-agent strategy in a computing network environment, and the method comprises the steps: obtaining the energy consumption and delay of a task in an unloading process from a local processor MSDn to an edge server ECs based on a communication model; constructing an optimization target of task unloading based on the calculation model; and according to the optimization target, adopting a DAPO algorithm based on a double-agent strategy to obtain an optimal unloading strategy of the cloud edge end cooperation task. According to the method, the stability of the algorithm in a dynamic and complex environment is enhanced through double-agent configuration, the complex system state, space and decision-making process can be managed more effectively, privacy is quantified through the information entropy model, the uncertainty of an unloading mode is quantified, and therefore the attack inference difficulty is quantified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cloud task offloading and reinforcement learning technology, and more specifically, to a task scheduling optimization method based on a dual-agent strategy in a computing network environment. Background Technology

[0002] Research on task offloading cost optimization has undergone several stages of development. Early research mainly employed traditional convex optimization and heuristic algorithms to solve resource allocation problems in static scenarios; however, this approach struggles to cope with highly dynamic and rapidly changing network environments. As edge computing environments become increasingly complex and dynamic, research focus has gradually shifted to more adaptive machine learning methods, especially reinforcement learning techniques. Recent studies have shown that while traditional reinforcement learning algorithms, such as DQN and DDQN, have achieved some success in task offloading decisions, they face problems such as slow convergence speed and unstable policy quality in highly dynamic mobile edge computing environments. The latest trend is the adoption of multi-agent methods, such as the multi-task offloading method proposed by Dai et al. in Seq2seq-based meta-reinforcement learning.

[0003] Despite these advances, existing methods still face several key limitations. Current reinforcement learning methods exhibit overly conservative policy updates in complex edge environments, leading to suboptimal solutions. Furthermore, these methods are poorly adaptable to rapidly changing network conditions, especially in highly mobile scenarios. Additional challenges include an insufficient balance between exploration and exploitation, inefficient parameter sharing in multi-agent systems, and suboptimal resource allocation across heterogeneous devices. Moreover, existing task offloading mechanisms often employ uniform privacy protections, resulting in predictable patterns that are vulnerable to inference attacks. Many methods treat privacy as a fixed cost, failing to adapt to actual computational demands and network dynamics. Furthermore, current methods often either sacrifice system performance for privacy or fail to provide sufficient privacy guarantees while optimizing performance. Most solutions also lack sufficient flexibility to cope with varying server loads, network conditions, and user mobility patterns. Summary of the Invention

[0004] In view of this, the present invention proposes a task scheduling optimization method based on a dual-agent strategy in a computing network environment to solve the problems existing in the prior art.

[0005] To achieve the above objectives, this invention proposes a task scheduling optimization method based on a dual-agent strategy in a computing network environment, comprising:

[0006] Based on the communication model, obtain the energy consumption and latency during the offloading process of the task from the local processor MSDn to the edge server ECs;

[0007] Optimization objectives for task unloading are constructed based on a computational model;

[0008] Based on the optimization objective, the DAPO algorithm based on a dual-agent strategy is used to obtain the optimal offloading strategy for cloud-edge-device collaborative tasks.

[0009] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0010] This invention designs a comprehensive resource allocation model that integrates multiple factors, including computational requirements, information entropy, and network dynamics, to optimize task offloading in dynamic edge computing environments. Furthermore, this invention proposes a dual-agent strategy optimization scheme for privacy-conscious task offloading. The framework employs two complementary agents: one focused on efficient offloading strategies, and the other emphasizing privacy protection, thus forming a balanced optimization approach. Unlike static privacy measures, this invention's system can dynamically adjust the protection level based on the computing environment, network conditions, and user preferences. Simultaneously, this invention quantifies privacy through an information entropy model, which quantifies the uncertainty of offloading patterns, thereby quantifying the difficulty of inference attacks.

[0011] The DAPO algorithm proposed in this invention offers significant advantages over traditional PPO and other methods. By using two agents, DAPO accelerates the convergence process and explores the policy space more effectively. The dual-agent configuration enhances the stability of the algorithm in dynamic and complex environments. By assigning the learning task to two agents, DAPO can more effectively manage complex system states, spaces, and decision-making processes, making it a more robust and adaptable solution to the mobile edge computing offloading challenge. Attached Figure Description

[0012] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings:

[0013] Figure 1 This is a schematic diagram of a computational offloading scenario considering information entropy in a mobile edge computing IoT architecture according to an embodiment of the present invention.

[0014] Figure 2 This is a framework diagram of the dual-proxy strategy optimization algorithm DAPO in this embodiment of the invention;

[0015] Figure 3 shows the performance of different algorithms in the iterative process in the embodiments of the present invention, where (a) is the reward, (b) is the latency, and (c) is the energy consumption;

[0016] Figure 4 shows the performance of different algorithms in the embodiments of the present invention on the average data volume (MB), where (a) is the reward, (b) is the latency, and (c) is the energy consumption;

[0017] Figure 5 shows the performance of different algorithms in the embodiments of the present invention under the maximum transmission power MSD (milliwatts), where (a) is the reward, (b) is the latency, and (c) is the energy consumption;

[0018] Figure 6 shows the performance of different algorithms in the user equipment computing power (per bit cycle) in the embodiments of the present invention, where (a) is the reward, (b) is the latency, and (c) is the energy consumption;

[0019] Figure 7 shows a comparison of the PPO algorithm and the DAPO algorithm in terms of action loss and value loss in the embodiments of the present invention, where (a) represents action loss and (b) represents value loss.

[0020] Figure 8 shows the performance of different algorithms before and after the introduction of information entropy in the embodiments of the present invention, where (a) is the DDQN algorithm, (b) is the DQN algorithm, (c) is the DAPO algorithm, (d) is the PPO algorithm, (e) is the A3C algorithm, and (f) is the SAC algorithm. Detailed Implementation

[0021] This embodiment designs a mobile edge computing system, such as Figure 1 As shown, computational tasks are generated by users, and the computational complexity of a task depends on its specific requirements, necessitating different computing resources. Generated tasks are randomly assigned to Mobile Smart Devices (MSDs), which possess certain computing resources and can handle some of the tasks. However, when computing power is insufficient, these tasks need to be offloaded to edge servers. Since edge servers have stronger computing capabilities, they can not only handle tasks offloaded from MSDs but also feed the completed tasks back to the MSDs. This scenario involves N Mobile Smart Devices (MSDs) and M Edge Computing Servers (ECs). Let N = {1,2,...,N} and M = {1,2,...,M} represent the sets of MSDs and ECs.

[0022] MSDs: Smart devices with network capabilities, such as smartphones, laptops, and car tablets. They acquire information about the external environment through sensors and collect data to perform simple calculations.

[0023] EC: Edge computing servers are generally deployed closer to user devices and data sources, and are characterized by high computing speed, large storage capacity, and high security.

[0024] like Figure 1As shown, the computational tasks of MSD are easily affected by various factors, such as user behavior, the computing power of hardware devices, and network speed. These factors are variable and uncertain, leading to the randomness of task arrival. In the simulation, the system is modeled as an environment based on a discrete time period.

[0025] Table 1

[0026]

[0027] For ease of understanding, the main variables are listed in Table 1. Time is divided into several time intervals, labeled t, with each interval having a length of T. The reason for using this modeling method is that the state of the moving average distance (MSD) is essentially determined within the time interval T. The calculation task is represented as A. n (t), which usually follows a Poisson distribution:

[0028] (1)

[0029] Where, λ n MSD n Calculate the average arrival rate of tasks in each time slot t, where G represents the number of tasks. This formula describes the probability of reaching a specific number of tasks within a given time period.

[0030] Problem Description

[0031] Different scenarios have different requirements for computation offloading and allocation, and therefore different strategies are adopted. In this embodiment, to protect data privacy, the offloading strategy is to perform computation on a multi-server cluster (MSD) whenever possible to prevent task data from being leaked through communication channels during the offloading process. If the computational demand exceeds the maximum capacity of the MSD, the excess portion is offloaded to ECs.

[0032] (2)

[0033] Assuming from MSD n To EC m The energy consumption and delay of the unloading process are E(m,n,t) and T(m,n,t), respectively. The weighted sum of energy consumption and delay, Q, can be defined to evaluate the unloading operation.

[0034] (3)

[0035] The specific calculations of E(m,n,t) and T(m,n,t) will depend on the communication mode and calculation model described below.

[0036] 1) Communication Model: The communication method should employ Orthogonal Frequency Division Multiple Access (OFDMA) between the multi-carrier (MSD) and spreading codes (ECs) to ensure that each modulated signal is strictly orthogonal, thereby enabling the receiver to completely separate the individual signals. Due to the strict orthogonality of the modulated signals, the interference between channels is very small, which allows for the application of ergodicity in practical applications.

[0037] In time slot t, the bandwidth allocated from the edge server ECm to MSDn can be represented as follows:

[0038] (4)

[0039] This embodiment found that in equation 4, when n∈ And ∀ When =1, EC m The communication resources will be in MSD n The bandwidth is evenly distributed across the system. Due to the ergodicity of the system, the bandwidth can be evenly distributed, thus simplifying the model:

[0040] (5)

[0041] In evaluating signal quality, this embodiment introduces the concept of signal-to-noise ratio (SINR), which is defined as the ratio of signal strength to noise strength. The definition is as follows:

[0042] (6)

[0043] Where P signal P represents signal power. noise Indicates noise power.

[0044] In the text, the MSD of time slot t on subchannel k. n To EC m The mathematical expression for SINR can be represented as:

[0045] (7)

[0046] In these cases, since the interference between units is ignored, it can be considered as 0. However, in practical applications, interference between units is sometimes unavoidable. Therefore, we can utilize ergodicity to replace the value of all time points with the result of a single measurement, or even directly set it as a constant.

[0047] Theorem 1: For an ergodic process, its statistical properties (all of which are statistical averages) can be completely replaced by the time average of any realization of the stochastic process.

[0048] Therefore, the above equation can be rewritten as:

[0049] (8)

[0050] Where c is the noise result obtained from a certain measurement.

[0051] Using the signal-to-noise ratio, the Shannon formula can be used to calculate the channel capacity:

[0052] (9)

[0053] Where C represents the channel capacity (i.e., the maximum data rate that the channel can transmit under given bandwidth and signal-to-noise ratio (SINR), and B represents the channel bandwidth. Therefore, during time slot t, EC m With MSD n The transmission rate between them can be expressed by the following mathematical expression:

[0054] (10)

[0055] At the same time, privacy protection must be ensured during data transmission, which consumes energy and time. Generally speaking, information with high privacy protection requirements requires more computing resources, latency, and energy consumption during processing.

[0056] The privacy entropy model proposed in this embodiment is based on the game theory framework established by Zhu and Bas'ar. Although their research mainly focuses on security strategies, this embodiment further extends this framework to quantify privacy requirements in task offloading scenarios. Specifically, the privacy entropy measurement in this embodiment considers: (i) the sensitivity of offloading data, which is reflected in the probability distribution of offloading decisions; (ii) the uncertainty of task allocation patterns, which affects the predictability of data flow; and (iii) the resource requirements for implementing privacy protection mechanisms. This quantitative approach allows us to incorporate privacy considerations into resource allocation and task scheduling decisions while maintaining system efficiency.

[0057] During data transmission on mobile devices, data security needs to be considered. Therefore, this embodiment proposes privacy entropy to measure the importance of data. The higher the privacy entropy, the more difficult it is to predict the data's uninstallation path, indicating that the data security requirements are higher.

[0058] The set of all MSD-ECs that performed unloading tasks is represented by X. k = {x k,1 x k,2 , ..., x k,m}, where x k,i MSD k The computing tasks were offloaded to EC i The probabilities of these uninstallation events are determined by the set P. k = {p k,1p k,2 , ..., p k,m} represents, where p k,i MSD k The computing tasks were offloaded to EC i The probabilities of these probabilities satisfy the following constraints:

[0059] (11)

[0060] This embodiment can express these two sets in the following way to more intuitively reflect the probabilistic model of mobile edge computing offloading.

[0061] (12)

[0062] According to MSD k Determine the privacy entropy of the uninstallation task.

[0063] (13)

[0064] Based on the above formula, the average privacy entropy of all transmitted information can be calculated as follows:

[0065] (14)

[0066] Define the privacy factor pr n for:

[0067] (15)

[0068] This coefficient reflects MSD k The greater the uncertainty of the unloading task, the higher the data security requirements. This coefficient can be used as a multiplier for the normal transmission time to obtain the transmission time that takes privacy protection into account.

[0069] Therefore, according to Formulas 10 and 15, the value from MSD can be calculated. n To EC m The transmission delay. The mathematical expression is as follows:

[0070] (16)

[0071] The time consumed when the task is executed locally is:

[0072] (17)

[0073] If the computation task is offloaded to an EC, the time required to execute the computation task within the EC is:

[0074] (18)

[0075] Local enforcement capacity can be represented as follows:

[0076] (19)

[0077] in It is the energy coefficient related to the chip architecture, C n It is MSD n Regarding the CPU workload, it's important to note that the allocation of MSD (Maximum Segment Depth) remains constant across the system from one time period to the next. Therefore, local energy consumption can be expressed as... (t), its calculation formula is as follows:

[0078] (20)

[0079] MSD n Convert to EC m The energy consumption is defined as follows:

[0080] (twenty one)

[0081] The energy consumption for executing tasks in ECs is:

[0082] (twenty two)

[0083] For a given time slot t, the local processor MSD n and edge server EC m The total latency of task transmission and computation can be expressed as:

[0084] (twenty three)

[0085] Under a given time slot t, task transfer and computation occur on the local processor MSD. n and edge server EC m The total energy consumption can be expressed as:

[0086] (twenty four)

[0087] 2) Computational model: Object constraints.

[0088] Its functions mainly include the following five parts:

[0089] a) Local processor CPU cycle frequency constraint: This represents the CPU cycle frequency of the local processor. The local processor can dynamically adjust this parameter to reduce power consumption, but this parameter is not unlimited and must be adjusted to avoid damaging the local processor. Therefore, the following limitations apply:

[0090] (25)

[0091] b) Computational resource constraints of edge processors: In EC m MSD allocated within the coverage area n The computing resources of edge processors are limited. Edge processors possess powerful computing capabilities, capable of handling computational tasks that local processors cannot complete. Unlike local processors, edge processors can handle multiple tasks. However, the computing resources of edge processors are also limited, and clearly, the following constraints must be met:

[0092] (26)

[0093] c) Bandwidth limitations: Indicates allocation to MSD n The communication resources, which are located in EC m Within the coverage area. Similarly, this allocation must not exceed the maximum bandwidth and must meet the following conditions:

[0094] (27)

[0095] Theorem 2: The bandwidth allocation problem among nodes in edge computing networks is an NP-hard problem.

[0096] Therefore, precise calculations are not required; it is sufficient to verify and adopt a reasonable solution.

[0097] 4) Wireless channel power constraints: P n (t) represents the wireless channel power of MSDn in time slot t. MSD n The maximum permissible wireless channel power during wireless communication in time slot t. To ensure no interference occurs between channels and to better reflect the system simplification conditions that satisfy ergodicity requirements, P must be... n (t) is constrained:

[0098] (28)

[0099] 5) Sub-channel constraints: This indicates that during time slot t, the EC server EC m To MSD device MSD n The bandwidth allocation coefficient must meet the following constraints. ∈ Since each subchannel is assigned to only one MSD, the following constraints must be satisfied:

[0100] (29)

[0101] Furthermore, each MSD is assigned to a separate sub-channel, and the following conditions are clearly met:

[0102] (30)

[0103] Based on task latency and energy consumption, Equation 3 was initially defined as the optimization objective. However, this is an objective function used to evaluate a single unloading operation. Considering the optimization of the entire system, the overall objective function C... total It can be defined as the sum of the Q(m,n,t) values ​​of all unload operations within the system.

[0104] (31)

[0105] Follow these constraints: [25-30]

[0106] The practical application of the privacy entropy model proposed in this embodiment directly affects system resource allocation and task offloading decisions, all of which are made within the aforementioned optimization framework. Specifically, the privacy coefficient pr defined in equation (15)... n It functions as a multiplication factor in transmission delay calculations.

[0107]

[0108] Higher privacy requirements automatically necessitate more communication resources. This approach embeds privacy protection into the resource allocation mechanism—higher entropy leads to a more conservative offloading strategy and increases resource reservation for security overhead. By integrating privacy considerations into the objective function and constraints, a comprehensive optimization framework is created that addresses both performance and security issues. The next section introduces the DAPO algorithm of this embodiment, which effectively solves this optimization problem by dynamically balancing privacy requirements with performance constraints, creating a privacy-aware offloading strategy that adapts to system conditions and evolving privacy requirements.

[0109] The PPO approach performed poorly, primarily due to the limited scope of policy updates. However, overly conservative update strategies can sometimes hinder the discovery of breakthrough offloading strategies. When sufficient time and computational resources are available, the following methods can be explored to identify breakthrough offloading strategies.

[0110] Figure 2The DAPO algorithm, an enhanced version of the PPO algorithm, is presented. Like all reinforcement learning algorithms, DAPO operates on the fundamental principle of agent-environment interaction. The agent interacts with the environment, updating its policy based on the resulting state and reward. This iterative process continues, optimizing the policy through repeated interactions. DAPO achieves its uniqueness by using two agents that interact with the same environment but employ slightly different policy update mechanisms. This dual-agent framework offers significant advantages over traditional PPO and other methods in mobile edge computing offloading tasks. By using two agents, DAPO accelerates the convergence process and explores the policy space more efficiently. This dual-agent configuration also enhances the algorithm's stability in dynamic and complex environments, such as the significant variations in network conditions and task loads common in mobile edge computing. By assigning the learning task to two agents, DAPO can more effectively manage complex system states, spaces, and decision-making processes, making it a more robust and adaptive solution to mobile edge computing offloading challenges.

[0111] The pseudocode for the policy update process is summarized in Algorithm 1, and includes the following five steps:

[0112] (1) Initialize the agents. Ensure that agent 1 and agent 2 are the same at the start of the algorithm (line 1).

[0113] (2) Agent 1 and Agent 2 interact with the environment respectively (lines 3-7).

[0114] (3) At specific intervals, the performance of Agent 1 and Agent 2 is compared. If the current interval is a multiple of 100 (including interval 0), the returns of the two are compared. If the return of Agent 1 is significantly higher than that of Agent 2, the policy parameters of Agent 1 are set to those of Agent 2; otherwise, the policy parameters of Agent 2 are set to those of Agent 1. In this way, the r(θ) of Agent 1 will not be pruned, while the r(θ) of Agent 2 will be pruned (lines 8-21).

[0115] (4) Next, continue with the regular policy update. If the current period is not a multiple of 100, both Agent 1 and Agent 2 will perform the regular policy update, and their r(θ) will not be pruned (lines 22-33).

[0116] (5) Algorithm convergence. If either agent 1 or agent 2 converges, it indicates that the algorithm has reached a satisfactory convergence state, and the algorithm will be terminated (lines 34-39).

[0117]

[0118] The aim of this approach is to ensure that while maintaining the inherent stability of PPOs, there is still an opportunity for significant policy updates. This is achieved through selection based on return value.

[0119] Theorem 3: Under the conditions of bounded policy space, reasonable learning rate scheduling, and dual-agent cooperation mechanism, the DAPO algorithm can converge to the local optimum policy of unbiased gradient estimation.

[0120] The following is a formula explanation for the PPO algorithm:

[0121] During execution, each state and action in each step can be recorded and represented as a set. This set represents the trajectory of the algorithm's execution:

[0122] (32)

[0123] The exact objective function J(θ) is expressed as:

[0124] (33)

[0125] The simplified J(θ) is:

[0126] (34)

[0127] The current goal is to find the optimal θ parameter that maximizes J(θ):

[0128] (35)

[0129] Calculate the gradient of J(θ):

[0130] (36)

[0131] After importance sampling, the strategy with parameter θ will be replaced by a new strategy with parameter θ′:

[0132] (37)

[0133] Here, πθ(τ) represents the policy corresponding to the current parameters, while πθ′(τ) represents the policy corresponding to the updated parameters. Under the PPO2 framework, policy updates should not be too aggressive and must adhere to specific constraints:

[0134] (38)

[0135] The objective function after the shearing operation is:

[0136] (39)

[0137] (40)

[0138] The following are the differences between the DAPO algorithm and the PPO algorithm:

[0139] In the proposed DAPO algorithm, r(θ) is re-evaluated every 100 cycles, as follows:

[0140] (41)

[0141] The new strategy interacts with the environment, and its J(θ) value is calculated for each interaction. The parameter set θmax, which increases the J(θ) value, is selected as the policy parameters. Subsequently, the new policy is set to π. θmax (τ) .

[0142] This embodiment employs extensive simulations to verify the performance of the proposed framework.

[0143] A. Data and Experimental Setup

[0144] In this embodiment, multiple EC servers and multiple MSDs are configured within a 500m × 500m area. The number of these servers and MSDs remains constant over short periods. Specifically, the number of EC servers is set to 5, and the number of MSDs is set to 100. The task size follows a Poisson distribution with a parameter value of 5. Furthermore, the computing power of the MSDs varies from 15 cycles / bit to 40 cycles / bit, and the computing resources of each MSD are randomly allocated to {0.5, 0.8, 1.0} GHz. The constant power factor of the MSDs is set to 10⁻²⁶ watts [s³ / cycles³]. The transmit power of the MSDs is set to 1000 milliwatts. The noise spectral density is set to 10⁻⁹.

[0145] W. The privacy factor for each MSD-EC pair is selected from the range [0, 2.32].

[0146] B. Comparison Algorithm

[0147] To demonstrate the effectiveness of the proposed DAPO in achieving computational offloading while protecting privacy, the following experiments were designed:

[0148] Experimental Setup: The experiments in this embodiment were conducted in an environment based on information entropy. The aim was to evaluate the ability of deep reinforcement learning algorithms to effectively allocate resources while considering the privacy factor of MSD-EC. Therefore, this embodiment compared the algorithm's performance in environments with and without information entropy.

[0149] Algorithm Comparison: DAPO is a code-level optimization of the standard PPO algorithm, and its theoretical advantages are difficult to quantify. However, by comparing its value loss and action loss metrics with those of the original PPO algorithm, its inherent advantages can be intuitively revealed.

[0150] Robustness assessment: Robustness is evaluated by comparing the performance of different algorithms in various environments containing information entropy.

[0151] Baseline Algorithms: To benchmark the proposed algorithm, six currently popular algorithms were selected as baselines: PPO, DDQN, DQN, SAC, A3C, and a randomized strategy.

[0152] C. Experimental Results

[0153] Here, this embodiment presents several experiments to verify the effectiveness of the proposed algorithm.

[0154] 1) Demonstration of Algorithm Convergence Process: Performance comparison analysis was conducted through simple iterations. Figure 3 shows the performance under different time steps and system costs. It can be seen that the DAPO algorithm in this embodiment outperforms other algorithms in all scenarios. Furthermore, the system cost, latency, and energy consumption of the six algorithms (DAPO, PPO, DDQN, DQN, A3C, and SAC) gradually decrease over time and eventually stabilize, while the random algorithm exhibits smaller fluctuations. In the early iteration stages, the system cost has not yet converged, and the system remains in a state of continuous fluctuation. As the number of iterations increases, the PPO algorithm begins to converge at approximately 1200 iterations. The DAPO algorithm also begins to converge at approximately 1500 iterations.

[0155] After 2000 iterations, the DDQN, DQN, A3C, and SAC algorithms began to converge. This is because, with deeper training, the agent accumulated more experience and knowledge, and the balance between exploration and exploitation became more reasonable, gradually forming a stable system reward. Meanwhile, the PPO and DAPO algorithms converged faster than DDQN, DQN, A3C, and SAC, and ultimately had lower system costs, making them more suitable for the scenario described in this paper. In contrast, the system cost of the stochastic algorithms remained consistently high. Figures 3(b) and 3(c) show the time-varying performance of different methods in terms of latency and energy consumption. It can be seen that, except for the stochastic algorithms, the latency and energy consumption of the other methods decreased.

[0156] This trend is particularly evident during the iterative process. Notably, the DAPO algorithm outperforms other algorithms in all aspects.

[0157] In terms of latency, the DAPO algorithm can reduce latency.

[0158] Compared to PPO, latency was reduced by 16.9%; compared to DDQN, latency was reduced by 52.1%; and compared to DQN, latency was reduced by 79.2%. The performance of the A3C and SAC algorithms was comparable to PPO, but still lagged behind DAPO. In terms of energy consumption, the DAPO algorithm reduced energy consumption by 52.1% compared to PPO and 86% compared to DDQN. Compared to DDQN, energy consumption was 1% lower, and compared to DQN, it was 87.5% lower, and it also significantly outperformed the A3C and SAC algorithms.

[0159] 2) Evaluation of the DAPO Algorithm: Performance Comparison Based on Data Size. Figures 4(a)-(c) show the comparison results of the seven algorithms in terms of total cost, energy consumption, and latency. By adjusting the data block size, we can control the average memory usage of the output task data. As memory usage increases, the required computing resources also increase. The results show that the DAPO algorithm significantly outperforms other methods in terms of total cost, latency, and energy consumption. The PPO algorithm performs second best, with relatively low performance growth as the data size increases. The A3C and SAC algorithms show moderate performance, but their efficiency decreases faster than DAPO and PPO as the data size increases. The DDQN and DQN algorithms show significant performance degradation with large data sizes, especially in terms of energy consumption. The random algorithm performs the worst in all metrics. In addition, the DAPO algorithm also shows higher stability compared to other algorithms, maintaining consistent performance even when the data size changes, with only a small increase in system cost, latency, and energy consumption even at the maximum data size tested.

[0160] Based on a performance comparison analysis of the maximum transmission power of the MSD, the impact of the maximum transmission power of the MSD on the offloading cost was further explored. The relevant results are shown in Figures 5(a) to (c). Key observations include:

[0161] (1) The DAPO algorithm consistently outperforms other algorithms in terms of total cost and energy consumption.

[0162] (2) The A3C and SAC algorithms exhibit moderate performance, with SAC showing higher stability under different transmission power levels. (3) The conservative offloading strategy adopted minimizes the impact of DAPO on the total cost when the maximum transmission power changes. This strategy not only effectively adapts to different offloading conditions but also enhances data privacy protection. (4) For the random algorithm, DDQN algorithm, and DQN algorithm, although the delay is largely unaffected by the increase in transmission power, their energy consumption fluctuates greatly, indicating that in these scenarios, energy consumption has a greater impact on the overall cost than delay.

[0163] Performance comparison analysis based on MSD computing power. The impact of MSD computing power on offloading costs was further explored; the relevant results are detailed in Figures 6(a) to (c). Key findings include:

[0164] (1) The DAPO algorithm is consistently superior to other algorithms;

[0165] (2) Regarding load balancing strategies, as the computing power of mobile devices decreases, more tasks will be transferred to ECs (edge ​​computing devices) with more powerful computing resources, thereby reducing operating costs. However, this also increases the risk of data privacy breaches. Therefore, in practical applications, it is recommended to improve the computing power of the MSD in system design to mitigate privacy risks.

[0166] (3) Revealing the intrinsic mechanism of DAPO: By calculating the value loss and action loss of the PPO and DAPO algorithms, as shown in Figure 7, after approximately 450 rounds, the value of the DAPO algorithm is lower than that of the PPO algorithm. This indicates that the DAPO algorithm outperforms the PPO algorithm. Lower and more stable loss generally implies better learning and optimization.

[0167] (4) Impact of Information Entropy: The system cost before and after introducing information entropy was analyzed for these six algorithms. As shown in Figure 8, the DAPO and PPO algorithms are not sensitive to the introduction of information entropy; their performance remains good even after entropy is introduced. A3C exhibits good adaptability similar to PPO, quickly converging to the optimal value even with increased complexity. SAC maintains relatively stable performance with and without information entropy, but its system cost is slightly higher than DAPO, PPO, and A3C. However, DDQN and DQN show drastically different performances—their performance deteriorates significantly after the introduction of information entropy.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A task scheduling optimization method based on a dual-agent strategy in a computing network environment, characterized in that, include: Based on the communication model, obtain the energy consumption and latency during the offloading process of the task from the local processor MSDn to the edge server ECs; Optimization objectives for task unloading are constructed based on a computational model; Based on the optimization objective, the DAPO algorithm based on a dual-agent strategy is used to obtain the optimal offloading strategy for cloud-edge-device collaborative tasks.

2. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The communication model includes: Calculate the bandwidth allocated to MSDn by the edge servers (ECs); The channel capacity is calculated using Shannon's formula, and then the data transmission rate between the local processor MSDn and the edge servers ECs is obtained. Construct a privacy entropy model, using a privacy coefficient to represent the uncertainty of task unloading; The probability of task offloading between the local processor MSDn and edge servers ECs is calculated and constrained; Based on the data transmission rate and the privacy factor, calculate the transmission latency from the local processor MSDn to the edge server ECs, the time consumed when the task is executed locally, the time required to execute the computing task in the edge server ECs, the task execution capability of the local processor MSDn, the energy consumption of transferring the task from the local processor MSDn to the edge server ECs, and the energy consumption of task execution on the edge server.

3. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The total latency for task transfer and computation on the local processor MSDn and edge servers ECs is shown below: , in, This represents the total latency of task transmission and computation. This indicates the time consumed when the task is executed locally. This indicates the time required for the task to execute on the edge servers (ECs). This indicates the transmission latency of a task from the local processor MSDn to the edge servers ECs.

4. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The total energy consumption for task transfer and computation on the local processor MSDn and edge servers ECs is shown below: , in, Indicates total energy consumption. This indicates the power consumption of the local processor during task execution. This indicates the energy consumption of edge server task execution. This indicates the energy consumption when a task is transferred from the local processor MSDn to the edge server ECs.

5. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The computational model includes CPU cycle frequency constraints on the local processor MSDn, computational resource constraints on the edge processors, communication bandwidth limits allocated to the local processor MSDn, wireless channel power constraints on the local processor MSDn, and sub-channel constraints on the edge servers ECs.

6. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The optimization objective is shown in the following formula: , Where Q(m,n,t) represents the weighted sum of energy consumption and delay, and C total This indicates the optimization objective.

7. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The DAPO algorithm based on the dual-proxy strategy includes: Policy updates are performed using two agents that interact with the same environment but employ different policy update mechanisms, with the policy update based on the agent with the better parameters.

8. The task scheduling optimization method based on a dual-agent strategy in a computing network environment according to claim 1, characterized in that, The policy update process of the DAPO algorithm based on the dual-agent strategy includes: Initialize proxy1 and proxy2; Agent 1 and Agent 2 interact with the environment respectively. The policy is updated based on the generated state and reward. When the current update cycle is a multiple of 100, the rewards of Agent 1 and Agent 2 are compared. The policy parameters of the other agent are updated based on the agent with the higher reward. The policy update process ends when the update process of either agent converges.

Citation Information

Cited By

  • Multi-agent-driven cloud-edge-end security task multi-target adaptive scheduling method

    CN121530763A