Task unloading optimization method based on adaptive particle swarm optimization and reinforcement learning

By combining adaptive particle swarm optimization and soft actor-critician reinforcement learning, dynamically adjusting the hyperparameters of the task offload algorithm, the problem of insufficient efficiency and adaptability of task offload decisions in the industrial Internet of Things environment is solved, and efficient, flexible and intelligent task offload optimization is achieved.

CN120104209APending Publication Date: 2025-06-06CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510171540.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In complex and dynamic industrial IoT environments, task offload decisions face dynamic changes in devices and servers, task complexity and limitations of optimization algorithms, resulting in insufficient efficiency and adaptability of prior art.

Method used

Combining adaptive particle swarm optimization (APSO) and soft actor-criticist (SAC) reinforcement learning, dynamically adjust the hyperparameters of particle swarm optimization algorithm, and optimize task offload decisions through real-time performance feedback and reward mechanisms.

Benefits of technology

It significantly improves the efficiency, adaptability and intelligence of task offload decisions, reduces task processing delays and energy consumption, optimizes the utilization of computing resources, and ensures that the system operates stably in a dynamic environment for a long time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005274195760000031
    Figure BDA0005274195760000031
  • Figure BDA0005274195760000041
    Figure BDA0005274195760000041
  • Figure BDA0005274195760000061
    Figure BDA0005274195760000061
Patent Text Reader

Abstract

The invention discloses a task unloading optimization method based on adaptive particle swarm optimization and reinforcement learning, and aims to improve the task unloading decision efficiency and the computing resource utilization rate in an industrial internet of things (IIOT) system. According to the method, through dynamic optimization of an unloading strategy and hyper-parameter adjustment, the flexibility and adaptability of task unloading are improved; firstly, a particle swarm optimization (PSO) algorithm is adopted to initialize a task unloading decision, and hyper-parameters are dynamically adjusted to adapt to different optimization stages; secondly, a soft actor-commentator (SAC) reinforcement learning algorithm is combined, and hyper-parameters in particle swarm optimization are finely adjusted according to real-time feedback, so that the precision of task unloading decision and the adaptive capacity of the system are effectively improved; then, dynamically adjusting a search strategy of a particle swarm based on a self-adaptive particle swarm optimization mechanism so as to optimize a task unloading path and maximize the utilization rate of computing resources; and an optimization mechanism based on real-time performance feedback is further adopted to monitor an unloading execution effect and dynamically adjust a task unloading strategy so as to ensure that the system stably and efficiently runs for a long time in a complex and dynamic IIOT environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of intelligent information processing and relates to a task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning. Background Art

[0002] With the development of industrial Internet of Things and mobile edge computing, the amount of data generated by devices and sensors has increased exponentially, which has brought great challenges to real-time task offloading and data processing. Mobile edge computing can reduce latency and bandwidth consumption by pushing computing tasks to network edge servers, and is an effective technology to solve this problem. However, in complex and dynamic industrial Internet of Things environments, task offloading decisions face many challenges: 1) Dynamic changes in devices and servers: In practical applications, the status of devices and networks as well as server loads change at any time, which makes task offloading decisions not only dependent on static geographic locations or server selections; 2) Complexity of task offloading: Factors such as the type of task, computing requirements, and network bandwidth have an important impact on offloading decisions. Traditional offloading methods, such as always selecting the nearest server, usually cannot provide the best solution in all cases; 3) Limitations of optimization algorithms: Particle swarm optimization (PSO) and reinforcement learning (RL) perform well in specific scenarios, but their performance is limited when dealing with dynamic environments and complex variables. For example, the fixed hyperparameters of particle swarm optimization lead to its poor adaptability in dynamic environments. Although reinforcement learning algorithms can dynamically adjust decisions, they usually require a large amount of training data and high computational costs.

[0003] Therefore, how to achieve efficient and flexible task offloading in the dynamic environment of IIoT has become a technical problem that needs to be solved urgently. This paper combines adaptive particle swarm optimization (APSO) with soft actor-critic (SAC) reinforcement learning to propose an efficient task offloading optimization method, aiming to overcome the limitations of existing technologies and improve the efficiency, adaptability and intelligence level of task offloading. Summary of the invention

[0004] In order to overcome the above defects, this application proposes a task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning. The specific steps of this application are as follows:

[0005] S1, initialize the particle swarm and particle swarm optimization parameters, create a particle swarm and initialize the characteristics of each particle (position, speed, personal best solution, etc.);

[0006] S2, particle update and position adjustment, searching the solution space by updating the particle speed and position;

[0007] S3, adaptive particle swarm optimization adjusts hyperparameters, dynamically adjusts the hyperparameters of the particle swarm to adapt to different optimization stages;

[0008] S4, soft actor-critic reinforcement learning fine-tuning hyperparameters, adaptively adjusts the hyperparameters of particle swarm optimization through reinforcement learning algorithm to improve the flexibility of the optimization process;

[0009] S5, performance feedback and optimization, monitors the offloading execution effect, collects performance data, optimizes system parameters through feedback mechanism, and ensures the long-term stable and accurate operation of task offloading strategy;

[0010] S6,Evaluate the efficiency of task offloading, and evaluate the performance of the optimization algorithm by calculating the cost and delay of task offloading.

[0011] The technical features and improvements of this application are:

[0012] For step S1, the present application uses the particle swarm optimization (PSO) algorithm for the preliminary stage of task offloading decision-making. The core task of this step is to initialize the particle swarm and its related hyperparameters. The particle swarm optimization algorithm essentially simulates multiple "particles" exploring in the solution space to find the global optimal solution; in this process, each particle represents a potential task offloading solution, and its position, speed, personal best solution and current fitness will jointly affect the optimization process; the initial position xi(0) and speed vi(0) of each particle are randomly initialized in the solution space to ensure that the particle swarm covers the entire solution space, thereby avoiding falling into the local optimum. The speed of the particle determines the amplitude of the particle's movement in the solution space, while its position indicates the choice of the current task offloading decision. The initialization formulas for the position and speed of the particle are as follows:

[0013] x i (0)=U(a,b), v i (0) = U(c, d) (1)

[0014] Among them, U(a, b) and U(c, d) are uniform random distribution functions, which are used to generate the initial values ​​of position and speed respectively. These initial values ​​enable the particle swarm to fully cover the possible solution space, ensuring that the algorithm has sufficient exploration ability. In addition, to guide the optimization process, the particle's personal best solution pbesti(0) is set to the fitness optimal solution corresponding to the particle's initial position, and the global best solution gbest(0) is the solution with the highest fitness in the particle swarm, representing the optimal task offloading decision of the entire particle swarm. In the initialization stage of the particle swarm, the initial global best solution is determined by calculating the fitness of each particle position, providing a diverse solution space for subsequent task offloading optimization, enabling the particle swarm to efficiently search in the solution space, laying the foundation for particle update and hyperparameter adjustment.

[0015] For step S2, the present application promotes the optimization process by continuously updating the speed and position of the particles. The particle update process is the core of the particle swarm optimization algorithm, which mainly adjusts the position of the particles in the solution space by calculating the speed change of the particles. The update process follows the following speed update formula:

[0016] v i (t+1)=w t ·v i (t)+c 1 ·U(a,b)·(pbest i (t)-x i (t))+c 2 ·U(a,b)·(gbest(t)-x i (t)) (2)

[0017] Among them, v i (t) represents the velocity of particle i at time step t, w t is the inertia weight, which determines the influence of the particle's previous velocity; c 1 and c 2 is the acceleration coefficient, which controls the speed at which the particle converges to its personal best solution pbesti(t) and the global best solution gbest(t); and U(a,b) is a uniformly distributed random number, which is used to increase the randomness of the solution space exploration and avoid premature convergence; the speed update formula comprehensively considers the cognitive part of particle self-adjustment and the social part of group collaboration, so that particles can effectively explore in the solution space. The position update of the particle is based on its current speed. The formula is as follows:

[0018] x i (t+1)=x i (t)+v i (t+1) (3)

[0019] Through this update mechanism, particles can move in the solution space and continuously optimize the current task offloading decision plan, thereby approaching the global optimal solution; the particle update process is carried out step by step, and each iteration will adjust the particle's solution space position according to the current speed and position. Through continuous updating, the particle swarm gradually converges to the optimal solution. This process will continue until the preset number of iterations or convergence conditions are met.

[0020] For step S3, the present application considers that the performance of the particle swarm optimization algorithm depends largely on the setting of its hyperparameters, especially the acceleration coefficient c. 1 、c 2 and the inertia weight w t, these hyperparameters directly affect the convergence speed of particles and the search efficiency of the global optimal solution; however, in a dynamic environment, fixed hyperparameter settings often cannot fully cope with the complexity and variability of the task offloading decision process. Therefore, this application proposes an adaptive particle swarm optimization mechanism, which adjusts these hyperparameters in real time to adapt to environmental changes, thereby significantly improving the adaptability and overall optimization efficiency of the particle swarm optimization algorithm. By dynamically adjusting the hyperparameters, the adaptive particle swarm optimization can flexibly balance the exploration and development capabilities of particles at different optimization stages, thereby ensuring better optimization results in complex and dynamic task offloading scenarios; in the adaptive particle swarm optimization, the hyperparameter c 1 、c 2 and w t It will be adjusted according to the search state of the particle swarm. By introducing the evolution factor to describe the distribution state of the particle swarm, the hyperparameter settings can be dynamically evaluated and changed. The evolution factor reflects the distribution of the particle swarm in the solution space. It can adjust the behavior of the algorithm according to the different search stages. Specifically, in the exploration stage, the particle swarm tends to search the solution space in a larger range. At this time, c can be increased. 1 and c 2 In order to promote the expansion of particles to a wider solution space. In the development stage, the particle swarm has gradually converged to a certain area, and c can be reduced. 1 and c 2 , to refine the search in the solution space; the inertia weight w t It is also dynamically adjusted according to the feedback of the evolution factor. Usually, a higher inertia weight is set in the early stage to promote the extensive search of particles, and the inertia weight is gradually reduced in the later stage so that particles can converge to the optimal solution more accurately. The specific formula is as follows:

[0021]

[0022] Among them, w 0 is the initial inertia weight, T is the maximum number of iterations. As the iteration proceeds, the inertia weight gradually decreases, thereby guiding the particles to gradually converge to the global optimal solution. The adaptive mechanism enables PSO to flexibly respond to various changes in the dynamic task offloading environment, greatly improving the effect of the optimization process.

[0023] For step S4, this application introduces the soft actor-critic (SAC) algorithm in reinforcement learning to fine-tune the hyperparameters in adaptive particle swarm optimization (APSO) to further optimize the task offloading decision; SAC is a policy-based reinforcement learning method that balances exploration and development by maximizing expected rewards and adding entropy regularization terms to achieve more efficient decision-making; SAC updates the strategy by optimizing the objective function, which considers both immediate rewards and the entropy term of the strategy, thereby encouraging more diverse decision choices to avoid premature convergence of the strategy. The objective function of SAC is as follows:

[0024]

[0025] Among them, R(s t ,a t ,s t+1 ) is the reward function, indicating that in state s t Take action a t Then transfer to state s t+1 The reward obtained; γ is the discount factor used to weigh the long-term reward; H(π(·|s t )) is the strategy entropy, which indicates the diversity of the current strategy, and α is the entropy coefficient that controls the balance between exploration and development; in this way, SAC can optimize the hyperparameters of the particle swarm (such as the acceleration coefficient c) according to the current optimization state. 1 、c 2 and the inertia weight w t ) to improve the adaptability and optimization efficiency of the particle swarm optimization algorithm in a dynamic environment; SAC automatically adjusts these hyperparameters after each iteration to ensure that the particle swarm can efficiently search for the optimal solution in the solution space while avoiding over-exploration or premature convergence.

[0026] The task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning in this application solves the problem of insufficient efficiency and adaptability of task offloading decision-making in dynamic environments in the prior art, and has the following advantages:

[0027] (1) The adaptive particle swarm optimization and reinforcement learning combined method proposed in this application can adaptively optimize task offloading decisions according to environmental changes by dynamically adjusting the hyperparameters in the PSO algorithm, significantly improving the adaptability of the algorithm in complex and dynamic environments. Unlike the traditional fixed hyperparameter setting, this method can flexibly adjust parameters at different stages of the task offloading process, effectively improving the ability to search for the global optimal solution, thereby significantly reducing task processing delays and energy consumption, and improving the utilization of edge computing resources;

[0028] (2) This application introduces a soft actor-critic (SAC) reinforcement learning mechanism to intelligently fine-tune the hyperparameters in the task offloading decision based on real-time performance feedback and reward mechanism. The entropy regularization property of SAC further enhances the diversity of strategies, avoids the problem of premature convergence, and ensures the stability and efficiency of the task offloading process. Through this adaptive fine-tuning mechanism, the system can flexibly respond to changes in different task types, network status, and computing resources, improve task offloading efficiency, and reduce the computing cost of the overall system;

[0029] (3) The task offloading optimization method based on APSO and SAC proposed in this application, by integrating the long-term reward optimization mechanism of reinforcement learning and the global search capability of particle swarm optimization, enables the system to make accurate task offloading decisions in a dynamically changing industrial Internet of Things environment. By adjusting the search strategy and exploration range of particles in real time, the system can always maintain efficient task offloading path planning in a changing network environment, maximize the success rate of task offloading, and ensure the long-term stable operation of the system in highly dynamic and complex industrial Internet of Things scenarios;

[0030] (4) This application combines adaptive particle swarm optimization and reinforcement learning mechanisms to dynamically adjust the task offloading path and server selection strategy. It can intelligently optimize task offloading decisions based on real-time information such as edge computing node load, network conditions, and task complexity, effectively avoiding problems such as computing node overload or insufficient bandwidth. Through load balancing and resource scheduling, it ensures the optimal allocation of edge computing resources, thereby improving the overall computing and transmission efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Flowchart of the task offloading optimization architecture for adaptive particle swarm optimization and reinforcement learning in this application.

[0032] Figure 2 Diagram of the particle swarm optimization mechanism adapted for this application.

[0033] Figure 3 Figure 2. Soft actor-critic (SAC) reinforcement learning combined with particle swarm optimization (PSO) in this application. DETAILED DESCRIPTION

[0034] The present application is further described in detail below with reference to the accompanying drawings and specific implementation methods:

[0035] A task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning, such as Figure 1 As shown, this is a task offloading optimization flow chart based on adaptive particle swarm optimization and reinforcement learning of this application, and the method includes:

[0036] S1, the key to initializing the particle swarm optimization algorithm is to create a particle swarm and assign initial characteristics to each particle. The basic principle of the particle swarm optimization algorithm is to gradually search the solution space through the cooperation and competition of multiple particles to find the global optimal solution. In the task offloading decision, each particle represents a potential task offloading decision scheme. The particle has a position x i (t), speed v i (t), personal best solution pbest i (t), the global best solution characteristic gbest(t), where the position represents the current solution of the particle in the solution space, that is, the selected solution for task offloading; the speed represents the rate at which the particle moves in the solution space, which is used to control the change of the particle position; the personal best solution represents the best solution found by the particle in the self-search process, reflecting the best task offloading decision of the particle in the optimization process; the global best solution represents the solution with the best fitness in the entire particle swarm, which is the global optimal task offloading solution in the current optimization process. First, the size N of the particle swarm needs to be set to determine the number of particles, which is usually determined according to the complexity of the problem. The initial position and speed of each particle are randomly initialized to ensure that the particle swarm can cover the solution space. Position x i (t) and speed v i Initialization of (t) can be done in the following ways:

[0037] x i (0)=U(a,b), v i (0) = U(c, d) (6)

[0038] Among them, U(a, b) and U(c, d) are uniform distribution functions to ensure the randomness of the initial position and velocity of the particle. Then, the fitness of the particle is calculated and the personal best solution and the global best solution pbest are updated. i (t) and gbest(t), and prepare for the subsequent particle update process.

[0039] S2, after the particle swarm is initialized, the optimization process of task offloading decision is promoted by continuously updating the speed and position of the particles. Each particle adjusts its position according to its current speed to explore the task offloading decision space. At each time step t, the speed and position of the particle are updated according to the following rules:

[0040] The speed of each particle is updated according to its current speed, personal best solution and global best solution, and the formula is as follows:

[0041] v i (t+1)=w t ·v i (t)+c 1 ·U(a,b)·(pbest i (t)-x i(t))+c 2 ·U(a,b)·(gbest(t)-x i (t))(7)

[0042] Where: w t is the inertia weight, which controls how much the particle depends on its previous velocity; c 1 and c 2 is the acceleration coefficient, which determines the degree of convergence of the particle to the personal optimal solution and the global optimal solution; U(a,b) is a uniform random distribution, which is used to increase the randomness of the search. The particle speed update process can make the particle gradually converge to the global optimal solution. After each update, the particle position x i (t+1) Adjust according to the updated speed:

[0043] x i (t+1)=x i (t)+v i (t+1) (8)

[0044] The particle’s position update brings it closer to the global optimal solution in the solution space, thereby gradually improving the task offloading decision.

[0045] S3, in order to overcome the limitations of traditional particle swarm optimization algorithm in dynamic environment, introduces adaptive particle swarm optimization mechanism to dynamically adjust the hyperparameter c in the algorithm 1 、c 2 and w t , enabling the particle swarm to flexibly optimize according to the complexity of task offloading decisions and environmental changes. In APSO, the evolution factor is introduced to reflect the distribution state of the particle swarm. The evolution factor E(t) represents the current exploration and development state of the particle swarm and determines the dynamic adjustment of the hyperparameters. Specifically, the evolution factor evaluates whether the particle swarm is in the exploration stage, development stage, or convergence stage based on its performance, and then adjusts the value of the hyperparameters:

[0046] c 1 (t+1)=c 1 (t)+δ 1 ·E(t) (9)

[0047] c 2 (t+1)=c 2 (t)+δ 2 ·E(t) (10)

[0048]

[0049] Among them, δ 1 and δ 2 is the adjustment coefficient, w 0is the initial inertia weight, T is the maximum number of iterations, and as the iterations proceed, the inertia weight gradually decreases, so that the particles gradually converge to the optimal solution; through this adaptive mechanism, the hyperparameters can be adjusted according to the distribution state of the particle swarm, ensuring that the algorithm can maintain good performance in both the exploration and development stages, thereby effectively improving the accuracy of task offloading decisions.

[0050] S4, in order to further improve the flexibility of task offloading optimization, the present invention introduces the soft actor-critic (SAC) reinforcement learning algorithm to fine-tune the hyperparameters in the adaptive particle swarm optimization. SAC ensures the diversity of strategies during the optimization process by maximizing long-term rewards and adding entropy regularization, avoiding premature convergence, and optimizing task offloading decisions through real-time feedback. The objective function of SAC combines the entropy of rewards and strategies to achieve balanced exploration and exploitation:

[0051]

[0052] Among them, R(s t ,a t ,s t+1 ) is the reward function, indicating that in state s t Take action a t Then transfer to state s t+1 The reward obtained; γ is the discount factor used to weigh the long-term reward; H(π(·|s t )) is the policy entropy, which indicates the diversity of the current policy, and α is the entropy coefficient that controls the balance between exploration and exploitation. SAC dynamically adjusts the hyperparameters in particle swarm optimization through policy updates, so that the system can accurately adjust the hyperparameters according to the requirements of the task and the changes in the environment, thereby further improving the efficiency and adaptability of task offloading decisions.

[0053] S5, in the process of task offloading, it is very important to monitor the performance of the system in real time and adjust the offloading strategy according to the feedback. The performance data of the task offloading process is collected through the feedback mechanism, and the system parameters are continuously optimized based on this data to ensure that the task offloading decision is always accurate, stable and efficient. By monitoring the performance indicators such as the delay, energy consumption, and computing resource consumption of task execution, the system can detect the performance degradation or bottleneck in time, and adjust the hyperparameters and task offloading strategy according to the feedback. This feedback mechanism can continuously adjust the hyperparameters in the particle swarm optimization process according to the real-time data, ensuring that the system can always run efficiently under different tasks and environmental conditions.

[0054] S6, by calculating the cost and delay of task offloading, evaluates the performance of the optimization algorithm and conducts a quantitative analysis of the task offloading efficiency. The cost of task offloading includes computing resource consumption and data transmission delay, and the evaluation indicators include the total time of task execution, system response time, energy consumption, etc.; through a detailed evaluation of the performance of the optimization algorithm, ensure that the task offloading process can achieve optimal performance under different network conditions and computing environments, the optimized task offloading decision can minimize the time and energy consumption of task processing while ensuring efficient resource utilization, thereby improving the overall performance of the system.

[0055] In summary, this application proposes a task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning, which effectively improves the efficiency and accuracy of task offloading decisions by dynamically adjusting hyperparameters, combining soft actor-critic reinforcement learning algorithms, and optimizing offloading strategies in real time. This method reduces task processing delays and energy consumption by flexibly adapting to environmental changes, optimizes the utilization of computing resources, ensures that the system can operate stably and long-term in a dynamic and complex Industrial Internet of Things (IIoT) environment, and achieves efficient task offloading performance.

[0056] Although the content of the present application has been described in detail through the above preferred embodiments, it should be appreciated that the above description should not be considered as a limitation of the present application. After reading the above content, it will be apparent to those skilled in the art that various modifications and substitutions of the present application can be made. Therefore, the protection scope of the present application should be limited by the appended claims.

Claims

1. A task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning, its characteristics and The specific steps are as follows: S1, initialize the particle swarm and particle swarm optimization parameters, create a particle swarm and initialize the characteristics of each particle (position, speed, personal best solution, etc.); S2, particle update and position adjustment, searching the solution space by updating the particle speed and position; S3, adaptive particle swarm optimization adjusts hyperparameters, dynamically adjusts the hyperparameters of the particle swarm to adapt to different optimization stages; S4, soft actor-critic reinforcement learning fine-tuning hyperparameters, adaptively adjusts the hyperparameters of particle swarm optimization through reinforcement learning algorithm to improve the flexibility of the optimization process; S5, performance feedback and optimization, monitors the offloading execution effect, collects performance data, optimizes system parameters through feedback mechanism, and ensures the long-term stable and accurate operation of task offloading strategy; S6,Evaluate the efficiency of task offloading, and evaluate the performance of the optimization algorithm by calculating the cost and delay of task offloading.

2. The task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning according to claim 1, characterized in that: For step S1, the present invention uses the particle swarm optimization (PSO) algorithm in the preliminary stage of task offloading decision-making. The core task of this step is to initialize the particle swarm and its related hyperparameters. The particle swarm optimization algorithm essentially searches for the global optimal solution by simulating multiple "particles" to explore in the solution space. In this process, each particle represents a potential task offloading solution, and its position, speed, personal best solution and current fitness will jointly affect the optimization process. The initial position x of each particle i (0) and speed v i (0) Randomly initialize in the solution space to ensure that the particle swarm covers the entire solution space, thereby avoiding falling into the local optimum. The particle speed determines the movement amplitude of the particle in the solution space, while its position indicates the choice of the current task offloading decision. The initialization formula of the particle position and speed is as follows: x i (0)=U(a,b),v i (0)=U(c,d)(1) Among them, U(a, b) and U(c, d) are uniform random distribution functions, which are used to generate the initial values ​​of position and speed respectively. These initial values ​​enable the particle swarm to fully cover the possible solution space, ensuring that the algorithm has sufficient exploration ability. In addition, to guide the optimization process, the particle's personal best solution pbesti(0) is set to the fitness optimal solution corresponding to the particle's initial position, and the global best solution gbest(0) is the solution with the highest fitness in the particle swarm, representing the optimal task offloading decision of the entire particle swarm. In the initialization stage of the particle swarm, the initial global best solution is determined by calculating the fitness of each particle position, providing a diverse solution space for subsequent task offloading optimization, enabling the particle swarm to efficiently search in the solution space, laying the foundation for particle update and hyperparameter adjustment.

3. The task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning according to claim 1, characterized in that: For step S2, the present invention promotes the optimization process by continuously updating the speed and position of the particles. The particle update process is the core of the particle swarm optimization algorithm, which mainly adjusts the position of the particles in the solution space by calculating the speed change of the particles. The update process follows the following speed update formula: v i (t+1)=w t ·v i (t)+c1·U(a,b)·(pbest i (t)-x i (t))+c2·U(a,b)·(gbest(t)-x i (t))(2) Among them, v i (t) represents the velocity of particle i at time step t, w t is the inertia weight, which determines the influence of the particle's previous speed; c1 and c2 are acceleration coefficients, which control the speed at which the particle converges to its personal best solution pbesti(t) and the global best solution gbest(t); and U(a,b) is a uniformly distributed random number, which is used to increase the randomness of the solution space exploration and avoid premature convergence; the speed update formula comprehensively considers the cognitive part of particle self-adjustment and the social part of group collaboration, so that particles can effectively explore in the solution space. The position update of the particle is based on its current speed. The formula is as follows: x i (t+1)=x i (t)+v i (t+1)(3) Through this update mechanism, particles can move in the solution space and continuously optimize the current task offloading decision plan, thereby approaching the global optimal solution; the particle update process is carried out step by step, and each iteration will adjust the particle's solution space position according to the current speed and position. Through continuous updating, the particle swarm gradually converges to the optimal solution. This process will continue until the preset number of iterations or convergence conditions are met.

4. The task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning according to claim 1, characterized in that: For step S3, the present invention takes into account that the performance of the particle swarm optimization algorithm depends largely on the settings of its hyperparameters, especially the acceleration coefficients c1, c2 and the inertia weight w t , these hyperparameters directly affect the convergence speed of particles and the search efficiency of the global optimal solution; however, in a dynamic environment, fixed hyperparameter settings often cannot fully cope with the complexity and variability of the task offloading decision process. Therefore, this application proposes an adaptive particle swarm optimization mechanism, which adjusts these hyperparameters in real time to adapt to environmental changes, thereby significantly improving the adaptability and overall optimization efficiency of the particle swarm optimization algorithm. By dynamically adjusting the hyperparameters, the adaptive particle swarm optimization can flexibly balance the exploration and development capabilities of particles at different optimization stages, thereby ensuring better optimization results in complex and dynamic task offloading scenarios; in the adaptive particle swarm optimization, the hyperparameters c1, c2 and w t It will be adjusted according to the search state of the particle swarm. By introducing the evolution factor to describe the distribution state of the particle swarm, the hyperparameter settings can be dynamically evaluated and changed. The evolution factor reflects the distribution of the particle swarm in the solution space. It can adjust the behavior of the algorithm according to the different search stages. Specifically, in the exploration stage, the particle swarm tends to search the solution space in a wider range. At this time, the values ​​of c1 and c2 can be increased to promote the expansion of particles to a wider solution space. In the development stage, the particle swarm has gradually converged to a certain area. At this time, c1 and c2 can be reduced to refine the search in the solution space. The inertia weight w t It is also dynamically adjusted according to the feedback of the evolution factor. Usually, a higher inertia weight is set in the early stage to promote the extensive search of particles, and the inertia weight is gradually reduced in the later stage so that particles can converge to the optimal solution more accurately. The specific formula is as follows: Among them, w0 is the initial inertia weight, T is the maximum number of iterations, and as the iteration proceeds, the inertia weight gradually decreases, thereby guiding the particles to gradually converge to the global optimal solution; the adaptive mechanism enables PSO to flexibly respond to various changes in the dynamic task offloading environment, greatly improving the effect of the optimization process.

5. The task offloading optimization method based on adaptive particle swarm optimization and reinforcement learning according to claim 1, characterized in that: For step S4, the present invention introduces a soft actor-critic (SAC) algorithm in reinforcement learning to fine-tune the hyperparameters in adaptive particle swarm optimization (APSO) to further optimize the task offloading decision; SAC is a policy-based reinforcement learning method that balances exploration and exploitation by maximizing expected rewards and adding entropy regularization terms to achieve more efficient decision making; SAC updates the strategy by optimizing the objective function, which takes into account both the immediate reward and the entropy term of the strategy, thereby encouraging more diverse decision choices to avoid premature convergence of the strategy. The objective function of SAC is as follows: Among them, R(s t ,a t ,s t+1 ) is the reward function, indicating that in state s t Take action a t Then transfer to state s t+1 The reward obtained; γ is the discount factor used to weigh the long-term reward; H(π(·|s t )) is the strategy entropy, which indicates the diversity of the current strategy, and α is the entropy coefficient that controls the balance between exploration and development; in this way, SAC can optimize the hyperparameters of the particle swarm (such as acceleration coefficients c1, c2 and inertia weight w) according to the current optimization state. t ) to improve the adaptability and optimization efficiency of the particle swarm optimization algorithm in a dynamic environment; SAC automatically adjusts these hyperparameters after each iteration to ensure that the particle swarm can efficiently search for the optimal solution in the solution space while avoiding over-exploration or premature convergence.

Citation Information

Cited By

  • Vehicle dynamic path planning method based on Harris eagle fusion deep reinforcement learning

    CN120333489A

  • Multi-terminal large model compression strategy selection method based on swarm intelligence optimization

    CN121683898A