Underwater Computing Task Offloading Method Based on Improved Deep Deterministic Policy Gradient

Through the improved depth deterministic strategy gradient algorithm, combined with the adaptive update amplitude method, the problems of high energy consumption and slow convergence speed of underwater computing tasks in traditional methods are solved, and faster convergence and lower energy consumption are achieved.

CN119255300BActive Publication Date: 2025-06-10NANJING UNIV

Patent Information

Application Number
CN202411768075.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-06-10
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

When traditional reinforcement learning is used to unload underwater computing tasks in the marine Internet of Things (OoT), there are problems such as high energy consumption and slow algorithm convergence.

Method used

The improved Deep Deterministic Policy Gradient (DDPG) algorithm is used to train by minimizing the average loss function, and an adaptive update amplitude method is used during the parameter update process to speed up the convergence speed of the algorithm and reduce energy consumption.

Benefits of technology

The improved DDPG algorithm significantly accelerates convergence speed and reduces the offload energy consumption of computing tasks for underwater sensor nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255300B_ABST
    Figure CN119255300B_ABST
Patent Text Reader

Abstract

The present invention discloses an underwater computing task offloading method based on an improved deep deterministic policy gradient. First, the state space of the task offloading problem is initialized, the action space is constructed by adding noise to the offloading mode and offloading quantity matrix, and the reward function is designed using the reciprocal of the task offloading energy consumption. Secondly, the deep deterministic policy gradient algorithm is improved, and the adaptive update amplitude method is used in the process of updating the algorithm parameters. The update amplitude is set to a relatively high value in the initial stage and gradually reduced during the training of the algorithm. Finally, the task offloading problem is solved by the improved deep deterministic policy gradient. The Markov decision model of the task offloading problem is input, and continuous training is iterated until the average loss function no longer decreases, and the task offloading mode and offloading quantity are output. The present invention can effectively accelerate the convergence speed of the algorithm and reduce the energy consumption of sensor node task offloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge computing, and more specifically, relates to an underwater computing task offloading method based on improved deep deterministic policy gradient. Background Art

[0002] The ocean covers more than 71% of the Earth's surface area. Building an Ocean of things (OoT) for three-dimensional observation of the ocean is becoming increasingly important. The networking architecture of OoT generally includes two layers: the sea surface and the underwater layer. The sea surface layer consists of buoys, and the underwater layer consists of various sensor nodes. The sensor nodes transmit the collected data to the buoy nodes, and then the buoy nodes process the data and send it to the onshore monitoring center. With the development of OoT, the coverage area of OoT is getting larger and larger, extending from coastal waters to the open ocean. At the same time, the computing services of sensor nodes in the network, such as data analysis and running network protocols, have increased sharply. The sensor nodes consume a large amount of energy to process these computing services. The traditional two-layer networking architecture of sea surface - underwater in OoT can no longer meet the requirements. Using satellite nodes to build a new three-layer space - sea surface - underwater OoT architecture and through mobile edge computing (MEC) task offloading technology, offloading the computing tasks generated by underwater sensor node data analysis and running network protocols to buoy or satellite edge servers for processing has become a research hotspot.

[0003] Using traditional reinforcement learning algorithms to solve the computing task offloading problem is a common method currently. In such methods, the reinforcement learning model can be deployed on the buoy nodes for training. After training, the computing task offloading strategy is sent to the sensor nodes within the communication range of the buoy. The sensor nodes can offload their own computing tasks to the relevant edge servers according to the offloading strategy. However, the action space of the underwater computing task offloading problem cannot be represented as a simple binary problem and needs to be further decomposed. Moreover, the influence of energy consumption is not considered in the design of the reward function of the reinforcement learning model, which easily leads to high energy consumption in computing task offloading. It should be noted that since satellite nodes and buoy nodes are usually equipped with solar panels for solar charging, while underwater sensor nodes usually cannot directly replenish energy underwater under the existing conditions and need to save power as much as possible during operation to maintain a longer working life cycle. Therefore, when performing underwater sensor task offloading, it is necessary to minimize the energy consumption overhead, including local computing energy consumption, the energy consumption of sending tasks to the buoy nodes and receiving the returned results. When the task needs to be offloaded, the underwater sensor nodes offload the data to the buoy nodes, and the buoy nodes process the tasks and return the results. If the buoy nodes are busy and unable to process, the computing tasks will be further offloaded to the satellite nodes for processing.

[0004] The applicant Tongji University discloses a collaborative task offloading method for edge computing based on multi-agent reinforcement learning in its patent application document "A Collaborative Task Offloading Method for Edge Computing Based on Multi-Agent Reinforcement Learning" (application date: May 30, 2024, application number: 202410687523.1, application publication number: CN118585263 A, and the content of this application can still be cited). The collaborative task offloading method for edge computing has the following deficiencies: the action space of its computing task offloading problem is represented as a simple 0 or 1 problem, that is, a binary problem, which cannot accurately describe the actions of computing task offloading; and the influence of energy consumption is not considered in the design of the reward function, resulting in a relatively high energy consumption for computing task offloading.

[0005] Using the Deep Deterministic Policy Gradient (DDPG) algorithm to solve the computing task offloading problem is an effective method to reduce the energy consumption of task offloading in the MEC system. The applicant Nanjing University of Science and Technology discloses a satellite edge computing task offloading and resource allocation method based on deep reinforcement learning in its patent application document "Satellite Edge Computing Task Offloading and Resource Allocation Method Based on Deep Reinforcement Learning" (application date: May 24, 2024, application number: 202410655864.0, application publication number: CN118250750 A, and the content of this application can still be cited). The satellite edge computing task offloading and resource allocation method has the following deficiencies: the deep reinforcement learning used is the DDPG algorithm. When using the DDPG algorithm for computing task offloading, the update amplitude of the DDPG algorithm uses a fixed value during training, resulting in a slow convergence speed of the algorithm when the computing task offloading volume is large, and the long-term training further exacerbates the energy consumption problem.

[0006] In summary, traditional reinforcement learning-based methods for OoT have problems of high energy consumption for computing task offloading and slow algorithm convergence speed. Summary of the Invention

[0007] Object of the Invention: When traditional reinforcement learning-based methods are used to solve the computing task offloading problem, the action space is represented as a simple binary problem, lacking randomness; the influence of energy consumption is not considered in the design of the reward function, resulting in high energy consumption for computing task offloading; the DDPG algorithm uses a fixed update amplitude during training, resulting in a slow convergence speed. The present invention proposes an underwater computing task offloading method based on improved deep deterministic policy gradient to make up for the above technical deficiencies.

[0008] To achieve the above object of the invention, the present invention adopts the following technical solutions to be realized:

[0009] An underwater computing task offloading method based on improved deep deterministic policy gradient, comprising the following steps:

[0010] S1: Establishing the Markov decision model for computing task offloading: Describe the underwater computing task offloading problem as a Markov decision process, initialize the state space of the task offloading problem, construct the action space by adding noise to the offloading mode and offloading quantity matrix, and design the reward function using the reciprocal of the energy consumption of computing task offloading;

[0011] S2: Improving the Deep Deterministic Policy Gradient algorithm: First, train by minimizing the average loss function instead of the traditional minimizing loss function. Second, in the process of updating the parameters of the current value function and the parameters of the target value function, use the adaptive update amplitude method to replace the traditional fixed update amplitude method to accelerate the convergence speed of the algorithm;

[0012] S3: Solving the computing task offloading problem by improving the Deep Deterministic Policy Gradient algorithm: Input the Markov decision model of the computing task offloading problem in step S1 into the improved Deep Deterministic Policy Gradient algorithm in S2, and continuously train the improved Deep Deterministic Policy Gradient algorithm in S2 until the average loss function no longer decreases, and output the computing task offloading mode and offloading quantity.

[0013] Furthermore, the specific steps of step S1 are as follows:

[0014] S1-1: Initializing the state space of the task offloading problem:

[0015] Deploy a three-layer network architecture of space-sea-underwater, including low-earth orbit satellites, denoted as ; buoy nodes, denoted as ; Each buoy node provides computing services for underwater sensor nodes, denoted as ; The th buoy node and the th underwater sensor node it manages are respectively denoted as and ; Assume that the low-earth orbit satellites and buoy nodes are both built-in with edge servers capable of data analysis and computing. The buoy nodes are also equipped with acoustic modules and electromagnetic wave modules, capable of communicating with underwater sensor nodes and satellite nodes simultaneously; The underwater sensor nodes process the generated computing tasks locally, or offload the computing tasks to the edge servers of buoy nodes or satellite nodes for computing. The computing tasks generated by the underwater sensor nodes include data analysis and running network protocols; Let and respectively represent the total amount of computing tasks owned by the underwater sensor nodes and the offloading quantity of the computing tasks; When it represents The computing task needs to be offloaded to the buoy node or satellite node, otherwise the computing task is locally processed by the underwater sensor node; define to indicate whether the task is offloaded to the buoy node or satellite node, expressed as:

[0016] (1),

[0017] The underwater sensor nodes form a network underwater. Each underwater sensor node is regarded as an agent. At the initial stage of the time slot , the agent will perceive the environmental information. The state at the time slot consists of the states of all agents in the network. The state space set is expressed as:

[0018] (2),

[0019] where represents the state of the th agent at the time slot and includes the following parts:

[0020] (3),

[0021] where represents the link state information, represents the geographical location information, represents the resource allocation information, represents the current computing task information, represents the computing resource information, represents the remaining energy information;

[0022] S1-2: Construct the action space by adding noise to the offloading mode matrix and offloading quantity matrix:

[0023] Let R Pol =[ r 1 Pol ,..., r i Pol ,..., r I Pol ] Τ represent the offloading mode matrix, and there is r i Pol =[u n i,1 ,u n i, j ,...,u n i, J ] ; R Vol =[ r 1 Vol ,..., r i Vol ,..., r I Vol ] Τ represents the offloading quantity matrix, and there is r i Vol =[ g i,1 , g i, j ,..., g i, J ] ; Define the offloading matrix as multiplied by the corresponding elements of , then the offloading matrix is expressed as:

[0024] (4),

[0025] The action space consists of the offloading matrix , where indicates 's computing task offloading mode, g i, j ∈ [0, G i, j ] indicates 's computing task offloading quantity, and when , it means the task is processed locally, and the corresponding position element of the offloading matrix is marked as a negative value, that is , , represents the total computing task volume owned by the th underwater sensor node managed by the th buoy node; The original action space set at time slot is expressed as:

[0026] (5),

[0027] where represents the th action that can be selected by the th at time slot , which is formed by splicing each row of the offloading matrix , is expressed as:

[0028] (6),

[0029] To increase the randomness of the action space, random noise is added to the original action space , and the final action space is expressed as:

[0030] (7),

[0031] where represents the original action space, represents random noise;

[0032] S1-3: Design the reward function using the reciprocal of the energy consumption of computing task offloading:

[0033] Using the reciprocal of the energy consumption of computing task offloading as the reward to guide the agent to gradually reduce the energy consumption of computing task offloading, the calculation formula of the reward function is:

[0034] (8),

[0035] where represents the energy consumption of processing computing tasks locally within the time slot within , represents the energy consumption of processing tasks on the buoy node within the time slot , represents the energy consumption of processing tasks on the satellite node within the time slot , and respectively represent the additional rewards for the buoy node service and the edge server service. When the buoy node is busy, is negative, and vice versa is positive; when the edge server is busy, is negative, and vice versa is positive.

[0036] Furthermore, the specific steps of step S2 are as follows:

[0037] S2-1: Calculate the current value function value of the improved deep deterministic policy gradient and perform algorithm training by minimizing the average loss function:

[0038] The calculation formula of the current value function of the improved deep deterministic policy gradient is:

[0039] (9),

[0040] Among them, represents the parameter of the current value function; represents the discount factor, whose value range is between 0 and 1, and the closer the value is to 1, the more the future value is emphasized; represents the current policy function; represents the reward function; represents the current value function of the next state, and the target value function is calculated as:

[0041] (10),

[0042] Among them, represents the parameter of the target value function; represents the target policy function; represents the target value function of the next state; the algorithm is trained by minimizing the average loss function, and the average loss function is calculated as:

[0043] (11),

[0044] Among them, represents the number of all agents in the network;

[0045] S2-2: Improve the parameters of the DDPG algorithm and use the adaptive update amplitude method during the update process. Set the update amplitude to a relatively large value in the initial stage and gradually reduce it during the training process of the improved DDPG algorithm:

[0046] Update the parameters in step S2-1 using the adaptive update amplitude method and , and the calculation formula is:

[0047] (12),

[0048] Among them, is the update amplitude, which is between 0 and 1, and the smaller the value, the slower the algorithm update; to stabilize the training process and ensure faster convergence while ensuring the stable update of the improved DDPG algorithm, set the update amplitude to a relatively large value in the initial stage of the improved DDPG algorithm training and gradually reduce it during the training process of the improved DDPG algorithm. The calculation formula is:

[0049] (13),

[0050] Among them, and represent the initial and later update amplitudes respectively. According to formula (13), the value of the adaptive update amplitude will gradually change from to . Once or during the change process, it will immediately switch to ; is the control threshold for the update amplitude size. After multiple experiments, it is determined that the value range is between 0.5 and 2.5. The smaller the value, the smaller the initial update amplitude; represents the switching time threshold. After multiple experiments, it is determined that the value range is between 1 and 10. The larger the value, the longer the update amplitude switching time.

[0051] Furthermore, in the step S3, the improved deep deterministic policy gradient algorithm is used to solve the computing task offloading problem, specifically as follows:

[0052] Solve the computing task offloading problem through the improved deep deterministic policy gradient algorithm in S2, and input the Markov decision model of the computing task offloading problem in S1; according to the current state and the current policy function select the action , calculate the current value function value of the deep deterministic policy gradient algorithm; according to the current value function value and the target value function value, calculate the average loss function, and update the parameters and ; repeat the above process, continuously reduce the average loss function and accumulate the reward function , until the average loss function no longer decreases, and output the computing offloading mode and the computing offloading quantity .

[0053] The advantages and technical effects of the present invention are as follows:

[0054] After simulation verification, in the present invention, the improved DDPG algorithm has a significantly faster algorithm convergence speed compared with the method of the fixed update amplitude in the traditional DDPG algorithm; compared with the Deep Q-Network (DQN) algorithm, the computing task offloading energy consumption of underwater sensor nodes is significantly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is the overall flowchart of an embodiment of the present invention;

[0056] Figure 2 It is the network architecture diagram of space - sea - underwater of an embodiment of the present invention;

[0057] Figure 3 It is the schematic diagram of the classification of underwater computing task offloading modes of an embodiment of the present invention;

[0058] Figure 4 It is the architecture diagram of underwater computing task offloading of an embodiment of the present invention;

[0059] Figure 5 It is the comparison simulation diagram of the loss function varying with the number of training rounds between the method of the present invention and the traditional fixed update amplitude method in an embodiment of the present invention;

[0060] Figure 6 It is the comparison simulation diagram of the energy consumption of computing task offloading varying with the task volume between the method of the present invention and the DQN method in an embodiment of the present invention. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0062] This embodiment proposes an underwater computing task offloading method based on improved deep deterministic policy gradient, and its overall flowchart is as Figure 1 shown, including the following steps:

[0063] S1: Establishing the Markov decision model for computing task offloading: Describing the underwater computing task offloading problem as a Markov decision process, initializing the state space of the task offloading problem, constructing the action space by adding noise to the offloading mode and offloading quantity matrix, and designing the reward function using the reciprocal of the energy consumption of computing task offloading. The specific steps are as follows:

[0064] S1 - 1: Initializing the state space of the task offloading problem:

[0065] As Figure 2 shown, deploying a three - layer network architecture of space - sea - underwater, including 2 satellite nodes, denoted as , 3 buoy nodes, denoted as , and each buoy node provides computing services for 3 underwater sensor nodes, denoted as , and all underwater sensor nodes are evenly distributed in the 1000×1000×1000 m 3 area. The th buoy node and the The underwater sensor nodes are respectively denoted as and , where and ; both the satellite node and the buoy node are built-in with edge servers, which can perform data analysis and calculation. The buoy node is also equipped with an acoustic module and an electromagnetic wave module, and can communicate with the underwater sensor nodes and the satellite node simultaneously. The schematic diagram of the underwater computing task offloading mode is as shown in Figure 3 . The underwater sensor nodes can process the generated computing tasks locally, such as data analysis and running network protocols, or offload the computing tasks to the edge servers of the buoy node or the satellite node for calculation. Let and respectively represent the total amount of computing tasks owned by the underwater sensor nodes and the offloading quantity of the computing tasks. When , it means that the computing tasks of need to be offloaded to the buoy node or the satellite node, otherwise it means that the computing tasks of are processed locally by the underwater sensor nodes. Define to indicate whether the task is offloaded to the buoy node or the satellite node, which is expressed as:

[0066] (1),

[0067] It should be noted that when is always 1, the task will only be offloaded to the buoy node. At this time, the three-layer network architecture of space-sea-underwater will become a two-layer network architecture of sea-underwater. Therefore, the present invention is also applicable to the sea-underwater two-layer architecture.

[0068] The underwater sensor nodes form a network underwater. Each underwater sensor node is regarded as an agent. At the initial stage of time slot , the agent will perceive the environmental information. The state at time slot is composed of the states of all agents in the network. The state space set is expressed as:

[0069] (2),

[0070] where represents the state of the th agent at time slot , includes the following parts:

[0071] (3),

[0072] where represents the link state information, represents the geographical location information, Represents resource allocation information, Represents the current computing task information, Represents computing resource information, Represents the remaining energy information;

[0073] S1-2: Construct the action space by adding noise to the offloading mode matrix and the offloading quantity matrix:

[0074] Let R Pol =[ r 1 Pol ,..., r i Pol ,..., r I Pol ] Τ Represents the offloading mode matrix, and there is r i Pol =[u n i,1 ,u n i, j ,...,u n i, J ] ; R Vol =[ r 1 Vol ,..., r i Vol ,..., r I Vol ] Τ Represents the offloading quantity matrix, and there is r i Vol =[ g i,1 , g i, j ,..., g i, J ] ; Define the offloading matrix as multiplied by the corresponding elements of , then the offloading matrix is expressed as:

[0075] (4),

[0076] The action space consists of the offloading matrix , where indicates the computing task offloading mode of , g i, j ∈ [0, G i, j ] indicates the computing task offloading quantity of , and when it means The task is processed locally, and the matrix is unloaded The elements at the corresponding positions Are marked as negative values, that is , Indicates the total amount of computing tasks owned by the th underwater sensor node managed by the th buoy node; time slot The original action space set at Is expressed as:

[0077] (5),

[0078] Among them, Represents the th action that can be selected at the time slot th At the time slot And is formed by splicing each row of the unloading matrix , Is expressed as:

[0079] (6),

[0080] To increase the randomness of the action space, random noise is added to the original action space , and the final action space Is expressed as:

[0081] (7),

[0082] Among them, Represents the original action space, Represents random noise.

[0083] S1-3: Design a reward function using the reciprocal of the energy consumption of computing task unloading:

[0084] Use the reciprocal of the energy consumption of computing task unloading as the reward to guide the agent to gradually reduce the energy consumption of computing task unloading. The calculation formula of the reward function is

[0085] (8),

[0086] Among them, Represents the energy consumption of processing the computing task locally within the time slot , Represents the energy consumption of processing the task on the buoy node within the time slot , Represents the energy consumption of processing the task on the satellite node within the time slot ​​​The energy consumption of the upper processing task and represent the additional rewards for the buoy node service and the edge server service respectively. When the buoy node is busy with services, the value is negative, and vice versa the value is positive; when the edge server is busy with services, the value is negative, and vice versa the value is positive.

[0087] S2: Improve the deep deterministic policy gradient algorithm: First, replace the traditional minimization of the loss function with the minimization of the average loss function for training. Second, use the adaptive update amplitude method to replace the traditional fixed update amplitude method in the process of updating the parameters of the current value function and the target value function to accelerate the convergence speed of the algorithm. The specific steps are as follows:

[0088] S2-1: Calculate the current value function value of the improved deep deterministic policy gradient and conduct algorithm training by minimizing the average loss function:

[0089] The calculation formula for the current value function of the improved deep deterministic policy gradient is:

[0090] (9),

[0091] where represents the parameter of the current value function; represents the discount factor, the value range is between 0 and 1, and the closer the value is to 1, the more it emphasizes the future value; represents the current policy function; represents the reward function; represents the current value function of the next state, and the calculation formula for the target value function is:

[0092] (10),

[0093] where represents the parameter of the target value function; represents the target policy function; represents the target value function of the next state; conduct algorithm training by minimizing the average loss function, and the average loss function calculation formula is:

[0094] (11),

[0095] wherein, represents the number of all agents in the network.

[0096] S2-2: Improve the parameters of the DDPG algorithm and use the adaptive update amplitude method during the update process. Set the update amplitude to a relatively high value in the initial stage and gradually reduce it during the training process of the improved DDPG algorithm:

[0097] Update the parameters in step S2-1 using the adaptive update amplitude method and , and the calculation formula is:

[0098] (12),

[0099] wherein, is the update amplitude, which is between 0 and 1. The smaller the value, the slower the algorithm update; to stabilize the training process and ensure faster convergence while ensuring stable update of the improved DDPG algorithm, set the update amplitude to a relatively large value in the initial stage of the improved DDPG algorithm training and gradually reduce it during the training process of the improved DDPG algorithm. The calculation formula is:

[0100] (13),

[0101] wherein, and represent the initial and later update amplitudes respectively. According to formula (13), the value of the adaptive update amplitude will change from to gradually. Once or during the change process, then immediately switch to ; is the update amplitude size control threshold. After multiple experiments, it is determined that the value ranges from 0.5 to 2.5. The smaller the value, the smaller the initial update amplitude. In this embodiment, is taken; represents the switching time threshold. After multiple experiments, it is determined that the value ranges from 1 to 10. The larger the value, the longer the update amplitude switching time. In this embodiment, hours are taken.

[0102] ​​​​​​​​​​​​​​S3: Solve the computing task offloading problem by improving the Deep Deterministic Policy Gradient algorithm: Input the Markov decision model of the computing task offloading problem in step S1 into the improved Deep Deterministic Policy Gradient algorithm in S2, and continuously train the improved Deep Deterministic Policy Gradient algorithm in S2 until the average loss function no longer decreases. Then output the computing task offloading mode and the offloading quantity. The specific steps are as follows:

[0103] Solve the computing task offloading problem through the improved Deep Deterministic Policy Gradient algorithm in S2. Input the Markov decision model of the computing task offloading problem in S1; according to the current state and the current policy function select an action , and calculate the current value function value of the Deep Deterministic Policy Gradient algorithm; according to the current value function value and the target value function value, calculate the average loss function, and update the parameters and ; repeat the above process, continuously reduce the average loss function and accumulate the reward function , until the average loss function no longer decreases, and output the computing offloading mode and the computing offloading quantity .

[0104] The simulation results of the loss function with the training steps using the method provided in this embodiment and the method with a traditional fixed update amplitude are compared as Figure 5 shown. This method uses an adaptive update amplitude. The initial update amplitude and the later update amplitude are taken as and respectively. Calculate the update amplitude according to formula (13) in the embodiment. The value of the adaptive update amplitude will gradually change from 0.7 to 0.1. Once or is less than the threshold during the change process, then immediately switch to ; the two comparison methods are the methods with traditional fixed update amplitudes. For comparison method 1, is fixed and does not change. For comparison method 2, is fixed and does not change. This embodiment completes the simulation in Matlab R2021a, uses a computer with 32GB of memory, and an Intel Core i5-12400F CPU processor based on x64. The specific parameters of all simulations in this embodiment are listed in Table 1.

[0105] Table 1 Simulation parameters

[0106]

[0107] From Figure 5 the simulation results, it can be seen that when using the scheme with a fixed value, the convergence speed of the algorithm is relatively slow, and the effect of using this method is better than that of the fixed value. Because in the initial stage of network training, a larger update amplitude can approach the optimal scheme faster; as the training progresses, reducing the value of the update amplitude can make the algorithm converge more stably. When using a varying update amplitude, Figure 5 the curve representing this method in reaches convergence after 450 rounds of training, and the final loss value is relatively stable; for Comparative Method 1 with a fixed update amplitude the curve reaches convergence after 1000 rounds of training; for Comparative Method 2 with a fixed update amplitude the curve reaches convergence after 600 rounds of training, and the final loss value oscillates back and forth. This method speeds up the convergence speed by 55% compared with Comparative Method 1 with a fixed update amplitude and speeds up the convergence speed by 25% compared with Comparative Method 2 with a fixed update amplitude, and can achieve faster convergence while ensuring the stable update of the network.

[0108] The simulation results of the network offloading delay varying with the task volume using the present invention and the traditional reinforcement learning method (DQN) are compared as Figure 6 shown. The DQN algorithm is a reinforcement learning algorithm that combines deep learning and is often used to solve the problem of excessive energy consumption in underwater computing task offloading.

[0109] From Figure 6 the simulation results, it can be seen that as the task size increases, the energy consumption of the computing task offloading of both methods shows an increasing trend, and the method provided in this embodiment is less affected. When the task size is 800 KB, the method provided in this embodiment reduces the energy consumption of the computing task offloading by about 19.2% compared with the DQN method.

[0110] In summary, the present invention can effectively accelerate the convergence speed of the algorithm and reduce the energy consumption of sensor node task offloading.

[0111] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions required to be protected by the present invention.

Claims

1. A method for offloading underwater computing tasks based on improved deep deterministic policy gradient, characterized in that: The following steps are involved: S1: The Markov decision model for computing task offloading is established as follows: S1-1: Initialize the state space of the task offloading problem: Deploy a three-layer network architecture of space-sea-underwater, including M satellite nodes, denoted as Γsa={Sa1,Sa2,...,Sa M }; N buoy nodes, denoted as Γbu={Bu1,Bu2,...,Bu N }; Each buoy node provides computing services for n′ underwater sensor nodes, denoted as The i-th buoy node and the j-th underwater sensor node it manages are denoted as Bu i and Un i,j Both satellite nodes and buoy nodes have built-in edge servers that can perform data analysis and calculations. Buoy nodes are also equipped with acoustic modules and electromagnetic wave modules, which can communicate with underwater sensor nodes and satellite nodes at the same time. The underwater sensor nodes process the computing tasks generated locally, or offload the computing tasks to the edge servers of the buoy nodes or satellite nodes for computing. The computing tasks generated by the underwater sensor nodes include data analysis and running network protocols. Let G i,j and g i,j They represent the total amount of computing tasks and the number of offloaded computing tasks owned by underwater sensor nodes respectively; when g i,j ≠0, represents Un i,j The computational tasks need to be offloaded to the buoy node Bu i Or satellite node Sa m , otherwise it represents Un i,j The computational tasks are processed locally by the underwater sensor nodes; define un i,j Indicates whether the task is offloaded to a buoy node or a satellite node, expressed as: Underwater sensor nodes are networked underwater. Each underwater sensor node is regarded as an intelligent agent. At the initial stage of time slot t, the intelligent agent perceives environmental information. The state at time slot t is composed of the states of all intelligent agents in the network. The state space s(t) is expressed as: s(t)={s1′(t),s2′(t),...,s e ′(t),...,s′ V (t)} (7) Where V represents the number of all agents in the network, s e ′(t) represents the state of the e-th agent at time slot t, s e ′(t) contains the following parts: s e ′(t)={S link (t),S local (t),S alloc (t),S task (t),S resou (t),S energ (t)} (8) Among them, S link (t) represents the link status information, S local (t) represents the geographic location information, S alloc (t) represents resource allocation information, S task (t) represents the current computing task information, S resou (t) represents computing resource information, S energ (t) represents the remaining energy information; S1-2: Use the unloading pattern matrix and the unloading quantity matrix to add noise to construct the action space: set up represents the uninstallation pattern matrix, and has represents the uninstall quantity matrix, and has The unloading matrix R is defined as and By multiplying the corresponding elements, the unloading matrix R is expressed as: The action space consists of the unloading matrix R, where un i,j ∈{0,1} indicates Un i,j The computing task offloading mode, g i,j ∈[0,G i,j ]Indicates Un i,j The number of computing tasks offloaded, and when g i,j =0, indicating Un i,j The task is processed locally, and the corresponding position element of the matrix R is offloaded Marked as negative value, i.e. G i,j represents the total amount of computing tasks owned by the jth underwater sensor node managed by the i-th buoy node; the original action space set a′(t) at time slot t is expressed as: a′(t)={a′1(t),a′2(t),...,a i ′(t),...,a′ N (t)} (10) Among them, a i ′(t) represents the jth Un managed by the i-th buoy node i,j The action that can be selected at time slot t is composed of each row of the unloading matrix R, a i ′(t) is expressed as: To increase the randomness of the action space, random noise f is added to the original action space noise , the final action space a(t) is expressed as: a(t)=a′(t)+f noise (12) Among them, a′(t) represents the original action space, f noise represents random noise; S1-3: Design reward function using the inverse of computing task offloading energy consumption: The inverse of the energy consumption of computing task offloading is used as a reward to guide the agent to gradually reduce the energy consumption of computing task offloading. The calculation formula of the reward function is: in, Indicates Un in time slot t i,j The energy consumption of local processing computing tasks, Indicates Un in time slot t i,j At the buoy node Bu i The energy consumption of the processing task, Indicates Un in time slot t i,j At the satellite node Sa m The energy consumption of the processing task, and Represents the additional rewards for buoy node business and satellite node business respectively. i When business is busy, The value is negative, otherwise The value is positive; when the satellite node Sa m When business is busy, The value is negative, otherwise The value is a positive number; S2: Improved deep deterministic policy gradient algorithm: The details are as follows: S2-1: Calculate the current Q-value function value of the improved deep deterministic policy gradient and train the algorithm by minimizing the average loss function: The calculation formula of the current Q-value function of the improved deep deterministic policy gradient is: Among them, s(t) represents the state space, a(t) represents the action space, r(s(t), a(t)) represents the reward function, Represents the parameter of the current Q-value function; δ represents the discount coefficient, and the value range of δ is between 0 and 1. The closer the value is to 1, the more attention is paid to the future Q-value; represents the current policy function; r(s(t),a(t)) represents the reward function; The current Q value function representing the next state and the target Q value function are calculated as: in, Represents the parameters of the target Q-value function; represents the target policy function; Represents the target Q value function of the next state; the algorithm is trained by minimizing the average loss function. The calculation formula of the average loss function Loss(Q) is: Among them, V represents the number of all agents in the network; S2-2: Improve DDPG algorithm parameters and The adaptive update amplitude method is used in the update process. The update amplitude is set to the maximum value in the initial stage and gradually reduced during the training process of the improved DDPG algorithm: Update the parameters in step S2-1 using the adaptive update amplitude method and The calculation formula is: Among them, κ is the update amplitude, which is between 0 and 1. The smaller the κ value, the slower the algorithm updates; In the initial stage of the improved DDPG algorithm training, the update amplitude κ is set to the maximum value, and it is gradually reduced during the improved DDPG algorithm training process. The calculation formula is: Among them, κ0 and κ1 represent the initial and later update amplitudes, respectively. According to formula (5), the κ value of the adaptive update amplitude will gradually change from κ0 to κ1. During the change process, once κ≤κ1 or t>0.5T yv , then immediately switch to κ=κ1; is the update amplitude control threshold, determining The value range is between 0.5 and 2.

5. The smaller the value, the smaller the initial update amplitude. yv Represents the switching time threshold, determining T yv The value range is between 1 and 10. A larger value means a longer update amplitude switching time. S3: Solve the computing task offloading problem by improving the deep deterministic policy gradient algorithm: input the computing task offloading Markov decision model established in step S1 into the improved deep deterministic policy gradient algorithm in step S2, and continuously train the improved deep deterministic policy gradient algorithm in S2 until the average loss function no longer decreases, and output the computing task offloading mode and offloading quantity.

2. The underwater computing task offloading method based on improved deep deterministic policy gradient according to claim 1, characterized in that: In step S3, an improved deep deterministic policy gradient algorithm is used to solve the problem of computing task offloading, as follows: The computational task offloading problem is solved by the improved deep deterministic policy gradient algorithm in S2, and the Markov decision model of the computational task offloading problem in S1 is input; according to the current state s(t) and the current policy function Select action a(t) and calculate the current Q-value function value of the deep deterministic policy gradient algorithm; calculate the average loss function based on the current Q-value function value and the target Q-value function value, and update the parameters and Repeat the above process, continuously reduce the average loss function Loss(Q) and accumulate the reward function r(s(t), a(t)) until the average loss function no longer decreases and the output calculation unloading mode is reached. i,j and calculate the number of unloads g i,j .

Citation Information

Patent Citations

  • Edge computing collaborative task unloading method based on multi-agent reinforcement learning

    CN118585263A

  • Cloud edge-end collaborative UAV auxiliary task unloading method based on maximum entropy reinforcement learning

    CN117421058A

  • Satellite edge computing task unloading and resource allocation method based on deep reinforcement learning

    CN118250750A

Cited By

  • Adaptive double time scale buoy satellite temperature-aware service deployment and task scheduling method

    CN122602175A