Service caching and task unloading collaborative optimization method in unmanned aerial vehicle assisted ocean network

By building a drone-assisted marine network architecture and a dual-time scale deep reinforcement learning algorithm, the service cache and task offloading strategies are optimized, and the limited resource problem of cache and offloading in the drone-assisted marine network is solved, and low-latency and low-energy service provision is achieved.

CN120390255APending Publication Date: 2025-07-29NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510420399.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-04
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In drone-assisted marine networks, how to reasonably make service cache decisions and calculate offload decisions to minimize task completion costs, especially when the storage resources of unmanned surface vehicles and drones are limited, how to optimize cache location and content, and how to reduce the communication overhead and energy consumption caused by frequent download of service content.

Method used

Build a drone-assisted marine network architecture, establish a multi-constraint optimization model, use a dual-time-scale deep reinforcement learning algorithm to convert the objective function into a hierarchical Markov decision-making process, optimize the service cache and task offload strategies through the service cache model and the task offload model, introduce discrete and continuous action networks, and use an action masking mechanism to block illegal actions.

Benefits of technology

It realizes the minimization of cache update costs, task execution delays and energy consumption in the drone-assisted marine network, reduces the overhead of cache updates, and optimizes the collaborative efficiency of service cache and task offloading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390255A_ABST
    Figure CN120390255A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of task unloading. The invention provides a service caching and task unloading collaborative optimization method in an unmanned aerial vehicle assisted ocean network. According to the embodiment of the invention, the service cache placement, the task unloading decision and the resource allocation strategy are jointly optimized, so that the cache updating cost, the task execution delay and the energy consumption are minimized. And cache decisions under different time scales are researched, so that the overhead of cache updating is reduced. Service caching and unloading decisions are modeled into a layered Markov decision process of double time scales, a long-time-scale agent is responsible for optimizing service caching decisions, and a short-time-scale agent is responsible for optimizing task unloading and resource allocation decisions. And considering that a discrete action space and a continuous action space exist in a target problem, a discrete action network and a continuous action network are introduced. Meanwhile, in consideration of coupling of service cache and task unloading decision, an action shielding mechanism is provided to shield illegal actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of task offloading, and in particular, to a method for collaborative optimization of service caching and task offloading in an unmanned aerial vehicle-assisted marine network. Background Art

[0002] In recent years, with the continuous increase in human demand for ocean perception, more and more ocean devices, such as unmanned surface vehicles, unmanned underwater vehicles, etc., have been widely deployed in the ocean environment. This development has given rise to a large number of ocean applications, such as ocean environmental monitoring, maritime navigation, etc., and has also promoted the rapid growth of computationally intensive applications. With the continuous progress of communication technology, marine edge computing is regarded as an effective paradigm to support low-latency and energy-efficient computing. By offloading computational tasks to marine edge nodes, ocean devices can significantly improve computational efficiency. In the traditional marine edge computing architecture, the communication process is usually divided into two transmission stages: the underwater transmission segment and the radio frequency transmission segment. Underwater acoustic transmission is widely used in underwater communication and is responsible for communication and data transmission between unmanned underwater vehicles and surface nodes (such as unmanned surface vehicles); while radio frequency communication is used for communication between surface nodes, such as data exchange between unmanned surface vehicles and marine base stations. However, in the traditional architecture, due to the low data rate, high propagation delay of underwater acoustic transmission, and long-distance communication between surface and underwater devices, these factors often make it difficult to meet the efficient requirements of applications. Therefore, in order to make up for the deficiencies of the traditional architecture, unmanned aerial vehicles equipped with edge servers have become an important solution to improve computational efficiency due to their high mobility and easy deployment, which has promoted the birth of the unmanned aerial vehicle-assisted marine network computing architecture.

[0003] Under this architecture, UUVs typically offload computational tasks to UUVs, UAVs, or offshore base stations for processing. However, effectively utilizing diverse marine devices to provide low-latency, low-energy services in complex marine environments remains a significant challenge. First, existing marine applications mostly rely on a data-driven approach, requiring the caching of some service content, such as libraries, code, and artificial intelligence models, on UUVs and UAVs. However, due to the limited storage resources of UUVs and UAVs, it is crucial to rationally determine which services to cache. Second, while UUVs and UAVs can download service content from offshore base stations to update their local caches, ensuring that computational tasks are completed on devices closer to them, frequently downloading service content from remote offshore base stations incurs significant communication overhead. Finally, due to the limited energy and computational resources of UUVs and UAVs, they cannot provide computational support for all tasks for extended periods of time. Therefore, in UAV-assisted marine networks, how to rationally make service caching decisions (including cache location and cache content) and computation offloading decisions (including offloading location) to minimize task completion costs remains an important issue that needs to be addressed.

[0004] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0005] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the invention

[0006] The purpose of the embodiments of the present disclosure is to provide a method for collaborative optimization of service caching and task offloading in a drone-assisted marine network, thereby overcoming one or more problems caused by the limitations and defects of related technologies to at least a certain extent.

[0007] According to an embodiment of the present disclosure, a method for collaborative optimization of service caching and task offloading in a drone-assisted marine network is provided, the method comprising:

[0008] Constructing a drone-assisted marine network architecture; wherein the drone-assisted marine network architecture includes a marine base station, U drones, S unmanned surface vehicles, and A unmanned underwater vehicles;

[0009] Based on a drone-assisted marine network architecture, a multi-constraint optimization model is established to minimize cache update cost, task execution delay, and energy consumption. The multi-constraint optimization model includes a communication model, a service cache model, a task offloading model, and an objective function.

[0010] Convert the optimization problem of the objective function into a hierarchical Markov decision process;

[0011] The hierarchical Markov decision process is solved using a double-time-scale deep reinforcement learning algorithm to obtain the optimal service caching and computing task offloading strategies, so as to complete the collaborative optimization of service caching and task offloading.

[0012] Furthermore, in the UAV-assisted marine network architecture, the unmanned surface vehicle serves as the gateway between the underwater acoustic network and the radio frequency network, receives the computing tasks of the unmanned underwater vehicle and converts them into radio frequency signals, and communicates with the UAV and the offshore base station;

[0013] Service content is cached on the offshore base station, UAV, and unmanned surface vehicle, and the computing tasks generated by the unmanned underwater vehicle are offloaded to the edge devices caching the corresponding services for execution.

[0014] Furthermore, the communication model specifically includes: realizing the uplink transmission between the unmanned underwater vehicle and the unmanned surface vehicle through underwater acoustic communication, and the channel gain is calculated based on the distance, acoustic signal attenuation, and noise power density; realizing the task offloading and service caching update between the unmanned surface vehicle, UAV, and offshore base station through radio frequency communication;

[0015] The service caching model specifically includes: defining binary caching variables, restricting the upper limit of the caching capacity, and calculating the communication delay and storage energy consumption of caching updates;

[0016] The task offloading model specifically includes: defining binary offloading variables, selecting to offload to the unmanned surface vehicle, UAV, or offshore base station according to the task type, and calculating the total delay and energy consumption of task execution;

[0017] The objective function includes: jointly optimizing the task offloading strategy, computing resource allocation strategy, and communication resource allocation strategy to minimize the average cost of all tasks and the caching update cost.

[0018] Furthermore, the expression of the objective function is:

[0019]

[0020] where, UAV unmanned surface vehicle unmanned underwater vehicle The time domain is divided into T time slots, denoted as The set of caching update time slots is Each caching update time slot contains τ system time slots, and the number of services is K, denoted as The size of each service is L k ; β is the weight coefficient; Cost1 a (t) is the average cost of all tasks, Cost2k (t′) is the cache update cost; is a binary cache variable, and indicates that service k is cached on the unmanned surface vehicle or the unmanned aerial vehicle, indicates that the service is not cached; C i represents the total cache capacity of the unmanned surface vehicle or the unmanned aerial vehicle; is the task offloading variable; is the task offloading variable, indicating whether the task of the unmanned underwater vehicle a at time t is offloaded to the unmanned surface vehicle s, is the task offloading variable, indicating whether the task of the unmanned underwater vehicle a at time t is offloaded to the unmanned aerial vehicle u, indicates whether service k is cached on the unmanned aerial vehicle u at time t; is the computing resource allocation ratio.

[0021] Further, in the step of converting the optimization problem of the objective function into a hierarchical Markov decision process, it includes:

[0022] The large time-scale agent generates an optimal cache update strategy according to the cache configuration, the cumulative cache gain, and the service popularity; among them,

[0023] The large time-scale state space is:

[0024]

[0025] In the formula,

[0026] The large time-scale action space is:

[0027]

[0028] In the formula,

[0029] The large time-scale reward function is the cumulative reward obtained in each time slot in the large time slot t ′ :

[0030]

[0031] The small time-scale agent generates a task offloading strategy and a resource allocation strategy according to the device location, the task information, and the cache status; among them,

[0032] The small time-scale state space is:

[0033] s(t) = {l(t), J(t), c(t)}

[0034] In the formula,

[0035] The small time-scale action space is as follows:

[0036] a(t) = {q(t), ∈(t), W(t)}

[0037] Wherein,

[0038] The small time-scale reward function is the opposite of the delay and energy consumption required to complete the computing task:

[0039]

[0040] Furthermore, in the step of using the double time-scale deep reinforcement learning algorithm to solve the hierarchical Markov decision process and obtain the optimal service caching and computing task offloading strategy, it includes:

[0041] Design a hierarchical Actor-Critic architecture based on the PPO algorithm. The hierarchical Actor-Critic architecture includes a large time-scale architecture and a small time-scale architecture; among them, the large time-scale architecture includes a first discrete actor network and a first critic network, and the network parameters are θ ′ d and φ ′ ; the small time-scale architecture includes a second discrete actor network, a continuous actor network and a second critic network, and the network parameters are θ d 、θ c and φ;

[0042] The large time-scale architecture outputs discrete caching actions, and the small time-scale architecture outputs discrete task offloading actions and continuous resource allocation actions;

[0043] Introduce an action masking mechanism to mask invalid offloading actions according to the caching state, ensuring that tasks are only offloaded to devices corresponding to the cached services.

[0044] Furthermore, for the first discrete actor network, the observation state is mapped to U+S heads through the hidden layer, and each head generates K numbers; the K numbers are input into the softmax function to generate the probability value of the caching action; the probability ratio is used to quantify the change before and after the policy update. At the same time, the learning process is stabilized and smoothed through generalized advantage estimation, and the clipping function is used to limit the update degree of the policy to obtain the objective of the first discrete actor network;

[0045] For the first critic network, the expected return from the preset observation state under the target policy is quantified to generate the state value function And obtain the objective of the first critic network, i.e., the optimal service cache;

[0046] For the second discrete actor network, the observed state s(t) is mapped to A heads through the hidden layer, and each head generates 1 + U + S numbers; the 1 + U + S numbers are input into the softmax function to generate probability values representing the selection of the task offloading location of the unmanned underwater vehicle;

[0047] For the continuous actor network, a random policy π is generated by outputting the mean and variance of the Gaussian distribution of all continuous actions c 。

[0048] For the second critic network, quantify the expected return from the preset observed state s(t) under the target policy to generate the state value function V φ (s(t)), and obtain the objective of the second critic network, i.e., calculate the task offloading policy.

[0049] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0050] In the embodiments of the present disclosure, through the above-mentioned collaborative optimization method for service caching and task offloading in the UAV-assisted ocean network, on the one hand, a UAV-assisted ocean network architecture is proposed. In this architecture, service content is cached on marine base stations, UAVs, and unmanned surface vehicles, and the computing tasks generated by unmanned underwater vehicles will be offloaded to edge devices that cache the content required for the execution of the tasks. By jointly optimizing service cache placement, task offloading decisions, and resource allocation strategies, the cache update cost, task execution latency, and energy consumption are minimized. And the cache decisions at different time scales are studied to reduce the cache update overhead. On the other hand, the present application models service caching and offloading decisions as a two-time-scale hierarchical Markov decision process, where the long-time-scale agent is responsible for optimizing service cache decisions, and the short-time-scale agent is responsible for optimizing task offloading and resource allocation decisions. And considering that there are both discrete action spaces and continuous action spaces in the target problem, the discrete action network and the continuous action network are introduced. At the same time, considering the coupling of service caching and task offloading decisions, an action masking mechanism is proposed to mask illegal actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0052] Figure 1 A step diagram showing the collaborative optimization method of service caching and task offloading in a UAV-assisted ocean network in an exemplary embodiment of the present disclosure;

[0053] Figure 2 A schematic diagram showing the UAV-assisted ocean network architecture in an exemplary embodiment of the present disclosure;

[0054] Figure 3 A simulation result diagram showing the steps of the collaborative optimization method of service caching and task offloading in a UAV-assisted ocean network in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0055] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0056] In addition, the drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0057] In this example embodiment, a collaborative optimization method of service caching and task offloading in a UAV-assisted ocean network is provided. Referring to Figure 1 as shown in, the collaborative optimization method of service caching and task offloading in the UAV-assisted ocean network may include:

[0058] Step S101: Construct a UAV-assisted ocean network architecture; wherein, the UAV-assisted ocean network architecture includes an offshore base station, a U-frame UAV, S unmanned surface vehicles, and A unmanned underwater vehicles;

[0059] Step S102: Based on the UAV-assisted ocean network architecture, establish a multi-constraint optimization model with the goal of minimizing cache update cost, task execution delay, and energy consumption; wherein, the multi-constraint optimization model includes a communication model, a service caching model, a task offloading model, and an objective function;

[0060] Step S103: Convert the optimization problem of the objective function into a hierarchical Markov decision process;

[0061] Step S104: Solve the hierarchical Markov decision process using the double-time-scale deep reinforcement learning algorithm to obtain the optimal service caching and computing task offloading policies, so as to complete the collaborative optimization of service caching and task offloading.

[0062] Through the above method for collaborative optimization of service caching and task offloading in the UAV-assisted ocean network, on the one hand, a UAV-assisted ocean network architecture is proposed. In this architecture, service content is cached on marine base stations, UAVs, and unmanned surface vehicles, and the computing tasks generated by unmanned underwater vehicles will be offloaded to edge devices that cache the content required for task execution. By jointly optimizing service caching placement, task offloading decisions, and resource allocation policies, the cache update cost, task execution delay, and energy consumption are minimized. And the caching decisions at different time scales are studied to reduce the cache update overhead. On the other hand, this application models service caching and offloading decisions as a double-time-scale hierarchical Markov decision process, where the long-time-scale agent is responsible for optimizing service caching decisions, and the short-time-scale agent is responsible for optimizing task offloading and resource allocation decisions. And this application takes into account that there are both discrete action spaces and continuous action spaces in the target problem, and introduces discrete action networks and continuous action networks. At the same time, considering the coupling of service caching and task offloading decisions, an action masking mechanism is proposed to mask illegal actions.

[0063] Next, reference will be made to Figures 1 to 3 to explain each step of the above method for collaborative optimization of service caching and task offloading in the UAV-assisted ocean network in this exemplary embodiment in more detail.

[0064] In step S101, a UAV-assisted ocean network architecture is constructed; among them, the UAV-assisted ocean network architecture includes a marine base station, U UAVs, S unmanned surface vehicles, and A unmanned underwater vehicles.

[0065] Specifically, this application considers a UAV-assisted ocean network composed of a marine base station (B), U rotary-wing UAVs and S unmanned surface vehicles The network adopts a two-stage underwater acoustic and radio frequency collaborative communication architecture, aiming to provide services for A unmanned underwater vehicles underwater. In this architecture, the unmanned surface vehicle acts as a gateway between the underwater acoustic network and the radio frequency network, receives the computing tasks of the unmanned underwater vehicle, and converts them into radio frequency signals, and then communicates with the UAV and the marine base station.

[0066] The time domain is divided into T time slots, denoted as The unmanned underwater vehicle collects data in the ocean and generates computing tasks, and each task needs to be processed through a specific service. Let the number of services be K, denoted as The size of each service is L k Since the computing resources carried by the unmanned underwater vehicle are limited, they need to offload the computing tasks to the unmanned surface vehicle covering them through underwater acoustic communication for computing.

[0067] The unmanned surface vehicle equipped with a small edge server can usually only cache part of the services due to the limitations of computing resources and buffer capacity. Therefore, the computing tasks beyond its processing capacity will be preferentially offloaded to the nearby unmanned aerial vehicle that has cached the corresponding services through radio frequency communication for processing, so as to achieve a lower transmission delay. Similarly, the unmanned aerial vehicle equipped with a small edge server can also only cache part of the service content. The unmanned aerial vehicle and the unmanned surface vehicle can download new service content from the marine base station to update the local cache. In addition, the unmanned surface vehicle can also offload the computing tasks to the marine base station that is far away, rich in computing resources and has cached all services through radio frequency communication. Although this method can reduce the computing delay, it will cause a higher transmission delay.

[0068] In step S102, based on the unmanned aerial vehicle-assisted ocean network architecture, a multi-constraint optimization model is established with the goal of minimizing the cache update cost, task execution delay, and energy consumption; among them, the multi-constraint optimization model includes a communication model, a service caching model, a task offloading model, and an objective function.

[0069] Specifically, the first is the underwater acoustic communication in the communication model. At any time slot t, the unmanned underwater vehicle a needs to transmit the computing task to the unmanned surface vehicle s covering it through underwater acoustic signals. According to Urick's model, the signal attenuation surface model in underwater acoustic communication is defined as:

[0070]

[0071] where f represents the center frequency of the acoustic signal, σ is the spreading factor, d a,s (t) represents the distance between the unmanned underwater vehicle a and the unmanned surface vehicle s at time t, and G(f) is the absorption coefficient. In the Cartesian coordinate system, the height of the sea level is defined as 0, and the positions of the unmanned underwater vehicle a and the unmanned surface vehicle s at time slot t are respectively represented as a (t)=(x a (t),y a (t),z a (t)) and l s (t)=(x s (t),y s (t),0). Based on the above, the distance d a,s (t)=||l a (t)-l s(t)||2. Based on Thorp's empirical formula, the absorption coefficient is expressed as:

[0072]

[0073] Different from traditional white noise, the transmission of underwater acoustic signals in the ocean is affected by various factors, including turbulence noise N1(f), shipping noise N2(f), wind and wave noise N3(f), and thermal noise N4(f). The influence of these noise sources can be expressed as:

[0074]

[0075] Among them, o1 ∈ (0, 1) represents the shipping activity factor, and o2 represents the wind speed. Based on the above definitions, the total noise power density can be calculated as:

[0076] N(f) = N1(f) + N2(f) + N3(f) + N4(f).

[0077] The underwater acoustic uplink channel gain from AUVa to USVs can be calculated as:

[0078]

[0079] Among them, W a represents the channel bandwidth allocated for acoustic communication from the unmanned underwater vehicle a to the unmanned surface vehicle s. When non-orthogonal multiple access (NOMA) and successive interference cancellation (SIC) techniques are adopted, the receiving end needs to sort the communication links based on the channel power gain relative to the unmanned surface vehicle s. Assume that the channel gains of the unmanned underwater vehicles are arranged as g 1,s (t) > g 2,s (t) >... > g A,s (t). According to Shannon's theorem, the data transmission rate from the unmanned underwater vehicle a to the unmanned surface vehicle s can be expressed as:

[0080]

[0081] Among them, p a is the transmission power of the unmanned underwater vehicle a.

[0082] Next is the radio frequency communication in the communication model. In this application, radio frequency communication is used between the unmanned surface vehicle, the unmanned aerial vehicle, and the offshore base. At the same time, it is assumed that the above devices communicate through orthogonal frequency division multiple access (OFDMA) to eliminate interference between different links. First is the communication between the unmanned surface vehicle and the unmanned aerial vehicle. At any time slot t, the unmanned surface vehicle s can offload the computing task to the unmanned aerial vehicle u that covers its range and caches the corresponding service. Assume that the unmanned aerial vehicle hovers at a fixed height, and its position is expressed as lu (t) = (x u (t), y u (t), z u ). According to Euclidean formula, the horizontal distance d between the unmanned surface vehicle and the unmanned aerial vehicle can be calculated s,u (t) = ||l s (t) - l u (t)||². Considering the mobility of the unmanned aerial vehicle and the characteristics of high-altitude communication, the communication link between the unmanned surface vehicle and the unmanned aerial vehicle is modeled as a line-of-sight (LoS) link. Therefore, at time slot t, the data transmission rate of the unmanned surface vehicle s unloading tasks to the unmanned aerial vehicle u can be expressed as:

[0083]

[0084] where W s,u (t) is the spectral bandwidth of the communication channel between the unmanned surface vehicle and the unmanned aerial vehicle, p s is the transmission power of the unmanned surface vehicle, g0 is the channel power gain at 1 m, G0 ≈ 2.2846, β1 is the path loss exponent, and δ 2 is the Gaussian white noise.

[0085] In addition, the unmanned surface vehicle s can also offload computing tasks to the offshore base station B. Let the position of the offshore base station be e B = (x B , y B , z B ). The distance between the unmanned surface vehicle and the offshore base station can be calculated as d s,B (t) = ||e s (t) - e B ||². Considering the occlusion factors in the marine environment, the channel state between the unmanned surface vehicle and the base station is modeled as a non-line-of-sight (NLoS) link. The data transmission rate between the unmanned surface vehicle s and the base station B can be expressed as:

[0086]

[0087] where W s,B (t) is the spectral bandwidth of the communication channel between the unmanned surface vehicle and the base station, and β2 is the additional attenuation factor caused by the NLoS link. Similarly, the unmanned surface vehicle can also download service content from the offshore base station to update the local cache, and the corresponding data rate is expressed as r B,s (t).

[0088] In this application, the drone needs to download service content from the marine base station to update its local cache. Since the drone usually hovers at a high altitude and the channel environment is relatively open, the communication link between the marine base station and the drone is modeled as a line-of-sight (LoS) link. Similarly, the data transmission rate between the marine base station and the drone can be expressed as:

[0089]

[0090] where W B,u (t) is the spectral bandwidth of the communication channel from the base station to the drone, and p B is the transmit power of the base station.

[0091] Next is the service caching model. To reduce the computational delay, it is considered to cache some services on the unmanned surface vehicle and the drone. However, due to the limited cache capacity of the unmanned surface vehicle and the drone, not all services can be stored. Therefore, it is necessary to optimize the caching strategy to decide which services to cache. In addition, considering the marine environment, downloading services from the remote marine base station to the unmanned surface vehicle and the drone incurs a large communication overhead. Therefore, the caching decision is updated every τ time slots. For this purpose, the set of cache update time slots is defined as where each cache update time slot contains τ system time slots. The binary caching variable is used to indicate whether service k is cached on the unmanned surface vehicle or the drone at cache update time slot t ′ , and is defined as follows:

[0092]

[0093] When indicates that service k is cached on the unmanned surface vehicle or the drone, indicates that the service is not cached.

[0094] In addition, the total service size cached on each unmanned surface vehicle and drone cannot exceed its cache capacity, and the constraint is as follows:

[0095]

[0096] where C i represents the total cache capacity of the unmanned surface vehicle or drone i. In particular, when the cache reaches its upper limit after multiple time slots on the unmanned surface vehicle and the drone, the first-in-first-out (FIFO) strategy is adopted to replace the old service to ensure the efficient operation of the system's cache management. Here, the time overhead of cache update is defined as the communication delay generated by downloading services from the base station, and is defined as follows:

[0097]

[0098] In the system initialization phase, it is assumed that all unmanned surface vehicles and unmanned aerial vehicles have not cached any services. For the indication function, specifically:

[0099]

[0100] Based on the cache update decision, the cache storage overhead is further defined, that is, the energy consumed to store the cache content, which is expressed as follows:

[0101]

[0102] Among them, η represents the energy consumption per unit of data stored.

[0103] Finally, there is the task offloading model. In each time slot t, each unmanned underwater vehicle a generates a computing task, which is expressed as Among them, represents the size of the uploaded data related to the task, represents the number of CPU cycles required for the task, represents the type of service required for the computing task.

[0104] Due to the limited local computing resources of the unmanned underwater vehicle, the computing task must be offloaded to an unmanned surface vehicle, an unmanned aerial vehicle, or a marine base station for execution. In addition, a binary offloading strategy is adopted, that is, each task can only select one computing node for execution. The task offloading variable is defined as:

[0105]

[0106] represents the task offloaded to different computing nodes in time slot t. Here, represents that the task is offloaded to the unmanned surface vehicle for computing, represents that the task is offloaded to the unmanned aerial vehicle for computing, represents that the task is offloaded to the base station for computing.

[0107] If the unmanned underwater vehicle a is within the coverage area of the unmanned surface vehicle s and the unmanned surface vehicle s has cached the service required for the computing task then the task can be offloaded to the unmanned surface vehicle s for computing. Since the unmanned surface vehicle s needs to process the computing tasks transmitted by multiple unmanned underwater vehicles simultaneously, the computing resource allocation ratio is defined to represent the proportion of the computing resources allocated by the unmanned surface vehicle s to the task Therefore, the total computing delay of executing the computing task on the unmanned surface vehicle s includes the task upload delay and the computing execution delay, which is expressed as:

[0108]

[0109] Similarly, to complete the task the total energy consumption required consists of the communication energy consumption for task uploading and the computing energy consumption for the task, expressed as:

[0110]

[0111] where φ is the correlation coefficient, representing the total available computing resources of the unmanned surface vehicle.

[0112] The computing tasks generated by the unmanned underwater vehicle can also be offloaded to the drone u that has cached the services required by the task for computing. Since the unmanned underwater vehicle cannot directly communicate with the drone, the offloading process is as follows: The unmanned underwater vehicle a transmits the task to the unmanned surface vehicle s covering it through underwater acoustic communication; the unmanned surface vehicle s decodes the data and uploads the task to the drone u covering it and caching the corresponding services through radio frequency communication for computing. Therefore, the task delay in the drone u for computing consists of three parts: the communication delay of the unmanned underwater vehicle uploading the task to the unmanned surface vehicle the delay of the unmanned surface vehicle decoding the data

[0113] the communication delay of the unmanned surface vehicle uploading the task to the drone and the task execution delay The total computing delay is expressed as:

[0114]

[0115] where ν s represents the number of CPU cycles required to process one bit of the underwater acoustic signal, and respectively represent the proportion of computing resources allocated to the task on the unmanned surface vehicle and the drone and represents the total available computing resources of the drone.

[0116] The corresponding total energy consumption consists of the following parts: the communication energy consumption of the task uploaded to the unmanned surface vehicle the energy consumption of the unmanned surface vehicle decoding the task the communication energy consumption of the task uploaded to the drone the computing energy consumption of the task

[0117] The total energy consumption is expressed as:

[0118]

[0119] When neither the unmanned surface vehicle nor the unmanned aerial vehicle caches the services required for the task, the computing task generated by the unmanned underwater vehicle a must be offloaded to the marine base station for computing. Similar to the computing of the unmanned aerial vehicle, this offloading process still requires the unmanned surface vehicle to act as a gateway to decode the underwater acoustic signal and upload it to the base station for computing through radio frequency communication. Therefore, the total computing delay for performing the computing on the base station is expressed as:

[0120]

[0121] The corresponding total energy consumption is expressed as:

[0122]

[0123] Based on the above, the time required to complete the task is:

[0124]

[0125] The energy consumption required to complete the task is:

[0126]

[0127] Therefore, the cost model for completing this task is defined as the weighted sum of the time and energy consumption required to complete this task, expressed as:

[0128]

[0129] where α is the weight coefficient. In addition, the cost of cache update is defined as the weighted sum of the cache communication time and the cache update energy consumption, expressed as:

[0130]

[0131] The goal of this application is to minimize the average cost of completing all tasks and the cache update cost by jointly optimizing the task offloading strategy, computing resource allocation strategy, and communication resource allocation strategy. Therefore, the problem can be formulated as:

[0132]

[0133] where β is the weight coefficient.

[0134] In step S103, the optimization problem of the objective function is converted into a hierarchical Markov decision process.

[0135] Specifically, first, the large-time-scale agent is responsible for learning the optimal cache update strategy, where the large-time-scale state, action, and reward are defined in detail as follows:

[0136] Large time-scale state: Since the service caching policy highly depends on the small time-scale offloading decisions, the state space consists of the current service cache configurations of each unmanned surface vehicle and unmanned aerial vehicle, the cumulative caching gain of each cached service, and the popularity of each service on the unmanned surface vehicle and unmanned aerial vehicle, denoted as where The cumulative caching gain of each cached service k represents the cost reduction achieved by the unmanned surface vehicle and unmanned aerial vehicle caching compared to the base station computing, denoted as:

[0138]

[0139] The popularity of a service is denoted as:

[0140]

[0141] Large time-scale action: The action is the service caching policy for each unmanned surface vehicle and unmanned aerial vehicle, denoted as where

[0142] Large time-scale reward: The large time-scale reward is the cumulative reward obtained in each time slot within the large time slot t ′ denoted as:

[0143]

[0144] Secondly, the small time-scale agent is responsible for learning the optimal computing task offloading policy, where the small time-scale state, action, and reward are defined in detail as follows:

[0145] Small time-scale state: The decision in time slot t depends on the positions, task information, and cache information of all devices in the current time slot, denoted as s(t) = {l(t), J(t), c(t)}, where

[0146] Small time-scale action: The action in time slot t includes computing task offloading decisions, computing and bandwidth resource allocation decisions, denoted as a(t) = {q(t), ∈(t), W(t)}, where

[0147] Small time-scale reward: The small time-scale reward is the negative of the delay and energy consumption required to complete the computing task, denoted as:

[0148]

[0149] In step S104, the hierarchical Markov decision process is solved using a double-time-scale deep reinforcement learning algorithm to obtain the optimal service caching and computing task offloading policies, so as to complete the collaborative optimization of service caching and task offloading.

[0150] Specifically, to improve the overall performance and overcome the non-stationarity in the UAV-assisted marine network environment, this application proposes a double-slot DRL method based on PPO. This method is based on a hierarchical-AC architecture and has an action masking mechanism to mask invalid actions.

[0151] First is the large-time-scale architecture design: Since the large-time-scale agent is responsible for generating discrete cache update decisions, in the architecture proposed in this application, it includes a first discrete actor network and a first critic network, with parameters θ ′ d and φ ′ . For the first discrete actor network, the observation state is mapped to U+S (representing UAVs and unmanned surface vehicles) heads through the hidden layer. Then each head generates K (representing the number of services) numbers, and these numbers are then input into the softmax function to generate the probability values of the cache actions, expressed as:

[0152]

[0153] Then the action taken at time slot t ′ can be sampled from the above distribution.

[0154] In PPO, the probability ratio is used to quantify the change before and after the policy update, expressed as

[0155]

[0156] To stabilize and smooth the learning process, the Generalized Advantage Estimation (GAE) is integrated into the network, expressed as:

[0157]

[0158] where λ∈[0,1] is the weighting parameter and γ∈[0,1] is the reward discount factor. A clipping function is used to limit the degree of policy update, expressed as:

[0159]

[0160] Therefore, the objective of the first discrete actor network can be expressed as:

[0161]

[0162] The first Critic network is used to generate the state value function to quantify the expected return from a given state under the target policy which is expressed as:

[0163]

[0164] The objective of the first Critic network is expressed as:

[0165]

[0166] where is the target value using GAE.

[0167] Secondly, there is the design of the small time-scale architecture: Different from the large time-scale agent which is only responsible for generating discrete cache update decisions, the small time-scale agent needs to generate both discrete computing task offloading strategies and continuous bandwidth and computing resource allocation strategies. Therefore, two independent actor networks (i.e., the second discrete actor network and the continuous actor network) are designed to output discrete and continuous actions respectively, and the network parameters are θ d and θ c respectively; Similarly, like the large time-scale network architecture, it includes a critic network. For this reason, the small time-scale policy π is decomposed into a discrete policy π d that outputs discrete actions and a continuous policy π c that outputs continuous actions. Similarly, for the discrete actor network, the observed state s(t) is mapped to A heads through the hidden layer (representing the tasks of A unmanned underwater vehicles), and then each head generates 1 + U + S numbers (corresponding to the base station, the unmanned aerial vehicle, and the unmanned surface vehicle respectively), and then these numbers are input into the softmax function to generate the probability values representing the selection of the task offloading location of the unmanned underwater vehicle. For the continuous actor network, the stochastic policy π c is generated by outputting the mean and variance of the Gaussian distribution of all continuous actions. The network parameter update and the objective function are similar to those of the large time-scale network architecture.

[0168] Finally, there is the design of the action masking mechanism: Since the computing task offloading location is related to the cache policy, because tasks cannot be offloaded to devices without caching the corresponding services. In traditional DRL methods, penalty terms are usually used to limit invalid actions, but it will reduce the convergence performance of the algorithm and cannot completely avoid invalid actions. For this reason, this application designs an action masking mechanism to mask invalid actions. Specifically, first use x i(t) represents the initial logit value of the action. To distinguish whether an action is valid, a binary mask vector is introduced, where 1 and 0 represent whether the corresponding service is cached at that position (which can be obtained from the large time-scale network). Then, by applying the masking mechanism, the masked logit value x i ′ (t) can be obtained, ensuring that the values at positions where the corresponding service is not cached are 0. Finally, the probability of an invalid action after softmax processing is 0, allowing the agent to randomly sample from the correct set of actions.

[0169] In a specific embodiment, as Figure 3 shown, in a system consisting of 1 base station, 3 unmanned surface vehicles, 3 unmanned aerial vehicles, and 10 unmanned underwater vehicles, when each unmanned surface vehicle and unmanned aerial vehicle caches c = 2 services respectively, it is tested that the method proposed in this application can converge to the optimal solution after about 600 rounds of training. When the cache is c = 3 services, the method of this application can still converge in about 1000 rounds, proving the effectiveness of the method proposed in this application.

[0170] Through the above-mentioned collaborative optimization method for service caching and task offloading in the UAV-assisted ocean network, on the one hand, a UAV-assisted ocean network architecture is proposed. In this architecture, service content is cached on marine base stations, UAVs, and unmanned surface vehicles, and the computing tasks generated by unmanned underwater vehicles will be offloaded to edge devices that cache the content required for the execution of this task. By jointly optimizing service caching placement, task offloading decisions, and resource allocation strategies, the cache update cost, task execution delay, and energy consumption are minimized. And the cache decision-making at different time scales is studied to reduce the cache update overhead. On the other hand, this application models service caching and offloading decisions as a two-time-scale hierarchical Markov decision process, where the long-time-scale agent is responsible for optimizing service caching decisions, and the short-time-scale agent is responsible for optimizing task offloading and resource allocation decisions. And this application takes into account that there are both discrete action spaces and continuous action spaces in the target problem, and introduces a discrete action network and a continuous action network. At the same time, considering the coupling of service caching and task offloading decisions, an action masking mechanism is proposed to mask illegal actions.

[0171] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this disclosure, "a plurality" means two or more unless otherwise specifically defined.

[0172] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0173] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A collaborative optimization method for service caching and task offloading in an unmanned aerial vehicle-assisted ocean network, characterized in that The method includes: Constructing a drone-assisted ocean network architecture; wherein, the drone-assisted ocean network architecture includes an offshore base station, U-frame drones, S unmanned surface vehicles, and A unmanned underwater vehicles; Based on the drone-assisted ocean network architecture, establishing a multi-constraint optimization model aiming to minimize cache update cost, task execution delay, and energy consumption; wherein, the multi-constraint optimization model includes a communication model, a service caching model, a task offloading model, and an objective function; Converting the optimization problem of the objective function into a hierarchical Markov decision process; Using a double time-scale deep reinforcement learning algorithm to solve the hierarchical Markov decision process, obtaining optimal service caching and computing task offloading strategies to achieve collaborative optimization of service caching and task offloading.

2. The collaborative optimization method for service caching and task offloading in the UAV-assisted marine network according to claim 1, wherein In the drone-assisted ocean network architecture, the unmanned surface vehicle is the gateway between the underwater acoustic network and the radio frequency network, receiving the computing tasks of the unmanned underwater vehicle and converting them into radio frequency signals, and communicating with the drones and the offshore base station; Service content is cached on the offshore base station, drones, and unmanned surface vehicles, and the computing tasks generated by the unmanned underwater vehicle are offloaded to the edge devices caching the corresponding services for execution.

3. The collaborative optimization method for service caching and task offloading in the UAV-assisted marine network according to claim 1, wherein The communication model specifically includes: realizing the uplink transmission between the unmanned underwater vehicle and the unmanned surface vehicle through underwater acoustic communication, and calculating the channel gain based on distance, acoustic signal attenuation, and noise power density; realizing task offloading and service cache update between the unmanned surface vehicle, drones, and the offshore base station through radio frequency communication; The service caching model specifically includes: defining binary caching variables, restricting the upper limit of the cache capacity, and calculating the communication delay and storage energy consumption of cache update; The task offloading model specifically includes: defining binary offloading variables, selecting to offload to the unmanned surface vehicle, drones, or the offshore base station according to the task type, and calculating the total delay and energy consumption of task execution; The objective function includes: jointly optimizing the task offloading strategy, computing resource allocation strategy, and communication resource allocation strategy to minimize the average cost of all tasks and the cache update cost.

4. The collaborative optimization method for service caching and task offloading in the UAV-assisted marine network according to claim 3, wherein The expression of the objective function is: Among them, the unmanned aerial vehicle unmanned surface vehicle unmanned underwater vehicle The time domain is divided into T time slots, denoted as The set of cache update time slots is Each cache update time slot contains τ system time slots, and the number of services is K, denoted as The size of each service is L k ; β is the weight coefficient; Cost1 a (t) is the average cost of all tasks, and Cost2 k (t′) is the cache update cost; is a binary cache variable, and indicates that service k is cached on the unmanned surface vehicle or the unmanned aerial vehicle, indicates that the service is not cached; C i represents the total cache capacity of the unmanned surface vehicle or the unmanned aerial vehicle; is the task offloading variable; is the task offloading variable, indicating whether the task of unmanned underwater vehicle a at time t is offloaded to unmanned surface vehicle s, is the task offloading variable, indicating whether the task of unmanned underwater vehicle a at time t is offloaded to unmanned aerial vehicle u, indicates whether service k is cached on unmanned aerial vehicle u at time t; is the computing resource allocation ratio.

5. The collaborative optimization method for service caching and task offloading in the UAV-assisted marine network according to claim 4, wherein In the step of converting the optimization problem of the objective function into a hierarchical Markov decision process, it includes: The large time-scale agent generates an optimal cache update strategy according to the cache configuration, cumulative cache gain, and service popularity; wherein, The large time-scale state space is: In the formula, The large time-scale action space is: In the formula, The large time-scale reward function is the cumulative reward obtained in each time slot within the large time slot t′: The small time-scale agent generates a task offloading strategy and a resource allocation strategy according to the device location, task information, and cache status; wherein, The small time-scale state space is: s(t) = {l(t), J(t), c(t)} In the formula, The small time-scale action space is: a(t) = {q(t), ∈(t), W(t)} In the formula, The small time-scale reward function is the negative value of the delay and energy consumption required to complete the computing task:

6. The collaborative optimization method for service caching and task offloading in a drone-assisted ocean network according to claim 5, wherein In the step of using a double time-scale deep reinforcement learning algorithm to solve the hierarchical Markov decision process and obtaining optimal service caching and computing task offloading strategies, it includes: Design a hierarchical Actor-Critic architecture based on the PPO algorithm. The hierarchical Actor-Critic architecture includes a large time-scale architecture and a small time-scale architecture. Among them, the large time-scale architecture includes a first discrete actor network and a first critic network, and the network parameters are θ' d and φ'; the small time-scale architecture includes a second discrete actor network, a continuous actor network and a second critic network, and the network parameters are θ d , θ c and φ; The large time-scale architecture outputs discrete caching actions, and the small time-scale architecture outputs discrete task offloading actions and continuous resource allocation actions; An action masking mechanism is introduced to mask invalid offloading actions according to the cache state, ensuring that tasks are only offloaded to the devices corresponding to the cache services.

7. The collaborative optimization method for service caching and task offloading in the UAV-assisted marine network according to claim 6, wherein For the first discrete actor network, the observed state is mapped to U+S heads through the hidden layer, and each head generates K numbers; the K numbers are input into the softmax function to generate the probability values of the cached actions; the change before and after the policy update is quantified using the probability ratio. At the same time, the learning process is stabilized and smoothed through generalized advantage estimation, and a clipping function is used to limit the degree of policy update to obtain the objective of the first discrete actor network; For the first critic network, quantify the expected return from a preset observation state under the target policy to generate a state value function and obtain the objective of the first critic network, that is, the optimal service cache; For the second discrete actor network, the observed state s(t) is mapped to A heads through the hidden layer, and each head generates 1 + U + S numbers; the 1 + U + S numbers are input into the softmax function to generate probability values representing the selection of the task offloading location of the unmanned underwater vehicle; For the continuous actor network, a stochastic policy π is generated by outputting the means and variances of Gaussian distributions for all continuous actions. c . For the second critic network, quantify the expected return from the preset observation state s(t) under the target policy to generate the state-value function V φ (S(t)), and obtain the target of the second critic network, that is, calculate the task offloading policy.

Citation Information

Cited By

  • Intelligent decision-making method for cross-sea-air transmission parameters based on reinforcement learning and related equipment

    CN121887340A