Caching, migration and unloading collaborative optimization method in satellite-assisted ocean network
By building a satellite-assisted ocean network architecture and adopting an attention-enhanced multi-agent reinforcement learning method, we optimized task caching, migration, and offloading decisions, solved the problem of high task latency in ocean networks, and maximized the system utility.
Patent Information
- Application Number
- CN202510983030.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies make it difficult to effectively coordinate and optimize task caching, migration, and offloading in marine networks, resulting in high task execution delays for autonomous underwater vehicles and an inability to meet the low-latency requirements of deep-sea scenarios.
A satellite-assisted ocean network architecture is constructed, and an attention-enhanced multi-agent reinforcement learning method is adopted to optimize task caching, migration, and offloading decisions, maximizing system utility through a partially observable Markov decision process.
Under multi-dimensional resource constraints, efficient collaborative optimization of task caching, migration and offloading in marine networks is achieved, which reduces task execution latency and improves system utility.
Smart Images

Figure CN120751407A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of marine networks, and in particular to a method for collaborative optimization of caching, migration and unloading in a satellite-assisted marine network. Background Art
[0002] With the increasing development and utilization of marine resources, a large number of marine equipment, such as autonomous underwater vehicles (AUVs), have been deployed to perform a variety of marine tasks, including subsea oil exploration and environmental monitoring. These applications place higher demands on high-speed marine communications and powerful computing capabilities. Unlike terrestrial networks, which have well-developed infrastructure, the harsh climate and construction conditions at sea severely restrict the deployment of offshore network resources. Considering the limitations of marine communications in terms of coverage and bandwidth, existing methods mainly focus on utilizing offshore relay nodes, such as unmanned surface vehicles (USVs), to extend the service range of shore-based networks at sea. However, the coverage of multi-hop networks is limited by the availability of USVs and cannot meet the low-latency requirements of marine applications in deep-sea scenarios, significantly reducing the quality of experience for AUVs.
[0003] Satellite communications, on the other hand, are crucial for expanding ocean network coverage and improving communication efficiency. Therefore, by offloading computational tasks from autonomous underwater vehicles (AUVs) to satellites equipped with powerful mobile edge computing servers, mission execution latency can be significantly reduced. Furthermore, when communication conditions are poor, such as insufficient underwater acoustic communication bandwidth resources, high bit error rates, or poor satellite-to-ground links, trading computational and cache capacity for communication capacity is a common approach to alleviate transmission pressure and reduce mission execution latency. By caching the services required for mission execution (such as programs and code libraries) on AUVs or USVs, and prioritizing computational tasks locally on AUVs or offloading them to USVs, latency associated with underwater acoustic transmission, transcoding, and frequent satellite-to-ground link transmission can be reduced. However, this network consists of two segments: underwater acoustic communication and surface radio frequency communication. AUVs cannot communicate directly with satellites or USVs outside of communication range. The USVs must act as gateways, transcoding underwater acoustic signals into radio frequency signals and forwarding them to other devices. Therefore, the task offloading path and cache update path become more complicated. Secondly, considering the scenario of task migration between unmanned surface vehicles, due to the limited local computing and cache resources of each unmanned surface vehicle, the tasks migrated to the unmanned surface vehicle will compete with the content unloaded by the autonomous underwater vehicle within the coverage area of the unmanned surface vehicle. When there are too many tasks on the unmanned surface vehicle and overload occurs, the delay in task processing will increase. Therefore, it becomes more challenging to coordinate the vertical offloading and horizontal migration of tasks. Finally, the service cache and task offloading decisions are coupled with each other, that is, tasks can only be offloaded to devices that have cached the services required by the task, and task migration will further expand the mutual influence of decisions between adjacent unmanned surface vehicles. How to comprehensively consider global information and generate optimal caching, migration and offloading decisions becomes more complicated.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0005] The present invention provides a method for collaborative optimization of caching, migration and offloading in a satellite-assisted ocean network, which is used to solve the problem of low-latency and high-utility task computing under a satellite-assisted ocean network architecture.
[0006] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.
[0007] According to a first aspect of the present invention, a method for collaborative optimization of caching, migration, and offloading in a satellite-assisted marine network is provided, the method comprising: Constructing a satellite-assisted ocean network architecture, which includes three layers: satellites, unmanned surface vehicles (USVs), and autonomous underwater vehicles (AUVs). The USVs and AUVs can cache some services. The AUVs generate computing tasks requiring corresponding services using a Poisson distribution. Tasks can be executed locally on the AUV or offloaded to the USVs and satellites directly covering it. Tasks can also be laterally migrated to neighboring UUVs for execution. According to the satellite-aided ocean network architecture, the task completion delay is mapped to the system utility, and the objective function is constructed with the goal of maximizing the system utility under multi-dimensional constraints. Convert the objective function optimization problem into a partially observable Markov decision process; Attention-enhanced multi-agent reinforcement learning solves the objective function optimization problem and obtains the optimal solution.
[0008] In some exemplary embodiments, the coverage areas of all the unmanned surface vehicles do not overlap in space and jointly cover the target area.
[0009] In some exemplary embodiments, constructing the objective function is specifically:
[0010]
[0011]
[0012] in, is the total utility of all tasks in the system at any time slot t; 、 Respectively represent that service k is cached in the underwater autonomous vehicle or unmanned surface vehicle ; and Autonomous underwater vehicles Unmanned Surface Vehicles The total cache capacity, Represents a cache service Required storage space; Respectively assigned to or Bandwidth ratio; Indicates the proportion of computing resources allocated to the task; represents the unloading decision of the underwater autonomous vehicle, Unmanned Surface Vehicle The collection of underwater autonomous vehicles covered, represents the collection of unmanned surface vehicles, represents the satellite set, Represents a collection of services, Indicates assignment to The proportion of computing resources.
[0013] In some exemplary embodiments, the objective function optimization problem is converted into a partially observable Markov decision process, specifically: In the satellite-assisted ocean network architecture, the state space, observation space, action space, and reward function in the caching, migration, and offloading problems are as follows: State space: At any time slot t, the global state Indicates the environmental status of the system, which is defined as follows:
[0014] in, Represents bandwidth and computing resource information of all satellites, Represents bandwidth and computing resource information for all unmanned surface vehicles, Represents the computing resource information of all underwater autonomous vehicles, Represents the connectivity status of all unmanned surface vehicles, Represents the service cache information of the previous time slot, Represents the computational mission information of all underwater autonomous vehicles; Observing Space: Unmanned Surface Vehicles In the time slot The observation space is:
[0015] in, Represents bandwidth and computing resource information of all satellites, represents the bandwidth and computing resource information of the unmanned surface vehicle u, Represents the computing resource information of all underwater autonomous vehicles under the coverage of unmanned surface vehicle u, Represents the connectivity status of the unmanned surface vehicle u with other unmanned surface vehicles, Represents the service cache information of the previous time slot, Represents the mission information of the underwater autonomous vehicle under the coverage of the unmanned surface vehicle u; Action space: After obtaining observations After that, the UUV needs to determine the execution location of the covered tasks, make service cache decisions, and formulate bandwidth and computing resource allocation strategies; In the time slot The action is defined as:
[0016] in, represents the execution position decision of all underwater autonomous vehicle tasks within the coverage area of the unmanned surface vehicle u, represents the service caching decisions of all underwater autonomous vehicles within the coverage area of the unmanned surface vehicle u and its own service caching decision, represents the ratio of bandwidth resources allocated by all USVs and satellites to each UUV within the coverage area of USV u, represents the ratio of computing resources allocated by all USVs and satellites to all UAVs within the coverage area of USV u; Reward function: In the time slot Execute an action Unmanned surface vehicle An immediate reward will be obtained from the environment; according to the optimization objective, each unmanned surface vehicle aims to maximize the sum of its utilities covering the underwater autonomous vehicle tasks; its reward is defined as:
[0017] in, is the utility function.
[0018] In some exemplary embodiments, the attention-enhanced multi-agent reinforcement learning solves the objective function optimization problem to obtain the optimal solution, specifically: In each decision slot , for each agent Perform a two-branch policy decision process, optimizing for discrete and continuous action spaces respectively: First, the agent observes the state of its local environment , which includes information about the bandwidth and computing resources of the satellite and itself, network connection relationships, service cache status of the previous time slot, and features of tasks generated by underwater autonomous vehicles within its coverage area; this observation information is input into the convolutional attention module to extract enhanced features:
[0019] Subsequently, the enhanced observed features Input to the discrete action decision network , which outputs the probability distribution of discrete actions and samples accordingly:
[0020] in, , Represents the parameters of the discrete policy network; enhances observation With discrete actions is concatenated into an enhanced observation vector ; This vector is then fed into the continuous action decision network The output includes continuous control actions such as bandwidth ratio and computing resource allocation:
[0021] in, , represents the parameters of the continuous policy network; ultimately, the unmanned surface vehicle In the time slot The joint action is expressed as .
[0022] According to a second aspect of the present invention, there is provided a storage medium having a computer program stored thereon, which, when executed by a processor, implements the collaborative optimization method for caching, migration and offloading in a satellite-assisted marine network as described in the first aspect.
[0023] According to a third aspect of the present invention, there is provided a computer program product having a computer program stored thereon, which, when executed by a processor, implements the collaborative optimization method for caching, migration and offloading in a satellite-assisted marine network as described in the first aspect.
[0024] According to a fourth aspect of the present invention, there is provided an electronic device, comprising: processor; and a memory for storing executable instructions of the processor; The processor is configured to implement the collaborative optimization method for caching, migration and offloading in a satellite-assisted marine network according to the first aspect above by executing the executable instructions.
[0025] The collaborative optimization method for caching, migration, and offloading in a satellite-assisted marine network provided by the embodiments of the present invention comprehensively considers service caching, task migration, and offloading decisions in a satellite-assisted marine network architecture to maximize the utility of all autonomous underwater vehicles. Compared with the existing technology, it has the following advantages: 1. The joint optimization problem of service caching, task migration and computation offloading is modeled in a satellite-assisted ocean network architecture, with the goal of maximizing system utility under multi-dimensional resource constraints.
[0026] 2. Convert service caching, task migration, and computation offloading in a satellite-assisted ocean network architecture into a partially observable Markov decision process.
[0027] 3. We propose an attention-enhanced multi-agent reinforcement learning algorithm. This approach uses a hybrid policy network to jointly learn discrete decision-making (including caching, migration, and offloading) and continuous resource allocation strategies (including bandwidth and computing resource allocation). Furthermore, we design a lightweight attention mechanism module to dynamically discover the criticality of different input states, thereby enhancing subsequent decision-making performance.
[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0030] Figure 1 Flowchart of a collaborative optimization method for caching, migration, and offloading in a satellite-assisted marine network according to an exemplary embodiment of the present invention; Figure 2 Schematic diagram of a satellite-assisted marine network architecture according to an exemplary embodiment of the present invention; Figure 3 Schematic diagram of converting task completion delay into piecewise utility function. DETAILED DESCRIPTION
[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0032] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0033] In view of the shortcomings and deficiencies of the existing technology, this example embodiment provides a method for collaborative optimization of caching, migration and offloading in a satellite-assisted marine network. Figure 1 As shown, the following steps may be specifically included: S1: Build a satellite-aided ocean network architecture consisting of three layers: satellites, unmanned surface vehicles (USVs), and autonomous underwater vehicles (AUVs). The USVs and AUVs can cache some services, and the AUVs generate computing tasks requiring specific services using a Poisson distribution. These tasks can be executed locally on the AUV or offloaded to directly overlaying USVs and satellites. Tasks can also be laterally migrated to neighboring USVs. S2: Based on the model constructed in S1, the task completion delay is mapped to the system utility. Under multi-dimensional constraints, the objective function is constructed with the goal of maximizing the system utility. S3: Considering the partial observation problem in the marine network environment, the objective function optimization problem with non-convex, high-dimensional state and action space is transformed into a partially observable Markov decision process; S4: Obtaining optimal solutions to problems using attention-enhanced multi-agent reinforcement learning.
[0034] Below, each step in this exemplary implementation will be described in more detail with reference to the accompanying drawings and embodiments.
[0035] Example 1 S1: Building a satellite-aided ocean network architecture Consider a A constellation of low-orbit satellites and The satellite-assisted ocean network system composed of unmanned surface vehicles provides cache and computing services for underwater autonomous vehicles. The specific system architecture is shown in the figure below. Figure 2 As shown. The satellite set and the unmanned surface vehicle set are respectively denoted as and Since each USV has a limited coverage area, we define Unmanned Surface Vehicle The set of underwater autonomous vehicles covered, including Indicates the number of underwater autonomous vehicles it covers. To ensure complete service coverage of the mission area, it is assumed that the coverage of all unmanned surface vehicles has no overlap in space and covers the target area together. The system time is discretized into consecutive time slots, denoted as .
[0036] In any time slot , each underwater autonomous vehicle will generate a computing task related to a specific service according to the Poisson distribution. Due to the limited cache and computing power of the underwater autonomous vehicle itself, it can only store some services. When the required service has been cached, the task can be processed locally; otherwise, the task will be offloaded to the unmanned surface vehicle covering it through acoustic communication. Each unmanned surface vehicle is equipped with a lightweight edge computing device, which can also only cache some services and has limited computing power. When the unmanned surface vehicle does not cache the required service locally or is overloaded due to a large number of task offloading, the system allows the task to be migrated to its directly connected and less loaded neighboring unmanned surface vehicle. We will unmanned surface vehicles The neighbor vector of ,in ,when Unmanned Surface Vehicle Unmanned Surface Vehicles Directly connected, otherwise In addition, USVs can establish direct communication links with satellites via very small aperture terminals. We assume that each satellite has sufficient cache and computing resources and can cover all USVs.
[0037] S2: Map the task completion delay to the system utility and construct an objective function with the goal of maximizing the system utility under multi-dimensional constraints; First is the service cache model. Since most computing tasks rely on related services (such as program code, database, etc.) when executing, we define the service set in the satellite-assisted ocean network as ,in Represents the total number of service types in the system. Given that the cache capacity of underwater autonomous vehicles and unmanned surface vehicles is limited, they need to decide which services to cache based on the current network status. To this end, we define The service caching decision is:
[0038] in, (or Representation Service Cached in underwater autonomous vehicles (or unmanned surface vehicle ); otherwise , indicating that it is not cached. Considering the cache capacity constraints of underwater autonomous vehicles and unmanned surface vehicles, the following storage constraints must be met:
[0039]
[0040] in, and Autonomous underwater vehicles Unmanned Surface Vehicles The total cache capacity, Represents a cache service The required storage space. At the same time, we assume that the satellite has sufficient cache resources to store all services.
[0041] Next, we will consider the task model. In the proposed system, we assume that each AUV randomly generates computing tasks, each with different service requirements and quality of experience requirements. The arrival of tasks from all AUVs follows a Poisson distribution. Each task can be represented as a five-tuple: ,in Autonomous underwater vehicle In the time slot The generated computing tasks. Specifically, Indicates the input data size of the task, Indicates the number of CPU cycles required, Indicates the type of service required by the task. Indicates the minimum delay threshold. When the delay is lower than this value, the quality of experience is not significantly improved. Indicates the maximum tolerable delay. If this threshold is exceeded, the quality of experience will be severely degraded and the task will be considered failed.
[0042] Next is the communication model. Since electromagnetic waves are severely attenuated in underwater environments, underwater acoustic communication is widely used for underwater data transmission. In the system model constructed by this invention, the underwater autonomous vehicle can offload computing tasks to the unmanned surface vehicle within its coverage area through the underwater acoustic link. Underwater acoustic channels are often interfered with by environmental noise, including turbulent noise. , shipping noise , wind and wave noise and thermal noise , the effects of these noise sources can be expressed as
[0043] in, is the center frequency of the sound signal, represents the shipping activity factor, represents the wind speed. Therefore, the comprehensive noise power spectrum density is:
[0044] Assuming the sea level is zero, then in the time slot , unmanned surface vehicle The location is , which covers underwater autonomous vehicles The location is , the Euclidean distance between the two is:
[0045] The corresponding underwater acoustic channel attenuation model is:
[0046] in is the expansion factor, is the frequency-dependent absorption coefficient, which is expressed as:
[0047] Based on this, the signal-to-noise ratio can be obtained as:
[0048] Ultimately, autonomous underwater vehicles Unmanned Surface Vehicle The data transmission rate is expressed as:
[0049] in Unmanned Surface Vehicle The total underwater communication bandwidth, To be allocated to The bandwidth ratio, Indicates the total efficiency of the transmitter power amplifier, transducer and other circuit modules.
[0050] The system also involves communication between unmanned surface vehicles. Unmanned surface vehicle directly connected to it Located respectively and , the Euclidean distance between the two is Considering the possible obstruction in the ocean environment, the communication channel between the two unmanned surface vehicles is modeled as a non-line-of-sight (NLoS) link. The corresponding data transmission rate is:
[0051] in Unmanned Surface Vehicle The total bandwidth, To be allocated to The bandwidth ratio, is the channel power gain per unit distance, , is the path loss exponent, is Gaussian white noise, Indicates the additional loss caused by the NLoS link.
[0052] This system also involves communication from unmanned surface vehicles to satellites. Indicates time slot Unmanned surface vehicle to satellite The main losses of satellite uplink include free space path loss, polarization loss and atmospheric loss. ,satellite The location is , unmanned surface vehicle With satellite The distance is , the corresponding free space path loss is:
[0053] in is the signal wavelength. Polarization loss usually occurs when the receiving antenna does not match the incident wave polarization, and is usually less than 3 Atmospheric loss is mainly caused by absorption and scattering of gas molecules, which can be estimated through statistics and actual measurements. Since the line-of-sight signal is the main component, the channel fading probability density function of the UUV-satellite link obeys the Rice distribution, which is expressed as:
[0054] in It represents the ratio of line-of-sight to non-line-of-sight signal power, is the average received power, is the zero-order modified Bessel function. Finally, the uplink data transmission rate is:
[0055] in and are the antenna gains of satellite and unmanned surface vehicle, For satellite Available spectrum resources in the Ka band, For allocation to unmanned surface vehicles bandwidth ratio.
[0056] Based on the above, the delay model is defined as follows. When the required services are cached on the AUV and it has sufficient computing resources, the task can be executed directly locally. In this case, the total task latency is only composed of the execution latency, which is expressed as follows:
[0057] in, Autonomous underwater vehicle total computing power.
[0058] If the required service is not cached on the AUV or its computing power is insufficient, the task will be unloaded onto the covering unmanned surface vehicle Execution. The total task delay is composed of two parts: the communication delay of task upload Execution delay on unmanned surface vehicles , the overall expression is:
[0059] in, Unmanned Surface Vehicle The total computing power of The ratio of computing resources allocated to the task.
[0060] If unmanned surface vehicle If the required service is not cached or the load is too heavy due to excessive task pressure, the task can be migrated to the neighboring unmanned surface vehicle that is directly connected to it. Considering that the communication between the underwater autonomous vehicle and the unmanned surface vehicle is an underwater acoustic link, and the communication between unmanned surface vehicles usually uses a radio frequency link on the water surface, the task needs to be converted from the underwater acoustic signal to the radio frequency signal before migration. In this scenario, the total delay of the task completion includes four parts: the initial underwater acoustic communication delay from the underwater autonomous vehicle to the unmanned surface vehicle , transcoding delay , migration communication delay between unmanned surface vehicles and execution delay on the target unmanned surface vehicle , the overall expression is:
[0061] in, Indicates the number of CPU cycles required to transcode each bit. and Unmanned Surface Vehicles and The total computing power of and The resource allocation ratio for the corresponding task.
[0062] When the link quality between UUVs is poor or the required services are not available in the UUV network, the UUV The task can be Offload to satellite At this time, the total delay includes: underwater acoustic communication delay from underwater autonomous vehicle to unmanned surface vehicle , transcoding delay on unmanned surface vehicles , uplink transmission delay from unmanned surface vehicle to satellite and execution latency on the satellite , totaling:
[0063] At time slot t, each unmanned surface vehicle Generate unloading decisions for all underwater autonomous vehicles it covers The dimension of each underwater autonomous vehicle's unloading decision is One-hot vector Indicates that For the dimensions. , indicating that the task is selected to be executed locally on the underwater autonomous vehicle; if , indicating that the task is directly offloaded to the unmanned surface vehicle it covers ;like , indicating that the mission is first transmitted to the unmanned surface vehicle via the underwater acoustic link , and then transferred to other unmanned surface vehicles through radio frequency links after transcoding Otherwise, the task will be offloaded to the unmanned surface vehicle , and then uplinked to the satellite after transcoding Execution. Therefore, the total delay of task completion can be uniformly expressed as:
[0064] Next is the utility model. Based on the utility function, Figure 3 As shown, the utility of all tasks is evaluated. The utility function is defined as follows:
[0065] Therefore, in any time slot , the total utility of all tasks in the system can be expressed as:
[0066] The goal of this invention is to maximize the utility of all tasks, and the objective function is expressed as
[0067]
[0068]
[0069] S3: Convert the objective function optimization problem with non-convex, high-dimensional state and action spaces into a partially observable Markov decision process; Under the satellite-assisted ocean network architecture, the state space, observation space, action space, and reward function in the caching, migration, and offloading problems are as follows.
[0070] First is the state space: at any time slot t, the global state Indicates the environmental status of the system, which is defined as follows:
[0071] The status information includes the bandwidth and computing resources of all satellites, unmanned surface vehicles, and underwater autonomous vehicles, the connection relationship of all unmanned surface vehicles, the cache status of all unmanned surface vehicles and satellites (the cache behavior of the previous time slot is used as the current status), and the mission information generated by each underwater autonomous vehicle in the current time slot.
[0072] Observation space: Since each unmanned surface vehicle can only perceive limited local information, it cannot directly obtain the global state , so its observation space is limited. Each unmanned surface vehicle can only observe the information related to the underwater autonomous vehicles within its coverage area. To facilitate task migration, we assume that all unmanned surface vehicles can broadcast their own resource availability and cache status. In actual deployment, the communication overhead brought by such state sharing can be ignored. Therefore, we assume that each unmanned surface vehicle can access global resources and cache information, which is feasible in real systems. Therefore, the unmanned surface vehicle In the time slot The observation space is:
[0073] Action space: After obtaining observations After that, the unmanned surface vehicle needs to determine the execution location of the covered tasks, make service cache decisions, and formulate bandwidth and computing resource allocation strategies. In the time slot The action is defined as:
[0074] Among them, offloading decision and service cache decision are discrete actions, and resource allocation strategy is a continuous action.
[0075] Reward function: In the time slot Execute an action Unmanned surface vehicle An immediate reward will be obtained from the environment. According to the optimization objective, each unmanned surface vehicle aims to maximize the sum of its utilities covering the underwater autonomous vehicle mission. Therefore, its reward is defined as:
[0076] S4: Obtaining optimal solutions to problems using attention-enhanced multi-agent reinforcement learning; Specific algorithms such as Figure 3 As shown, in each decision time slot , for each agent A two-branch policy decision process is performed, optimizing for discrete and continuous action spaces respectively.
[0077] First, the agent observes the state of its local environment , which includes information about the bandwidth and computing resources of the satellite and itself, network connection relationships, service cache status of the previous time slot, and features of tasks generated by underwater autonomous vehicles within its coverage area. This observation information is input into the convolutional attention module to extract enhanced features:
[0078] Subsequently, the enhanced observed features Input to the discrete action decision network , which outputs a probability distribution over discrete actions (such as caching, offloading, and migration options) and samples accordingly:
[0079] in, , Represents the parameters of the discrete policy network. Enhanced observation With discrete actions is concatenated into an enhanced observation vector This vector is then fed into the continuous action decision network The output includes continuous control actions such as bandwidth ratio and computing resource allocation:
[0080] in, , represents the parameters of the continuous policy network. Finally, the unmanned surface vehicle In the time slot The joint action is expressed as .
[0081] It should be noted that, as another aspect, the present application also provides a storage medium, which may be included in an electronic device or may exist independently without being incorporated into the electronic device. The storage medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.
[0082] In one embodiment, the present application provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0083] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0084] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.
[0085] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings and that various modifications and variations can be made without departing from the scope thereof, which is limited only by the appended claims.
Claims
1. A method for collaborative optimization of caching, migration and offloading in a satellite-assisted marine network, characterized in that: The method comprises: Constructing a satellite-assisted ocean network architecture, which includes three layers: satellites, unmanned surface vehicles (USVs), and autonomous underwater vehicles (AUVs). The USVs and AUVs can cache some services. The AUVs generate computing tasks requiring corresponding services using a Poisson distribution. Tasks can be executed locally on the AUV or offloaded to the USVs and satellites directly covering it. Tasks can also be laterally migrated to neighboring UUVs for execution. According to the satellite-aided ocean network architecture, the task completion delay is mapped to the system utility, and the objective function is constructed with the goal of maximizing the system utility under multi-dimensional constraints. Convert the objective function optimization problem into a partially observable Markov decision process; Attention-enhanced multi-agent reinforcement learning solves the objective function optimization problem and obtains the optimal solution.
2. The method according to claim 1, characterized in that The coverage areas of all the unmanned surface vehicles do not overlap in space and jointly cover the target area.
3. The method according to claim 1, characterized in that The objective function is constructed as follows: in, is the total utility of all tasks in the system at any time slot t; 、 Respectively represent that service k is cached in the underwater autonomous vehicle or unmanned surface vehicle ; and Autonomous underwater vehicles Unmanned Surface Vehicles The total cache capacity, Represents a cache service Required storage space; Respectively assigned to or Bandwidth ratio; Indicates the proportion of computing resources allocated to the task; represents the unloading decision of the underwater autonomous vehicle, Unmanned Surface Vehicle The collection of underwater autonomous vehicles covered, represents the collection of unmanned surface vehicles, represents the satellite set, Represents a collection of services, Indicates assignment to The proportion of computing resources.
4. The method according to claim 1, wherein The objective function optimization problem is transformed into a partially observable Markov decision process, specifically: In the satellite-assisted ocean network architecture, the state space, observation space, action space, and reward function in the caching, migration, and offloading problems are as follows: State space: At any time slot t, the global state Indicates the environmental status of the system, which is defined as follows: in, Represents bandwidth and computing resource information of all satellites, Represents bandwidth and computing resource information for all unmanned surface vehicles, Represents the computing resource information of all underwater autonomous vehicles, Represents the connectivity status of all unmanned surface vehicles, Represents the service cache information of the previous time slot, Represents the computational mission information of all underwater autonomous vehicles; Observing Space: Unmanned Surface Vehicles In the time slot The observation space is: in, Represents bandwidth and computing resource information of all satellites, represents the bandwidth and computing resource information of the unmanned surface vehicle u, Represents the computing resource information of all underwater autonomous vehicles under the coverage of unmanned surface vehicle u, Represents the connectivity status of the unmanned surface vehicle u with other unmanned surface vehicles, Represents the service cache information of the previous time slot, Represents the mission information of the underwater autonomous vehicle under the coverage of the unmanned surface vehicle u; Action space: After obtaining observations After that, the UUV needs to determine the execution location of the covered tasks, make service cache decisions, and formulate bandwidth and computing resource allocation strategies; In the time slot The action is defined as: in, represents the execution position decision of all underwater autonomous vehicle tasks within the coverage area of the unmanned surface vehicle u, represents the service caching decisions of all underwater autonomous vehicles within the coverage area of the unmanned surface vehicle u and its own service caching decision, represents the ratio of bandwidth resources allocated by all USVs and satellites to each UUV within the coverage area of USV u, represents the ratio of computing resources allocated by all USVs and satellites to all UAVs within the coverage area of USV u; Reward function: In the time slot Execute an action Unmanned surface vehicle An immediate reward will be obtained from the environment; according to the optimization objective, each unmanned surface vehicle aims to maximize the sum of its utilities covering the underwater autonomous vehicle tasks; its reward is defined as: in, is the utility function.
5. The method according to claim 1, wherein The multi-agent reinforcement learning based on attention enhancement solves the objective function optimization problem and obtains the optimal solution, specifically: In each decision slot , for each agent Perform a two-branch policy decision process, optimizing for discrete and continuous action spaces respectively: First, the agent observes the state of its local environment , which includes information about the bandwidth and computing resources of the satellite and itself, network connection relationships, service cache status of the previous time slot, and the characteristics of the tasks generated by the underwater autonomous vehicles within its coverage area; This observation information is fed into the convolutional attention module to extract enhanced features: Subsequently, the enhanced observed features Input to the discrete action decision network , which outputs the probability distribution of discrete actions and samples accordingly: in, , Represents the parameters of the discrete policy network; enhances observation With discrete actions is concatenated into an enhanced observation vector ; This vector is then fed into the continuous action decision network The output includes continuous control actions such as bandwidth ratio and computing resource allocation: in, , represents the parameters of the continuous policy network; ultimately, the unmanned surface vehicle In the time slot The joint action is expressed as .
6. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for collaborative optimization of caching, migration and offloading in a satellite-assisted marine network according to any one of claims 1 to 5 is implemented.
7. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for collaborative optimization of caching, migration and offloading in a satellite-assisted marine network according to any one of claims 1 to 5 is implemented.
8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the method for collaborative optimization of caching, migration and offloading in a satellite-aided marine network according to any one of claims 1 to 5 by executing the executable instructions.
Citation Information
Cited By
TD3-based cross-domain ocean network task unloading and computing resource allocation method
CN121985036A