Satellite-ground integrated cache updating and cooperative transmission method and device and storage medium

By adopting deep reinforcement learning and behavioral cloning technologies in the integrated caching network of the satellite-earth integrated caching network, the cache decision-making solution is designed, and the problem of inefficient energy efficiency of the satellite-earth integrated caching network is solved, and the system's energy efficiency is significantly improved and the overall performance is improved.

CN120110487AActive Publication Date: 2025-06-06HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510245536.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Due to the limited edge cache space, the integrated caching network faces the challenge of inefficient energy efficiency. Unreasonable caching strategies will lead to reduced system resource utilization efficiency and affect overall performance.

Method used

The deep reinforcement learning method is used to design a cache decision-making solution. By building a satellite-ground integrated cache network system, three cooperative transmission links, base station-user link, satellite-base station link and gateway-satellite shared link are established, optimization problems are established based on channel models, and the deep deterministic policy gradient algorithm (DDPG) and behavioral cloning technology are used for solving them, which are decomposed into the cache content update sub-problems of base stations and satellites, realizing asynchronous updates.

Benefits of technology

Through iterative training, the deep deterministic strategy gradient algorithm is converged, and the optimal integrated star-ground cooperative cache strategy is obtained, which significantly improves the energy efficiency of the system and improves overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110487A_ABST
    Figure CN120110487A_ABST
Patent Text Reader

Abstract

The invention discloses a satellite-ground integrated cache updating and cooperative transmission method and a related device, and the method comprises the steps: building three cooperative transmission links, namely a base station-user link, a satellite-base station link and a gateway-satellite shared link, and building a channel model based on a probability density function; based on the channel model, establishing an optimization problem which takes satellite and base station layer cache contents as optimization variables and takes maximization of long-term system energy efficiency as a target; and an optimization problem is solved based on a depth deterministic strategy gradient algorithm, a behavior cloning technology is introduced to improve the training performance of the depth deterministic strategy gradient algorithm, and finally an optimal satellite-ground integrated cooperative caching strategy is obtained. According to the method, the energy efficiency of the satellite-ground integrated network system is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a satellite-ground integrated cache update and collaborative transmission method, device and storage medium. Background Art

[0002] As the development direction of the mainstream network architecture in the future, the satellite-ground integrated cache network has edge cache devices that can effectively share the pressure of network traffic and reduce system overhead. However, due to the limited edge cache space, the satellite-ground integrated network faces the challenge of low energy efficiency. Unreasonable cache strategies will not only reduce the utilization efficiency of system resources, but also further reduce the cache efficiency due to dynamic changes in user demand and channel parameters in long-term operation, thereby affecting the overall performance of the system. Summary of the invention

[0003] The present invention provides a satellite-ground integrated cache update and collaborative transmission method and related devices, aiming to effectively improve the energy efficiency of the system.

[0004] The present invention provides a satellite-ground integrated cache update and collaborative transmission method, the method comprising the following steps:

[0005] Constructing a satellite-ground integrated cache network system, establishing three cooperative transmission links, namely, a base station-user link, a satellite-base station link, and a gateway-satellite shared link, based on the satellite-ground integrated cache network system, and establishing a channel model for the three cooperative transmission links based on a probability density function;

[0006] Based on the channel model, an optimization problem P1 is established with the cache contents at the satellite and base station layers as optimization variables and the goal of maximizing the long-term system energy efficiency;

[0007] The optimization problem P1 is solved based on the deep deterministic policy gradient algorithm, and the behavior cloning technology is used to improve the training performance of the deep deterministic policy gradient algorithm. During the solution process, according to the asynchronous update mechanism of the satellite and the base station, the optimization problem P1 is decomposed into the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2. Among them, subproblem P1.1 is based on the predicted value of user request information within the base station update cycle, and the deep deterministic policy gradient algorithm is used to obtain the cache decision strategy under the current base station cache content distribution; subproblem P1.2 is based on the current base station cache strategy and the predicted value of user requests within the update cycle, and the deep deterministic policy gradient algorithm is used to obtain the satellite cache decision strategy;

[0008] Through iterative training, the deep deterministic policy gradient algorithm for the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 converges, and finally the optimal satellite-ground integrated collaborative caching strategy is obtained.

[0009] Optionally, the specific implementation steps of the method include:

[0010] Step 1: Based on the constructed satellite-ground integrated cache network system, the system energy efficiency expression is calculated using the channel parameters;

[0011] Step 2: Initialize the network parameters of the deep deterministic policy gradient algorithm at the satellite and base station layers;

[0012] Step 3: collect a section of historical user request information, select an expert strategy to generate a corresponding base station layer expert cache solution and a satellite expert cache solution, assemble them into {user request information, expert cache solution}, and store them in the base station layer expert experience cache pool and the satellite expert experience cache pool respectively;

[0013] Step 4: Randomly extract a group of samples from the base station expert experience cache pool and the satellite expert experience cache pool, and use the back propagation mechanism in supervised learning to pre-train the Actor network in the deep deterministic policy gradient algorithm;

[0014] Step 5: Determine whether the training round is greater than the set expert training round. If not, repeat step 3. If it reaches the set expert training round, the Actor network pre-training is considered complete, and the parameters of the Actor network are copied to the target Actor network.

[0015] Step 6: Determine whether the current phase is a satellite update phase or a base station update phase. If it is a satellite update phase, proceed to step 7; if it is a base station update phase, proceed to step 11;

[0016] Step 7: Collect the predicted value of user request information in the current satellite stage, as well as the current satellite deep deterministic policy gradient algorithm network parameters, satellite experience cache pool, and base station cache decision strategy;

[0017] Step 8: Generate a base station cache plan based on the current base station cache decision strategy and the predicted value of the user request information, and form the input state of the satellite deep deterministic policy gradient algorithm with {current satellite cache plan, base station cache plan, user request information}, and obtain the output action, and observe the long-term network system energy efficiency as a reward after taking the action;

[0018] Step 9: Put the experience samples {state, action, reward, next state} into the satellite experience buffer pool, and randomly extract a group of samples from it to train the deep deterministic policy gradient algorithm and update the deep deterministic policy gradient algorithm network parameters;

[0019] Step 10: Determine whether the number of training times reaches the set maximum number. If not, repeat step 8. If the maximum number of training times is reached, the current satellite deep deterministic policy gradient algorithm training is considered complete, the current satellite cache decision content is output, and jump to step 15.

[0020] Step 11: Collect the predicted value of user request information in the current base station stage, as well as the current base station deep deterministic policy gradient algorithm network parameters, base station experience cache pool and satellite cache solution;

[0021] Step 12: {current satellite cache solution, base station cache solution, user request information} is used as the input state of the base station deep deterministic policy gradient algorithm, and the output action is obtained. After taking the action, the short-term network system energy efficiency is observed as a reward;

[0022] Step 13: Put the experience sample {state, action, reward, next state} into the base station experience buffer pool, and randomly extract a group of samples from it to train the deep deterministic policy gradient algorithm, and update the deep deterministic policy gradient algorithm network parameters;

[0023] Step 14: determine whether the number of training times reaches the set maximum number of times. If not, repeat step 11; if it reaches the set maximum number of times, it is determined that the current base station deep deterministic policy gradient algorithm training is completed, output the current base station cache decision content, and jump to step 15;

[0024] Step 15: Determine whether there is no new update phase. If there is a new update phase, repeat step 6; if not, end.

[0025] Optionally, the optimization problem P1 with the goal of maximizing the long-term system energy efficiency is specifically expressed as:

[0026]

[0027] Where T is the index of the satellite update cycle, and the satellite update cycle T contains W base station update cycles t, that is, represents the cache matrix of the base station within the satellite update period T and the base station update period t, represents the cache vector of the satellite within the base station update period t within the satellite update period T, Indicates the total amount of user request data in the network, represents the total energy consumption of the system, C1 represents the cache space limit of the base station. represents the maximum cache space of the base station, C2 represents the cache space limit of the satellite, Indicates the maximum cache space of the satellite, and C3 indicates that the maximum cache data volume of the file in any cache device is sufficient to restore the source file.

[0028] Optionally, the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 are expressed as follows:

[0029]

[0030] stC1,C2,C3

[0031]

[0032] stC1,C2,C3.

[0033] Optionally, the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 are solved by continuously looping the following steps until the algorithm gradually converges:

[0034] fixed Decision-making strategy to solve Make maximum;

[0035] fixed Decision-making strategy to solve Make maximum.

[0036] Optionally, a deep deterministic policy gradient algorithm is used to obtain a base station cache decision strategy under the current satellite cache content distribution, including:

[0037] Base station status: It consists of the current satellite cache decision content and the predicted value of the user request information, namely: in is the predicted value of the file request probability of file i in cell n in satellite period T and base station period t, Refers to the cache vector of the satellite within the base station update period t within the satellite update period T, Refers to the cache matrix of the base station in the previous update period t-1 within the satellite update period T;

[0038] Base station action: The action is represented by the ratio of the cached data volume of each file stored in each base station, that is: in

[0039] Base station reward: The reward is expressed using the short-term energy efficiency of the system, namely: Indicates the total amount of user request data in the network, Represents the total energy consumption of the system.

[0040] Optionally, a satellite cache decision strategy is obtained using a deep deterministic policy gradient algorithm, including:

[0041] Satellite status: in Representation based on state As input, the base station cache update decision at the satellite update time slot is obtained by using the deep deterministic policy gradient algorithm output, where represents the cache vector of the base station in the last base station update cycle within the satellite cycle T-1, represents the cache vector of the satellite within the base station update period t within the satellite update period T-1, where in

[0042] Satellite action: The action is represented by the ratio of the amount of cached data per file stored in the satellite, i.e.:

[0043] Satellite Rewards: Rewards are expressed in terms of the long-term energy efficiency of the system, namely: Indicates the total amount of user request data in the network, Represents the total energy consumption of the system.

[0044] A second aspect of the present invention provides a satellite-ground integrated cache update and collaborative transmission device, the device comprising:

[0045] A network system construction module is used to construct a satellite-ground integrated cache network system, establish three cooperative transmission links, namely, a base station-user link, a satellite-base station link, and a gateway-satellite shared link, based on the satellite-ground integrated cache network system, and establish a channel model for the three cooperative transmission links based on a probability density function;

[0046] The target optimization problem building module is used to establish the optimization problem P1 based on the channel model, taking the cache contents of the satellite and base station layers as optimization variables and aiming at maximizing the long-term system energy efficiency;

[0047] The target optimization problem solving module is used to solve the optimization problem P1 based on the deep deterministic policy gradient algorithm, and the behavior cloning technology is used to improve the training performance of the deep deterministic policy gradient algorithm. During the solution process, according to the asynchronous update mechanism of the satellite and the base station, the optimization problem P1 is decomposed into the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2. Among them, subproblem P1.1 is based on the user request information prediction value within the base station update cycle, and the deep deterministic policy gradient algorithm is used to obtain the cache decision strategy under the current base station cache content distribution; subproblem P1.2 is based on the current base station cache strategy and the user request prediction value within the update cycle, and the deep deterministic policy gradient algorithm is used to obtain the satellite cache decision strategy;

[0048] The cache strategy acquisition module is used to converge the deep deterministic policy gradient algorithm of the base station cache content update subproblem P1.1 and the star cache content update subproblem P1.2 through iterative training, and finally obtain the optimal satellite-ground integrated collaborative cache strategy.

[0049] A third aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the satellite-ground integrated cache update and collaborative transmission method as described in any one of the first aspects of the present invention.

[0050] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed, the satellite-ground integrated cache update and collaborative transmission method as described in any one of the first aspects of the present invention is implemented.

[0051] It can be seen from the above technical solutions that the present invention has the following advantages:

[0052] The present invention uses coding cache technology to process cache files, thereby improving cache efficiency; the present invention models the channel model based on the probability density function, fully considering the dynamic changes of the channel model during long-term network operation; the present invention uses behavioral cloning technology to pre-train the deep deterministic policy gradient algorithm (DDPG algorithm), thereby solving the reward sparseness problem faced by the DDPG algorithm in the optimization of the satellite-ground integrated network; the present invention uses a dual DDPG algorithm to make cache decisions on the cache contents of the base station layer and the satellite, and combines them using the idea of ​​iteration, thereby realizing the asynchronous update problem of base stations and satellites, which is rarely concerned in the current satellite-ground integrated network research using deep reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0054] Figure 1 A flowchart of the implementation steps of a satellite-ground integrated cache update and collaborative transmission method provided in an embodiment of the present invention;

[0055] Figure 2 A communication topology diagram of a satellite-ground integrated cache network provided in an embodiment of the present invention;

[0056] Figure 3 A structural block diagram of a satellite-ground integrated cache update and collaborative transmission device provided in an embodiment of the present invention;

[0057] Figure 4 A performance comparison diagram of cache decision algorithms in a small-scale ISTN provided for verifying an embodiment of the present invention;

[0058] Figure 5 Energy efficiency performance comparison diagram of various cache decision algorithms under different file numbers provided for verification of the embodiment of the present invention

[0059] Figure 6 A comparison chart of energy efficiency performance of various cache decision algorithms under different numbers of base stations provided for verifying the embodiment of the present invention. DETAILED DESCRIPTION

[0060] The embodiment of the present invention provides a satellite-ground integrated cache update and collaborative transmission method and related devices, aiming to solve the optimization problem of two-level cache devices at the base station layer and the satellite layer in a satellite-ground integrated network under a dynamic environment.

[0061] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] As the development direction of the mainstream network architecture in the future, the satellite-ground integrated cache network has edge cache devices that can effectively share the pressure of network traffic and reduce system overhead. However, due to the limited edge cache space, the satellite-ground integrated network faces the challenge of low energy efficiency. Unreasonable cache strategies will not only reduce the utilization efficiency of system resources, but also further reduce the cache efficiency due to dynamic changes in user demand and channel parameters in long-term operation, thereby affecting the overall performance of the system.

[0063] In response to this complex energy efficiency optimization problem, the present invention uses a deep reinforcement learning method to design a cache decision solution, dynamically adapting to environmental changes and improving cache efficiency in an intelligent way. In addition, to further improve the utilization of cache resources, the scheme introduces a coding cache technology based on LT code (Luby Transform Code) to optimize cache performance and enhance the resource allocation efficiency of the system under limited cache conditions.

[0064] The satellite-ground integrated cache network involves collaborative caching and delivery of two-level cache devices. Unreasonable cache content distribution will reduce the offloading performance of limited cache space for network traffic. Existing research mostly focuses on static cache design based on known and fixed content popularity, which is difficult to adapt to dynamic actual scenarios. In addition, current research mainly focuses on complete file transmission, and pays insufficient attention to the application of encoding cache technology. In terms of cache updates, most methods based on deep reinforcement learning only consider the situation where base stations and satellites are updated synchronously. To address these problems, this study proposes a dynamic caching strategy based on deep reinforcement learning, which supports asynchronous updates of base stations and satellites to maximize system energy efficiency and improve overall performance.

[0065] A satellite-ground integrated cache update and collaborative transmission method provided in Embodiment 1 of the present invention includes the following steps:

[0066] Constructing a satellite-ground integrated cache network system, establishing three cooperative transmission links, namely, a base station-user link, a satellite-base station link, and a gateway-satellite shared link, based on the satellite-ground integrated cache network system, and establishing a channel model for the three cooperative transmission links based on a probability density function;

[0067] Based on the channel model, i.e., the data transmission volume and power consumption of the three links, an optimization problem P1 is established with the cache contents at the satellite and base station layers as optimization variables and the goal of maximizing the long-term system energy efficiency;

[0068] The optimization problem P1 is solved based on the deep deterministic policy gradient algorithm, and the behavior cloning technology is used to improve the training performance of the deep deterministic policy gradient algorithm. During the solution process, according to the asynchronous update mechanism of the satellite and the base station, the optimization problem P1 is decomposed into the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2. Among them, subproblem P1.1 is based on the predicted value of user request information within the base station update cycle, and the deep deterministic policy gradient algorithm is used to obtain the cache decision strategy under the current base station cache content distribution; subproblem P1.2 is based on the current base station cache strategy and the predicted value of user requests within the update cycle, and the deep deterministic policy gradient algorithm is used to obtain the satellite cache decision strategy;

[0069] Through iterative training, the deep deterministic policy gradient algorithm for the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 converges, and finally the optimal satellite-ground integrated collaborative caching strategy is obtained.

[0070] Specifically, Figure 1 As shown, the specific implementation steps of the method include:

[0071] Step 1: Based on the constructed satellite-ground integrated cache network system, the system energy efficiency expression is calculated using the channel parameters;

[0072] Step 2: Initialize the network parameters of the deep deterministic policy gradient algorithm at the satellite and base station layers;

[0073] Step 3: collect a section of historical user request information, select an expert strategy to generate a corresponding base station layer expert cache solution and a satellite expert cache solution, assemble them into {user request information, expert cache solution}, and store them in the base station layer expert experience cache pool and the satellite expert experience cache pool respectively;

[0074] Step 4: Randomly extract a group of samples from the base station expert experience cache pool and the satellite expert experience cache pool, and use the back propagation mechanism in supervised learning to pre-train the Actor network in the deep deterministic policy gradient algorithm;

[0075] Step 5: Determine whether the training round is greater than the set expert training round. If not, repeat step 3. If it reaches the set expert training round, the Actor network pre-training is considered complete, and the parameters of the Actor network are copied to the target Actor network.

[0076] Step 6: Determine whether the current phase is a satellite update phase or a base station update phase. If it is a satellite update phase, proceed to step 7; if it is a base station update phase, proceed to step 11;

[0077] Step 7: Collect the predicted value of user request information in the current satellite stage, as well as the current satellite deep deterministic policy gradient algorithm network parameters, satellite experience cache pool, and base station cache decision strategy;

[0078] Step 8: Generate a base station cache plan based on the current base station cache decision strategy and the predicted value of the user request information, and form the input state of the satellite deep deterministic policy gradient algorithm with {current satellite cache plan, base station cache plan, user request information}, and obtain the output action, and observe the long-term network system energy efficiency as a reward after taking the action;

[0079] Step 9: Put the experience samples {state, action, reward, next state} into the satellite experience buffer pool, and randomly extract a group of samples from it to train the deep deterministic policy gradient algorithm and update the deep deterministic policy gradient algorithm network parameters;

[0080] Step 10: Determine whether the number of training times reaches the set maximum number. If not, repeat step 8. If the maximum number of training times is reached, the current satellite deep deterministic policy gradient algorithm training is considered complete, the current satellite cache decision content is output, and jump to step 15.

[0081] Step 11: Collect the predicted value of user request information in the current base station stage, as well as the current base station deep deterministic policy gradient algorithm network parameters, base station experience cache pool and satellite cache solution;

[0082] Step 12: {current satellite cache solution, base station cache solution, user request information} is used as the input state of the base station deep deterministic policy gradient algorithm, and the output action is obtained. After taking the action, the short-term network system energy efficiency is observed as a reward;

[0083] Step 13: Put the experience sample {state, action, reward, next state} into the base station experience buffer pool, and randomly extract a group of samples from it to train the deep deterministic policy gradient algorithm, and update the deep deterministic policy gradient algorithm network parameters;

[0084] Step 14: determine whether the number of training times reaches the set maximum number of times. If not, repeat step 11; if it reaches the set maximum number of times, it is determined that the current base station deep deterministic policy gradient algorithm training is completed, output the current base station cache decision content, and jump to step 15;

[0085] Step 15: Determine whether there is no new update phase. If there is a new update phase, repeat step 6; if not, end.

[0086] In the specific implementation process, three cooperative transmission links are introduced into the satellite-ground integrated cache network to provide file delivery services for users to ensure that user file requests can be met, such as Figure 2 As shown. After that, by deriving the formulas for the data transmission volume and power consumption of the three links, an optimization problem is established with the cache content of the satellite and base station layers as optimization variables and the goal of maximizing the long-term system energy efficiency, and it is solved based on the deep deterministic policy gradient (DDPG) algorithm. In order to deal with the problem of sparse rewards for effectively improving energy efficiency in complex network systems, behavioral cloning technology is used to improve the algorithm training performance. At the same time, according to the asynchronous update mechanism of satellites and base stations, the optimization problem is decomposed into sub-problems of base station cache content update and satellite cache content update. Among them, the base station layer uses the DDPG algorithm to obtain the cache decision strategy under the current base station layer cache content distribution based on the predicted value of user request information within its update cycle. Similarly, the satellite DDPG algorithm optimizes the satellite cache decision strategy based on the current base station layer cache strategy and the predicted value of user requests within the update cycle. Finally, through iterative training of multiple update cycles, the two-level DDPG algorithm finally converges to an efficient satellite-ground integrated collaborative cache strategy.

[0087] Build a channel model:

[0088] The encoding cache technology is used to cache popular files, that is, files are allowed to be cached in the cache device in any proportion. All file requests in the defined scenario belong to the file library containing F files, that is, The file index is represented by i, and each file has the same size. Let k re To recover the original file, Refers to the cache matrix of base station n within the satellite update period T and the base station update period t, where Similar order Refers to the cache vector of the satellite within the satellite update period T, where

[0089] Base station-user link: In the downlink of the ground network, the effects of large-scale fading and small-scale fading are considered at the same time. Large-scale fading can be modeled as path loss fading. For small-scale fading, Rayleigh channel can be used for modeling. Therefore, the channel gain of the base station delivery channel can be expressed as:

[0090] g n =d -α |h n | 2 (1)

[0091] where h n represents the coefficient of the Rayleigh channel, which obeys the cumulative distribution k of the Rayleigh channel re Distribution function (CDF): Therefore, the effective transmission rate of base station n to users in the base station service area can be expressed as:

[0092]

[0093] Where P n ,W b respectively refer to the transmission power and channel bandwidth of base station n, N b Refers to the power spectral density of the noise in the ground network.

[0094] Note that the update scenario considered involves multiple time slots, h n Therefore, in order to ensure that the file encoding packets broadcast by the edge cache device through the wireless channel in the satellite-ground integrated cache network can be successfully transmitted to the user within a limited time, it is stipulated that the success rate of file transmission must be higher than the threshold δ, that is, Pr(t s2u ≤t 0 )≥δ, where t s2u Indicates the time required for the base station cache device to deliver all cached coded packets of a file, t 0 represents the maximum time allowed for a single file delivery. Therefore, according to equation (2), base station n is the time required for successfully transmitting all the coded packets of file i in the cache. The minimum power required is:

[0095]

[0096] in express The inverse function of Indicates the data volume of the coded packets of file i cached by base station n.

[0097] Satellite-user link: Satellite communication channels are affected by free space loss, antenna gain, shadow fading, and many other factors. Energy loss and shadow fading are the most important aspects. Specifically, the Wilbu model can be used to model energy loss, so the energy loss in the satellite-user link can be modeled as follows:

[0098]

[0099] Among them G t , G r are the antenna gains of the satellite and the user respectively, λ refers to the wavelength of the electromagnetic wave, H l Refers to the height of the satellite from the ground, f rain is the parameter of Wilbull distribution. For the shadow fading of satellite-user link, the shadow Rayleigh dispersion fading model takes into account both line-of-sight shadow fading (LOS) and multipath shadow fading, and is often used in wireless channel environments of various frequency bandwidths. Therefore, |h s | 2 The cumulative distribution function can be expressed as:

[0100]

[0101] where |h s | 2 represents the channel gain of shadow fading, Where 2b represents the average power of the dispersion part, Ω represents the average power of the LOS loss part, and c represents the Nakagami fading coefficient.

[0102] Considering that in ISTN, user equipment can directly communicate with satellites, the total gain of the satellite-user link is g s The combined effects of energy loss and shadow dispersion should be considered:

[0103] g s =G s |h s | 2 . (6)

[0104] Therefore, similar to the base station-user link, the satellite successfully transmits all the coded packets about file i in the cache. The minimum power required is:

[0105]

[0106] in express The inverse function of represents the data volume of the coded packet of satellite cache file i, W s Respectively represent the satellite transmission power and channel bandwidth, N s Represents the power spectral density of noise in satellite communications.

[0107] Gateway-user link: For the gateway-user link, in order to simplify the problem, only the downlink channel overhead is considered, and it is assumed that there is no packet loss during the uplink forwarding process. Therefore, similar to the satellite-user link, successful transmission The minimum power required is:

[0108]

[0109] where ε represents the processing and forwarding delay in the gateway-satellite shared link, represents the amount of data that file i needs to forward in the gateway-user link, that is,

[0110] Setting up the optimization problem

[0111] Based on the coordinated delivery of the three links, all user requests in the network can be guaranteed to be met. Therefore, the data transmission volume in the network is equivalent to the total user request data volume in the network, that is:

[0112]

[0113] in It represents the expected number of user requests in the kth file request phase in the satellite period T and the base station period t.

[0114] The energy consumption of the system only focuses on the power consumption of downlink file transmission, so the total energy consumption of the system can be expressed as the sum of the power consumption of the three links, that is:

[0115]

[0116] in and They represent the power overhead derived in the satellite period T and the base station period t respectively. P i sat and

[0117] This establishes the energy efficiency optimization problem in the satellite-ground integration scenario:

[0118]

[0119] Where T is the index of the satellite update cycle, and the satellite update cycle T contains W base station update cycles t, that is, Formula (12) represents the cache space limitation of the base station, formula (13) represents the cache space limitation of the satellite, and formula (14) represents that the maximum cache data amount of each file in any cache device is sufficient to restore the data amount of the source file, so as to reduce the invalid cache in the cache and further improve the cache utilization.

[0120] Optimization problem decomposition:

[0121] Considering that in ISTN, due to the different update cycles of base stations and satellites, the energy efficiency optimization problem can be regarded as a long-term and short-term update problem. Therefore, problem P(1) can be decomposed into two sub-problems according to the time scale:

[0122]

[0123] st(12),(13),(14)(16) and

[0124]

[0125] st(12),(13),(14)(18)

[0126] In this regard, an iterative algorithm inspired by the coordinate descent method (BCD) is designed to solve the update problem, that is, continuously following the following steps:

[0127] (1) Fixed Decision-making strategy to solve Make maximum.

[0128] (2) Fixed Decision-making strategy to solve Make maximum.

[0129] By continuously executing the above steps, the algorithm will gradually converge and obtain the optimization solution that can achieve the highest energy efficiency.

[0130] DDPG algorithm design:

[0131] 1) DDPG algorithm design for ground network (DDPG-SBS):

[0132] State: It consists of the current cache decision content and the predicted value of the user request information, namely:

[0133]

[0134] in is the predicted value of the file request probability of file i in cell n in satellite period T and base station period t.

[0135] Action: Action is the ratio of the amount of cached data for each file stored in each SBS. Combination, namely:

[0136]

[0137] in

[0138] Reward: The reward is expressed in terms of the short-term energy efficiency of the system, namely:

[0139]

[0140] 2) DDPG algorithm design for satellite networks (DDPG-SAT):

[0141] state:

[0142]

[0143] in Refers to state-based As input, the DDPG-SBS algorithm outputs the base station cache update decision during the satellite update time slot. in

[0144] Action: Action is represented by the ratio of the amount of cached data for each file stored in the satellite, i.e.:

[0145]

[0146] Reward: The reward is expressed in terms of the long-term energy efficiency of the system, namely:

[0147]

[0148] Behavioral cloning design:

[0149] The behavior cloning technology is used to pre-train the policy network of the DDPG algorithm. By imitating the expert strategy, DDPG can quickly obtain initial decision-making capabilities and provide a higher starting point for subsequent training. In order to use the behavior cloning technology in the DDPG algorithm training process, it is necessary to rewrite the input and output of the expert strategy into the form of state action. represents the cached decision returned by the expert strategy, where the input state in the pre-training of the DDPG-SBS and DDPG-SAT algorithms is:

[0150]

[0151] The DDPG-SBS algorithm, DDPG-SAT algorithm, and satellite-ground collaborative cache update algorithm combined with pre-training technology are shown in Tables 1, 2, and 3, where E sbs and E sat is the maximum number of training episodes.

[0152] Table 1 DDPG-SBS-BC algorithm

[0153]

[0154]

[0155] Table 2 DDPG-SAT-BC algorithm

[0156]

[0157] Table 3 DDPG-SBS-SAT-BC algorithm

[0158]

[0159] See also Figure 3 , Figure 3 The structural block diagram of a satellite-ground integrated cache update and cooperative transmission device according to an embodiment of the present invention is shown. The device 400 includes:

[0160] The network system construction module 410 is used to construct a satellite-ground integrated cache network system, establish three cooperative transmission links, namely, a base station-user link, a satellite-base station link, and a gateway-satellite shared link, based on the satellite-ground integrated cache network system, and establish a channel model for the three cooperative transmission links based on a probability density function;

[0161] The target optimization problem building module 420 is used for the channel model, i.e., based on the data transmission volume and power consumption of the three links, to establish the optimization problem P1 with the buffer contents of the satellite and base station layers as optimization variables and the goal of maximizing the long-term system energy efficiency;

[0162] The target optimization problem solving module 430 is used to solve the optimization problem P1 based on the deep deterministic policy gradient algorithm, and to quote the behavior cloning technology to improve the training performance of the deep deterministic policy gradient algorithm. During the solving process, according to the asynchronous update mechanism of the satellite and the base station, the optimization problem P1 is decomposed into the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2, wherein the subproblem P1.1 is based on the user request information prediction value within the base station update cycle, and the deep deterministic policy gradient algorithm is used to obtain the base station cache decision strategy under the current satellite cache content distribution; the subproblem P1.2 is based on the current base station cache strategy and the user request prediction value within the update cycle, and the deep deterministic policy gradient algorithm is used to obtain the base station cache decision strategy under the current satellite cache content distribution.

[0163] The algorithm obtains the satellite cache decision strategy;

[0164] The cache strategy acquisition module 440 is used to converge the deep deterministic policy gradient algorithm for the base station cache content update subproblem P1.1 and the star cache content update subproblem P1.2 through iterative training, and finally obtain the optimal satellite-ground integrated collaborative cache strategy.

[0165] In addition to the upper module, the device 400 may also include other components. However, since these components are irrelevant to the content of the embodiment of the present disclosure, their illustration and description are omitted here.

[0166] The other specific working processes of the satellite-ground integrated cache update and cooperative transmission device 400 refer to the description of the above-mentioned satellite-ground integrated cache update and cooperative transmission method embodiment, which will not be repeated here.

[0167] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the satellite-ground integrated cache update and collaborative transmission method as in any embodiment of the present invention.

[0168] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed, the satellite-ground integrated cache update and collaborative transmission method as described in any embodiment of the present invention is implemented.

[0169] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0170] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0171] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0174] In order to verify the effectiveness of the method of the present invention, a simulation experiment was carried out:

[0175] The scenario parameters are: Unless otherwise specified, a multi-base station distribution scenario with 6 base stations is considered, where the file library size is 20, the SBS cache size is set to 10% of the total file library data, and the satellite cache size is set to 20%.

[0176] DDPG algorithm parameters at the base station layer: The actor network is built using five fully connected layers, including one input layer, three hidden layers, and one output layer. The dimension of each layer is 800. The critic network is also built using five fully connected layers, but in order to improve the convergence speed, the dimensions from the input layer to the output layer are 800, 400, 200, 100, and 1 respectively. The learning rate of the actor network is kept at 0.00001, while the learning rate of the critic network is kept at 0.001. Cache pool R sbs We choose the capacity as 10000, the sampling batchsize as 32, the soft update parameter ρ as 0.005, the discount factor γ as 0.99, and the number of pre-training episodes as 500.

[0177] Satellite DDPG algorithm parameters: Similarly, the same parameters as the base station layer DDPG are used to build the satellite actor network and critic network, and the learning rate and other parameters are also kept consistent. Other scenario simulation parameters are shown in Table 4:

[0178] Table 4 Simulation parameters

[0179]

[0180] Simulation comparison and analysis:

[0181] Uniform Caching (UC): The cache device evenly distributes the amount of encoded data for each file to fill its cache space.

[0182] Most Popular File Placement (MP): The cache device places the most popular files based on the predicted value. The relative proportions of each file are stored.

[0183] DDPG: The DDPG-SBS-BC and DDPG-SAT-BC algorithms without the behavior cloning pre-training step are used as baseline algorithms, named DDPG-SBS and DDPG-SAT respectively.

[0184] Trust Region Constrained Optimization Algorithm (TR): It is identified as the most ideal algorithm to determine the upper bound of the performance of the proposed algorithm.

[0185] Algorithm combination: It is described in the format of "algorithm-base station layer / satellite-pre-training technology". For example, DDPG-SBS-BC+UC-SAT means that the DDPG-SBS-BC algorithm that has been pre-trained with behavioral cloning is applied at the base station layer, and the equal placement algorithm is adopted at the satellite layer.

[0186] Figure 4 The optimization performance of DDPG-SBS and DDPG-SAT algorithms is demonstrated, where the number of base stations is 2, the file library contains 5 files, and the cache capacity of the base station and satellite is 25% and 50% of the total file library, respectively. This experiment does not consider the variability of environmental parameters and real-time request information, and can be regarded as a pure cache allocation problem. Different from the optimization scenario of the traditional DDPG algorithm, ISTN involves the joint decision of two layers of cache devices. Therefore, when updating the allocation of one layer of cache devices, the influence of the cache strategy of the other layer of cache devices must be considered. The results show that no matter whether the other layer of devices adopts the most popular file placement strategy or the equal placement strategy, the proposed DDPG-SBS and DDPG-SAT algorithms can complete convergence within 800 screens. Compared with the MP-SBS+UC-SAT algorithm, the two algorithms achieve 9% and 4% performance improvements, respectively. In addition, compared with the TR algorithm, the performance of the DDPG-SBS algorithm is very close to the TR level, while the DDPG-SAT algorithm can achieve 70% of the gain of the TR algorithm.

[0187] Figure 5The performance comparison of algorithms such as DDPG-SBS-SAT-BC under different file library sizes is shown. The most popular file placement is adopted as the expert strategy imitated by the base station DDPG algorithm, and the equal-sharing algorithm is adopted as the expert strategy imitated by the satellite. The results show that with the increase of the total number of files in the file library, the energy efficiency optimization level of each decision algorithm generally shows a downward trend. The reason is that the storage space of the edge cache device is limited and it is difficult to meet the cache requirements of more files. It is worth noting that when the total number of file libraries is low (that is, the number of files is less than or equal to 10), the most popular file placement strategy adopted by the satellite layer can achieve higher energy efficiency optimization than the equal-sharing placement scheme. However, as the total number of file libraries increases, the energy efficiency performance of the equal-sharing placement strategy gradually outperforms the most popular file placement strategy. The reason for this phenomenon is consistent with the analysis of the base station cache space. When the number of files is large, the equal-sharing placement strategy can allocate cache resources more evenly, thereby improving the overall energy efficiency.

[0188] from Figure 5 It can be seen that compared with the UC-SBS+UC-SAT algorithm combination, the DDPG-SBS-SAT-BC, DDPG-SBS-BC+UC-SAT and MP-SBS+UC-SAT algorithm combinations achieved 49% to 59%, 31% to 48% and 26% to 30% performance improvements respectively, showing the significant advantages of these optimization algorithms under different conditions. In general, as the file library size increases, the DDPG-based algorithm shows strong adaptability and can better handle complex cache optimization problems.

[0189] Figure 6 The performance of algorithms such as DDPG-SBS-SAT-BC under different numbers of base stations is demonstrated. It can be seen that compared with the UC-SBS+UC-SAT algorithm combination, the DDPG-SBS-SAT-BC, DDPG-SBS-BC+UC-SAT and MP-SBS+UC-SAT algorithm combinations can maintain stable optimization capabilities under different numbers of base stations, achieving performance improvements of about 53%, 46% and 28% respectively.

[0190] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A satellite-ground integrated cache update and collaborative transmission method, characterized in that: The method comprises the following steps: Constructing a satellite-ground integrated cache network system, establishing three cooperative transmission links, namely, a base station-user link, a satellite-base station link, and a gateway-satellite shared link, based on the satellite-ground integrated cache network system, and establishing a channel model for the three cooperative transmission links based on a probability density function; Based on the channel model, an optimization problem P1 is established with the cache contents at the satellite and base station layers as optimization variables and the goal of maximizing the long-term system energy efficiency; The optimization problem P1 is solved based on the deep deterministic policy gradient algorithm, and the behavior cloning technology is used to improve the training performance of the deep deterministic policy gradient algorithm. During the solution process, according to the asynchronous update mechanism of the satellite and the base station, the optimization problem P1 is decomposed into the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.

2. Among them, subproblem P1.1 is based on the predicted value of user request information within the base station update cycle, and the deep deterministic policy gradient algorithm is used to obtain the cache decision strategy under the current base station cache content distribution; subproblem P1.2 is based on the current base station cache strategy and the predicted value of user requests within the update cycle, and the deep deterministic policy gradient algorithm is used to obtain the satellite cache decision strategy; Through iterative training, the deep deterministic policy gradient algorithm for the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 converges, and finally the optimal satellite-ground integrated collaborative caching strategy is obtained.

2. The method according to claim 1, characterized in that The specific implementation steps of the method include: Step 1: Based on the constructed satellite-ground integrated cache network system, the system energy efficiency expression is calculated using the channel parameters; Step 2: Initialize the network parameters of the deep deterministic policy gradient algorithm at the satellite and base station layers; Step 3: collect a section of historical user request information, select an expert strategy to generate a corresponding base station layer expert cache solution and a satellite expert cache solution, assemble them into {user request information, expert cache solution}, and store them in the base station layer expert experience cache pool and the satellite expert experience cache pool respectively; Step 4: Randomly extract a group of samples from the base station expert experience cache pool and the satellite expert experience cache pool, and use the back propagation mechanism in supervised learning to pre-train the Actor network in the deep deterministic policy gradient algorithm; Step 5: Determine whether the training round is greater than the set expert training round. If not, repeat step 3. If it reaches the set expert training round, the Actor network pre-training is considered complete, and the parameters of the Actor network are copied to the target Actor network. Step 6: Determine whether the current phase is a satellite update phase or a base station update phase. If it is a satellite update phase, proceed to step 7; if it is a base station update phase, proceed to step 11; Step 7: Collect the predicted value of user request information in the current satellite stage, as well as the current satellite deep deterministic policy gradient algorithm network parameters, satellite experience cache pool, and base station cache decision strategy; Step 8: Generate a base station cache plan based on the current base station cache decision strategy and the predicted value of the user request information, and form the input state of the satellite deep deterministic policy gradient algorithm with {current satellite cache plan, base station cache plan, user request information}, and obtain the output action, and observe the long-term network system energy efficiency as a reward after taking the action; Step 9: Put the experience samples {state, action, reward, next state} into the satellite experience buffer pool, and randomly extract a group of samples from it to train the deep deterministic policy gradient algorithm and update the deep deterministic policy gradient algorithm network parameters; Step 10: Determine whether the number of training times reaches the set maximum number. If not, repeat step 8; if the maximum number is reached, the current satellite deep deterministic policy gradient algorithm training is considered complete, output the current satellite cache decision content, and jump to step 15; Step 11: Collect the predicted value of user request information in the current base station stage, as well as the current base station deep deterministic policy gradient algorithm network parameters, base station experience cache pool and satellite cache solution; Step 12: {current satellite cache solution, base station cache solution, user request information} is used as the input state of the base station deep deterministic policy gradient algorithm, and the output action is obtained. After taking the action, the short-term network system energy efficiency is observed as a reward; Step 13: Put the experience sample {state, action, reward, next state} into the base station experience buffer pool, and randomly extract a group of samples from it to train the deep deterministic policy gradient algorithm, and update the deep deterministic policy gradient algorithm network parameters; Step 14: determine whether the number of training times reaches the set maximum number of times. If not, repeat step 11; if it reaches the set maximum number of times, it is determined that the current base station deep deterministic policy gradient algorithm training is completed, output the current base station cache decision content, and jump to step 15; Step 15: Determine whether there is no new update phase. If there is a new update phase, repeat step 6; if not, end.

3. The method according to claim 1, characterized in that The optimization problem P1, which aims to maximize the long-term system energy efficiency, is specifically expressed as: P(1): s.t.C1: C2: C3: Where T is the index of the satellite update cycle, and the satellite update cycle T contains W base station update cycles t, that is, Refers to the cache matrix of base station n within the satellite update period T and the base station update period t, represents the cache vector of the satellite within the base station update period t within the satellite update period T, Indicates the total amount of user request data in the network, represents the total energy consumption of the system, k re To recover the original file, represents the cache space of file i, C1 represents the cache space limit of the base station, represents the maximum cache space of the base station, C2 represents the cache space limit of the satellite, represents the maximum cache space of the satellite, C3 represents the maximum cache data volume of the file in any cache device is enough to restore the source file, N is the number of base stations, and all file requests in the definition scenario belong to the file library containing F files.

4. The method according to claim 3, characterized in that The base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 are expressed as follows: P1.1: stC1,C2,C3 P1.2: stC1,C2,C3.

5. The method according to claim 4, characterized in that The solution of the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.2 is to repeat the following steps until the algorithm gradually converges: fixed Decision-making strategy to solve Make maximum; fixed Decision-making strategy to solve Make maximum.

6. The method according to claim 1, characterized in that The deep deterministic policy gradient algorithm is used to obtain the base station cache decision strategy under the current satellite cache content distribution, including: Base station status: It consists of the current satellite cache decision content and the predicted value of the user request information, namely: in is the predicted value of the file request probability of file i in cell n in satellite period T and base station period t, Refers to the cache vector of the satellite within the base station update period t within the satellite update period T, Refers to the cache matrix of the base station in the previous update period t-1 within the satellite update period T; Base station action: The action is determined by the ratio of the cached data volume of each file stored in each base station. Combination, namely: in Base station reward: The reward is expressed using the short-term energy efficiency of the system, namely: Indicates the total amount of user request data in the network, Represents the total energy consumption of the system.

7. The method according to claim 1, characterized in that The satellite cache decision strategy is obtained using a deep deterministic policy gradient algorithm, including: Satellite status: in Representation based on state As input, the base station cache update decision at the satellite update time slot is obtained by using the deep deterministic policy gradient algorithm output. represents the cache vector of the base station in the last base station update cycle within the satellite cycle T-1, represents the cache vector of the satellite within the base station update period t within the satellite update period T-1, in is the mean of the predicted values ​​of the file request probability; Satellite Action: Action is the ratio of the amount of cached data per file stored in the satellite. Combination, namely: Satellite Rewards: Rewards are expressed in terms of the long-term energy efficiency of the system, namely: Indicates the total amount of user request data in the network, Represents the total energy consumption of the system.

8. A satellite-ground integrated cache update and collaborative transmission device, characterized in that: The device comprises: A network system construction module is used to construct a satellite-ground integrated cache network system, establish three cooperative transmission links, namely, a base station-user link, a satellite-base station link, and a gateway-satellite shared link, based on the satellite-ground integrated cache network system, and establish a channel model for the three cooperative transmission links based on a probability density function; The target optimization problem building module is used to establish the optimization problem P1 based on the channel model, taking the cache contents of the satellite and base station layers as optimization variables and aiming at maximizing the long-term system energy efficiency; The target optimization problem solving module is used to solve the optimization problem P1 based on the deep deterministic policy gradient algorithm, and the behavior cloning technology is used to improve the training performance of the deep deterministic policy gradient algorithm. During the solution process, according to the asynchronous update mechanism of the satellite and the base station, the optimization problem P1 is decomposed into the base station cache content update subproblem P1.1 and the satellite cache content update subproblem P1.

2. Among them, subproblem P1.1 is based on the user request information prediction value within the base station update cycle, and the deep deterministic policy gradient algorithm is used to obtain the cache decision strategy under the current base station cache content distribution; subproblem P1.2 is based on the current base station cache strategy and the user request prediction value within the update cycle, and the deep deterministic policy gradient algorithm is used to obtain the satellite cache decision strategy; The cache strategy acquisition module is used to converge the deep deterministic policy gradient algorithm of the base station cache content update subproblem P1.1 and the star cache content update subproblem P1.2 through iterative training, and finally obtain the optimal satellite-ground integrated collaborative cache strategy.

9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the satellite-ground integrated cache update and collaborative transmission method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the satellite-ground integrated cache update and collaborative transmission method as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Satellite cooperative caching and user access method based on reinforcement learning

    CN117879680A

  • NOMA enabled air-to-ground content delivery network track and resource optimization method

    CN118741532A

  • Satellite-ground collaborative edge network resource allocation method based on depth deterministic strategy gradient

    CN119109504A

  • Computation offloading method and communication apparatus

    US20230081937A1