A dynamic cache updating method for a star-ground fusion network

By constructing a three-layer caching architecture of LEO satellite-base station-core network gateway in the space-ground converged network, and combining a multi-content sub-library model and a communication link model, and using meta-reinforcement learning and the MATD3 algorithm to optimize cache updates, the transmission capacity and cooperation problems of traditional caching strategies in the face of data traffic growth are solved, and efficient cache updates and service quality improvement are achieved.

CN119997054BActive Publication Date: 2025-11-25HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510091871.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-11-25
Estimated Expiration
2045-01-21

Smart Images

  • Figure CN119997054B_ABST
    Figure CN119997054B_ABST
Patent Text Reader

Abstract

The application provides a kind of dynamic cache updating method for star-ground fusion network, comprising: step one: according to the position information of LEO satellite, the time-varying association matrix of LEO satellite-base station is established, and the cache architecture network model of three layers of ground base station-LEO satellite-core network gateway is constructed based on the time-varying association matrix;Step two: construct a multi-content sub-library model to represent the popularity of content under different satellite coverage areas;Step three: construct a communication link model between satellites and base stations in the network, and establish a performance index to measure the pros and cons of the satellite-ground cache strategy;Step four: propose a star-ground and inter-satellite adaptive threshold collaborative cache strategy based on meta-reinforcement learning, learn the change characteristics of content popularity between different coverage areas, and realize cache updating during area switching based on the content feature similarity between satellite coverage areas and the candidate content index set at the satellite node.The beneficial effects of the application are: improving the generalization ability of deep reinforcement learning and improving network energy efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of edge caching technology, and in particular to a dynamic cache update method for satellite-ground fusion networks. Background Technology

[0002] With the development of communication technology, wireless communication has brought users increasingly efficient and convenient network services, accompanied by an explosive growth in data content on the network. It is predicted that mobile data traffic will grow at a compound annual growth rate of 23% between 2023 and 2030. By the end of 2030, mobile user penetration is expected to reach 74% of the global population, with the number of global mobile users projected to increase to 6.3 billion, and data traffic exceeding 465 EB per month. This unprecedented growth in mobile data traffic poses a significant challenge to traditional cellular networks in providing high throughput and multi-access services with limited backhaul link capacity. Due to the limitations of terrestrial network spectrum resources and energy efficiency, traditional terrestrial networks with buffering capabilities cannot cope with the anticipated surge in data. Satellite networks offer advantages such as broadcast transmission and vast coverage; therefore, the concept of STIN, combining terrestrial and satellite networks, was proposed, viewing satellite networks as a supplement to terrestrial networks to fill the gaps in buffered content transmission. However, satellite network resources are very limited, and onboard payloads are constrained by hardware conditions, such as limited spectrum and communication energy, making it difficult to increase network transmission capacity. At the same time, STIN faces numerous challenges in handling a large volume of user content requests, placing significant pressure on maintaining service quality. Therefore, providing high-quality content delivery services has become a critical issue that has attracted widespread attention.

[0003] Most existing studies consider pre-known content popularity to characterize the request preferences of ground users. Since user request preferences vary spatially due to the dynamic coverage of geographical areas by satellites, and also change dynamically over different time intervals (i.e., exhibiting temporal variability), constant content popularity models are likely to lead to mismatches between theory and practice. Therefore, traditional caching strategies are insufficient to provide good edge caching performance. On the other hand, most existing studies adopt a centralized approach to update caching decisions, such as using Medium Earth Orbit (MEO), Geostationary Earth Orbit (GEO) satellites or ground stations as centralized controllers to make caching decisions for all nodes in the network. Centralized methods use a single controller, which, given the highly dynamic connectivity and limited coverage of STIN, may lead to unreliable caching decision update processes, generating additional communication overhead and significant computational costs. However, considering the heterogeneity of network nodes, each node only observes a local environment. In this case, a fully distributed structure cannot fully demonstrate the advantages of inter-node collaboration. Therefore, when updating the cache state of nodes in STIN, information asymmetry can be eliminated first by synchronizing information between nodes, and then a distributed approach can be adopted to carry out the cache state update work, so as to take into account both the local characteristics of each node and the overall collaborative needs. Summary of the Invention

[0004] To address the problems in the existing technology, this invention provides a dynamic cache update method for satellite-ground converged networks, comprising the following steps: Step 1: Establish a time-varying correlation matrix between LEO satellites and base stations based on the location information of LEO satellites, and construct a three-layer cache architecture network model of ground base station-LEO satellite-core network gateway based on the time-varying correlation matrix between LEO satellites and base stations.

[0005] Step 2: Construct a multi-content sub-library model to represent the popularity of content under different satellite coverage areas; each content sub-library contains multiple contents, representing a certain type of content. Users in different regions will prefer the content in different sub-libraries. The content popularity distribution in each region follows a Zipf distribution. There are differences in preferences between different regions. Some sub-libraries' content is only popular in some regions.

[0006] Step 3: Construct a communication link model between satellites and base stations in the network, and establish performance indicators to measure the effectiveness of satellite-to-ground caching strategies;

[0007] Step 4: First, propose a satellite-to-ground and inter-satellite adaptive threshold collaborative caching strategy based on meta-reinforcement learning to quickly learn the changing characteristics of content popularity between different coverage areas. Then, based on the similarity of content features between satellite coverage areas and the set of alternative content indexes at satellite nodes, realize fast cache updates when switching regions.

[0008] As a further improvement of the present invention, step one further includes:

[0009] Step 1: Use a size S J ×B I matrix G t To describe the association status between each time slot satellite and the base station, it is represented as:

[0010]

[0011] in Indicates that in time slot t, satellite S j With base station B i The connection relationship is shown in Figure 1, indicating a connection. Satellite S j With base station B i A value of 0 indicates an associated state, while 0 indicates the opposite.

[0012] Step 2: Deploy caches on both the base station and satellite sides, with a cache capacity of [capacity to be specified] for each base station. The buffer capacity of each satellite is The base station's buffer state is represented as The satellite's cache status is represented as and These represent base station B. i and LEO satellite S j The c-th content is stored in the cache.

[0013] As a further improvement of the present invention, step two specifically involves: dividing the content set into M content sub-libraries to represent different types of content, the set being represented as:

[0014]

[0015] Users in each region may be interested in content from one or more content sub-repositories, a content sub-repository There are multiple types of files, and the popularity ranking of each piece of content changes over time, but the probability of a user requesting content in each sub-library always follows probability P. m , is represented as:

[0016]

[0017] Content popularity It follows a Zipf distribution, where i is the popularity ranking of the content;

[0018] In step two, time is divided into multiple discrete time slots, denoted as t = 1, 2, ..., T. In each time slot t, users within the region generate content requests based on their content preferences. Satellite S j The maximum number of user requests received is the sum of the number of users in each coverage area.

[0019] In step two, the method of transmitting content to the user uses a layered service architecture. First, the local base station responds. If there is no corresponding content, the request is forwarded to the associated satellite. Finally, the remote gateway provides the corresponding content. The base station and LEO satellite update the cached content through the gateway connected to the remote core network.

[0020] As a further improvement of the present invention, step three includes:

[0021] Step 1: Model the downlink feeder link, uplink feeder link, and terrestrial wireless transmission link, and obtain the unit content transmission delay τ of the base station. b The unit content transmission delay τ of LEO satellites s The unit content transmission delay τ of the gateway g and uplink power loss P UL and the power loss P of the terrestrial wireless propagation link BS ;

[0022] Step 2: Based on the cache architecture network model in Step 1 and the multi-content sub-library model in Step 2, calculate the latency reduction gain for ground base stations and LEO satellites;

[0023] Step 3: Calculate the symmetric difference between the cache states of the two time slots to obtain the number of updated contents. The expression is:

[0024] D t =|x t Δ x t-1 | (22)

[0025] Subsequently, LEO satellite S was obtained. j and ground base station set The formula for calculating cache cost is:

[0026]

[0027] Then use P respectively local,t and E local,t This represents the caching cost and utility of the local network to which the current agent belongs:

[0028]

[0029] Where ξ represents the proportionality coefficient. Indicates the buffer status of the base station. This indicates the satellite's cache status.

[0030] As a further improvement of the present invention, step 2 specifically includes: first, defining a function I(x,y) to determine whether element x is in vector y, the expression being:

[0031]

[0032] Base Station B i User requests received in time slot t Cache status is Therefore, base station B is obtained. i The delay reduction gain is:

[0033]

[0034] according to The result is related to satellite S j Cooperative base station set After a user's request receives a service response from the base station layer, any content requests that do not receive a response are forwarded to the satellite S. j ,Right now From Equation 11, we get:

[0035]

[0036] LEO satellite cache status is Therefore, satellite S is obtained. j The delay reduction gain is:

[0037]

[0038] Subsequently, the satellite S in time slot t was obtained. j With base station set The delay reduction gain is:

[0039]

[0040] As a further improvement of the present invention, in step four, during the meta-training of meta-reinforcement learning, the outer loop iterates from the task set. A batch of tasks were sampled from the middle. The task is defined as user requests generated based on changes in the popularity of different types of content, simulating the regional popularity differences that the satellite will face, and the inner layer uses each sampled task in a loop. To train the model, the inner loop obtains K experience trajectories for each task, denoted as τ. j ={(s j,0,a j,0 ,r j,0 ,s j,1 ),(s j,1 ,a j,1 ,r j,1 ,s j,2 ),...,(s j,K-1 ,a j,K-1 ,r j,K-1 ,s j,K Finally, the outer loop performs cross-task learning based on the experience trajectories collected by the inner loop, and updates the agent's meta-parameters. and

[0041] As a further improvement of the present invention, in step four, when a satellite needs to switch regions, the satellite in the next region will accumulate the request status. Send to satellite S j The satellite will calculate the Manhattan distance between the cumulative request state vectors of the two regions, using this distance as a measure of the difference in content features between the regions. The expression is as follows:

[0042]

[0043] This reflects the differences in content characteristics between the two regions. The largest difference occurs when the user request states in the two regions are completely different, and the difference value in this case is denoted as W. max ,when When υ∈[0,1], v is a scaling parameter; the satellite's actions tend to explore, decreasing exploration as the reward value increases. f This represents the user's request, and f represents the content. The meaning is related to satellite S j The number of requests for different content received by the associated base station from users during the time interval from t0 to t is represented by a vector.

[0044] As a further improvement of the present invention, in step four, a set is used.

[0045] This represents the set of candidate content, which corresponds to the index stored on the satellite side. The upper limit of the size is defined as F alt The LEO satellite stores the corresponding content indexes of all previously received user requests in a collection. Then, a first-in-first-out (FIFO) algorithm is used as the set. The update algorithm, i.e., set When the size reaches the limit, the earliest content index that entered the collection is discarded first.

[0046] As a further improvement of this invention, the MATD3 algorithm is used in the inner loop of the meta-training to perform few-sample learning on different distributions. The specific steps are as follows:

[0047] Step S1: Initialize the meta-policy parameters and the Critic network parameters θ π θ 1 and θ 2 Discount factor γ, outer learning rate η, inner learning rate ε, task distribution

[0048] Step S2: Sample each task using the outer loop. Training the model;

[0049] Step S3: Use strategy parameters Collect experience trajectory τ j ;

[0050] Step S4: Through formula R t =E agent,t +γ1E local,t +γ2E total,t Calculate the reward R on the experience trajectory according to the formula Calculate the loss function of the Critic network;

[0051] Step S5: Using stochastic gradient descent Update parameter θ 1 and θ 2 ε represents the inner learning rate. Represents the loss function. This represents the network parameters within the inner loop;

[0052] Step S6: According to the formula and Calculate gradient update parameters θ π , Indicates the strategy parameters, ε represents the gradient, and ε represents the inner learning rate.

[0053] Step S7: Perform soft updates on the target network parameters, where k = 1, 2, These are the meta-parameters of the agent. These are the parameters of the target network. Since there are two Critic networks, Critic 1 and Critic 2, each with its own network parameters corresponding to θ. 1 and θ 2 Here, k represents 1 and 2.

[0054] As a further improvement to this invention, the Reptile method, improved based on meta-learning, is used to train the network model, as expressed in:

[0055]

[0056] Where m is The number of tasks in the middle layer, η is the outer learning rate, a is the actor network, and c is the critic network.

[0057] The beneficial effects of this invention are: the dynamic cache update method of this invention fully considers the cooperation between satellites and between satellites and ground in scenarios with highly heterogeneous content popularity, uses meta-learning to improve the generalization ability of deep reinforcement learning, uses multi-agent reinforcement learning to improve the cooperation of distributed systems, and makes full use of the local environmental information observed by the agents to construct an adaptive threshold, effectively improving network energy efficiency. Attached Figure Description

[0058] Figure 1 This is a flowchart of the dynamic cache update method of the present invention;

[0059] Figure 2 This is the STIN network model of the present invention;

[0060] Figure 3 This is a schematic diagram of the heterogeneous user content preference area of ​​the present invention. Detailed Implementation

[0061] To improve network service quality with limited caching resources, this invention focuses on researching multi-node collaborative caching technology for Satellite-Terrestrial Integrated Networks (STINs) in scenarios with spatiotemporally heterogeneous content popularity. First, considering the spatiotemporal heterogeneity of content popularity, the invention primarily considers caching optimization strategies for subnetworks composed of one satellite and multiple base stations over a given period. Then, considering the dynamic nature of connections between satellites and base stations, as well as the richness of data content and the regional heterogeneity of popularity, the invention investigates satellite-to-ground and inter-satellite collaborative caching strategies in multi-satellite scenarios to alleviate backhaul link pressure, reduce service latency, and improve overall network energy efficiency.

[0062] In the context of satellite-ground converged network scenarios, this invention provides an algorithm for collaborative caching strategy design based on Meta-RL to address the caching strategy design problem for satellites and base stations. The aim is to design a distributed caching strategy for satellites and base stations when STIN provides content services to users under highly heterogeneous content popularity, thereby improving the generalization ability of the caching strategy and reducing network service latency and cache update costs.

[0063] like Figure 1As shown, this invention discloses a dynamic cache update method for satellite-ground fusion networks, comprising the following steps:

[0064] Step 1: Establish a time-varying correlation matrix between LEO satellites and base stations based on the location information of LEO (low Earth orbit) satellites. Construct a three-layer caching architecture network model of ground base station-LEO satellite-core network gateway based on the time-varying correlation matrix between LEO satellites and base stations. Since the caching capacity of satellites and base stations is limited, a caching model for satellites and base stations is constructed, and the corresponding cached content is represented by maintaining a fixed-size content set.

[0065] Step Two: Construct a multi-content sub-library model to represent content popularity across different satellite coverage areas. Since the amount of data content in the real world is massive and distributed across different regions, each content sub-library contains multiple pieces of content, representing a specific type of content. Users in different regions will prefer content from different sub-libraries, and the content popularity distribution in each region follows a Zipf distribution. Preferences vary significantly between different regions, and content from some sub-libraries may only be popular in certain regions.

[0066] Step 3: Construct a communication link model between satellites and base stations in the network, and then establish performance indicators to measure the effectiveness of the satellite-to-ground caching strategy. These performance indicators mainly include content transmission latency, cache update cost, and network energy efficiency. When a user sends a content request to a base station, the base station first checks its local cache for the requested content. If the content is cached locally, the base station can directly send the content to the user with low latency. In this case, the content transmission latency only includes the transmission latency from the base station to the user; this situation is also called a cache hit. If the base station cannot find the corresponding content in its local cache, it checks the caches of associated satellites. If an associated satellite has the requested content cached, it sends the corresponding content to the user. If neither the base station's cache nor its associated satellites have cached the requested content, the content needs to be retrieved from the core network server to serve the user, which leads to greater latency. For these three scenarios, the corresponding content transmission latency is calculated to obtain the network service latency. Satellites and base stations update cached content via uplink feeder links and ground backhaul links, respectively. The uplink feeder link mainly considers free space propagation loss and assumes that the base station is located in a remote access network. Both the satellite and the base station will incur non-negligible energy loss when updating the cache, and this energy loss will be regarded as the cost of cache update.

[0067] Step 4: First, a satellite-to-ground and inter-satellite adaptive threshold collaborative caching strategy based on meta-reinforcement learning (Meta-RL) is proposed to quickly learn the changing characteristics of content popularity between different coverage areas. Then, based on the similarity of content features between satellite coverage areas and the set of alternative content indexes at satellite nodes, fast cache updates are achieved when switching regions.

[0068] Due to the dynamic nature of satellites and the high heterogeneity of content popularity across regions over a large timescale, satellite caching strategies constantly face changing content popularity distributions. Retraining the neural network model every time a satellite reaches a new region would require a large number of training rounds. Therefore, Meta-RL is used to quickly learn the changing characteristics of content popularity across different coverage regions. In the inner loop of the meta-training, the Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm is used for few-sample learning of different distributions, while the outer loop implements cross-task learning. Subsequently, co-orbital satellite collaboration is used to calculate the content popularity difference between two regions before and after a region switch for the LEO satellite, determining the similarity of content features between the two regions. An edge-side candidate content index set is then constructed for the LEO satellite, reducing the dimensionality of the action space and enabling rapid cache updates during region switches.

[0069] Specifically, the algorithm models the collaborative caching strategy optimization problem as a Markov decision process and constructs a reward function based on the objective function. The reward function is built based on the local information of each cache node and the global network information, minimizing service latency and cache update costs for users in the system, and maximizing network energy efficiency, i.e., system reward. The algorithm first enhances the model's generalization ability based on Model-Agnostic Meta-Learning (MAML), enabling edge nodes to flexibly adapt to various content popularity changes, and then uses MATD3 to complete the collaborative caching strategy optimization. Utilizing information sharing between satellites in the same orbit, the algorithm calculates the content popularity difference between the two regions before and after a LEO satellite region switch, determining the similarity of content features between the two regions. To reduce the spatial dimension of action selection and better reflect reality, a satellite-side content index library with fewer contents than the information source and tailored to the needs of the satellite service area is constructed, thereby improving network energy efficiency and enabling rapid convergence of the caching strategy.

[0070] This invention considers the algorithm design of a satellite-ground and inter-satellite cooperative caching strategy based on meta-reinforcement learning in satellite-ground fusion network scenarios:

[0071] I. System Model

[0072] (1) Network Model

[0073] The research scenario of this invention includes multiple regions, such as Figure 2 As shown, a network model is constructed based on this scenario. The set of base stations is represented as... LEO satellite set is represented as To maintain their orbital altitude, LEO satellites orbit the Earth at a speed of approximately 7.8 km / s, typically completing one orbit in 1.5 to 2 hours. During their orbit, LEO satellites only cooperate with ground stations within their coverage area to serve users; that is, the ground stations currently being served by the satellite S... j The covered area will be covered by satellite S in the future. j+1 Coverage. Therefore, the relationship between LEO satellites and ground base stations is constantly changing. This invention uses a space of size S. J ×B I matrix G t To describe the association status between each time slot satellite and the base station, it is represented as:

[0074]

[0075] in Indicates that in time slot t, satellite S j With base station B i The connection relationship is shown in Figure 1, indicating a connection. Satellite S j With base station B i A value of 0 indicates an associated state, while 0 indicates an unassociated state. To compare the similarity of user preferences across different regions, LEO satellites must communicate not only with associated ground base stations but also with satellites on the same orbit that are one hop apart.

[0076] Caches are deployed on both the base station and satellite sides, with a cache capacity of [capacity] for each base station. The buffer capacity of each satellite is Because the network service range of this invention expands from a local area to the satellite's orbital coverage area, and the coverage areas of each satellite in the same orbital period do not overlap with those of other satellites, this means that each user can only be served by one satellite at any given time. Since user preferences differ across regions, and the data content is complex and massive, when a satellite enters an area where user preferences differ significantly from before, there will be a "blind spot" for the caching strategy, making it impossible to predict in advance. Therefore, the caching status of the base station and satellite is represented as follows: and and These represent base station B. i and LEO satellite S j The c-th content is stored in the cache.

[0077] There is no need to consider the connection between the satellite and the ground gateway separately, because the time-varying correlation matrix between the satellite and the base station includes the correlation with the ground gateway.

[0078] (2) Content Request Model

[0079] Because LEO satellites orbit the Earth at a high speed of approximately 7.8 km / s, they periodically provide content services to different target areas along their orbits. As the satellites orbit, they periodically provide content services to different target areas. The popularity of content is often strongly correlated with regional location. For example... Figure 3 As shown, for ease of representation, regular hexagons are used to represent coverage cells, and a satellite coverage area consists of multiple cells. This spatial diversity is simulated by dividing the coverage area along the orbit into N regions with different content popularity, where different colors represent different user-preferred content. Note that a large satellite beam may cover multiple regions with different local content popularity distributions.

[0080] In the real world, network data accumulates over time, but much of this content becomes "cold content" over time—content that has been rarely or no longer requested by users recently. Therefore, this invention reasonably assumes that all data requested by users at any given time is stored in the cloud and can be accessed through the core network; the content library is represented as a set. To facilitate the representation of content types and user preferences in different regions, this invention divides the content set into M content sub-libraries to represent different types of content. The set can be represented as follows:

[0081]

[0082] Users in each region may be interested in content from one or more content sub-libraries. A content sub-library There are multiple types of files, and the popularity ranking of each piece of content changes over time, but the probability of a user requesting content in each sub-library always follows probability P. m , is represented as:

[0083]

[0084] Popularity of the content of this invention It follows a Zipf distribution, where i is the popularity ranking of the content.

[0085] Time is divided into multiple discrete time slots, denoted as t = 1, 2, ..., T. In each time slot t, users within the area generate content requests based on their content preferences. Assume the number of users in the coverage area of ​​each base station is U. i Each user initiates a request once per time slot, base station B i The received user request is represented as in This refers to the content requested by user u, which may belong to a content library. Any sub-library in the dataset. Based on the satellite-base station time-varying correlation matrix G. t The cooperative relationship between LEO satellites and base stations is represented as follows: It can be represented by satellite S j The number of cooperating base stations, i.e., the area covered. Satellite S j The maximum number of user requests that can be received is the sum of the number of users in each coverage area, which can be expressed as: Therefore, the user request received by the satellite is represented as in Indicates user u i The content requested.

[0086] The method of delivering content to users uses a layered service architecture. First, the local base station responds. If no corresponding content is available, the request is forwarded to the associated satellite. Finally, the remote gateway provides the corresponding content. The base station and LEO satellite can update the cached content through the gateway connected to the remote core network.

[0087] (3) Communication model

[0088] This communication model primarily studies information transmission resulting from content transmission between users, base stations, and satellites, as well as information transmission resulting from content transmission and cache state updates between base stations, satellites, and the core network. Therefore, the links involved include the terrestrial wireless transmission link between users and base stations, the power supply link between satellites and ground stations, and the backhaul link from the subnetwork to the core network.

[0089] The downlink loss of LEO satellites mainly considers free-space propagation loss and atmospheric loss. The altitude of the LEO satellite is regarded as the distance between the satellite and the user, and the specific expression is as follows:

[0090]

[0091] Since free-space propagation loss accounts for the majority of downlink feeder link loss, and atmospheric loss is mainly caused by absorption and scattering of gas molecules in the Earth's atmosphere, as well as meteorological factors such as rain attenuation and fog attenuation, the expression for downlink loss is:

[0092] L atm =A gas +A rain +A fog +... (41)

[0093] L DL =L FS +L atm (42)

[0094] Therefore, the received signal-to-noise ratio of the downlink feed link can be obtained as follows:

[0095] SNR DL (dB)=EIRP-L DL +G r -kT total -B s (43)

[0096] Where EIRP is the satellite's equivalent isotropic radiated power, G r It is the ground station receiving antenna gain, T total It is the total noise temperature, B s This refers to the satellite channel bandwidth. Using Shannon's formula, the maximum transmission rate per unit of content from a LEO satellite to a user can be obtained, expressed as:

[0097] R DL =B s log(1+SNR DL (44)

[0098] In addition to free-space propagation loss, terrestrial wireless transmission links must also consider large-scale fading and small-scale fading. Large-scale fading is typically described using the path loss exponential model, with the specific expression as follows:

[0099]

[0100] Where n is the path loss exponent and d is the distance between the base station and the user. Common small-scale fading includes Rayleigh fading and Ricean fading. Here, we consider Rayleigh fading, where the amplitude of the received signal follows a Rayleigh distribution, specifically expressed as:

[0101] P r,R =|h| 2 P r,LS (46)

[0102] Where h is the fading factor. This allows us to obtain the received power after fading, and the maximum transmission rate from the base station to the user can be calculated using Shannon's formula:

[0103]

[0104] Among them B b It is the channel bandwidth of the base station.

[0105] This invention is based on modeling the downlink feed link, uplink feed link, and terrestrial wireless propagation link. This invention can obtain the unit content transmission delay τ of the base station for each of these components. b The unit content transmission delay τ of LEO satellites s The unit content transmission delay τ of the gateway g and uplink power loss P UL and the power loss P of the terrestrial wireless propagation linkBS Since the information transmitted between satellites is feature information extracted from historical user request data, the amount of data is very small compared to the content requested by the user, so the latency and energy loss during inter-satellite information transmission are not considered.

[0106] Subsequently, based on the network model and the multi-content sub-library model, this invention calculates the latency reduction gain for ground base stations and LEO satellites. The latency reduction gain is the difference between the maximum transmission latency (all content is sent to the user through the core network gateway) and the transmission latency currently using edge caching nodes (base stations and satellites). The specific steps are as follows:

[0107] First, define a function I(x,y) to determine whether element x is in vector y, with the expression:

[0108]

[0109] Base Station B i User requests received in time slot t Cache status is From this, we can obtain base station B. i The delay reduction gain is:

[0110]

[0111] according to It can be concluded that the relationship with satellite S j Cooperative base station set After a user's request receives a service response from the base station layer, any content requests that do not receive a response are forwarded to the satellite S. j ,Right now This can be obtained from equation (49):

[0112]

[0113] LEO satellite cache status is From this, we can obtain satellite S j The delay reduction gain is:

[0114]

[0115] Subsequently, the satellite S in time slot t can be obtained. j With base station set The delay reduction gain is:

[0116]

[0117] The cache update cost primarily considers the energy loss incurred by satellites and base stations when retrieving new content from the core network. This mainly considers the uplink power supply link of the satellite and the terrestrial wireless transmission link from the base station to the core network. Taking the orbital altitude of the LEO satellite as the communication distance between the satellite and the core network, the energy loss mainly considers free-space propagation loss, expressed as:

[0118]

[0119] When updating the cache, both satellites and base stations must complete the update within a time slot interval t0. The scenario with the greatest energy consumption is when the current cache state is completely inconsistent with the content in the cache decision, i.e., a full replacement. Therefore, we calculate the maximum rate at which the core network transmits content to the satellite under the worst-case scenario, expressed as:

[0120]

[0121] Using Shannon's formula, the uplink power loss P can be solved. UL for:

[0122]

[0123] Therefore, the energy consumption for each satellite cache update can be calculated as follows:

[0124]

[0125] in This represents the amount of new content to be acquired. Similarly, the energy consumed by the base station to acquire new content is:

[0126]

[0127] Cache cost, or energy loss incurred during cache updates, is only related to the amount of content updated by the cache node, i.e., the difference between the cache state before and after. Since the elements in the cache state are unordered, and each piece of content is not cached repeatedly at the same time, the amount of updated content can be obtained by calculating the symmetric difference between the cache states of two consecutive time slots. The expression is:

[0128] D t =|x t Δ x t-1 | (58)

[0129] Subsequently, LEO satellite S can be obtained separately. j and ground base station set The formula for calculating cache cost is:

[0130]

[0131] Use P respectivelylocal,t and E local,t This represents the caching cost and utility of the local network to which the current agent belongs:

[0132]

[0133] ξ represents the proportionality coefficient.

[0134] A satellite will collaborate (associate) with several ground base stations over a period of time. Each of these edge nodes can be called an agent in the context of reinforcement learning, and these edge nodes form a small network. The local network to which the agent currently belongs is the network in which the agent currently resides.

[0135] Different resource attributes and performance indicators in a network have different impacts on the system, which are considered as system benefits and costs. The utility function of the target system can be expressed as the difference between system benefits and costs to measure the system. Specifically, G t This refers to the revenue generated from cached traffic, P. t total This refers to the power consumption of the caching system. This invention aims to maximize the long-term utility of the sub-network by using the difference between the network latency reduction gain and the network cache update cost as the instantaneous network utility, and then summing these over time to obtain the long-term utility.

[0136] (4) Optimization problem

[0137] Finally, this invention aims to maximize the long-term utility of STIN (Short-Range Infrared) systems. STIN typically consists of a large constellation of low-Earth orbit (LEO) satellites and a terrestrial network. The Walker constellation is a large constellation of polar-orbiting LEO satellites, with an equal number of satellites in each orbit and equidistant distribution. Therefore, research on caching strategies for LEO satellites in one orbit can be extended to other orbits. To achieve this goal, this invention constructs a utility optimization problem and considers the caching states of base stations and LEO satellites. As an optimization variable:

[0138]

[0139] Where ξ is the proportionality coefficient, and constraints C1 and C2 are the satellite's buffer capacity limits.

[0140] II. Problem Analysis and Solutions

[0141] (1) Problem Analysis

[0142] In scenarios with diverse content sub-libraries and heterogeneous user preferences, how can LEO satellites better perform content and model preheating when crossing top and cross regions?

[0143] Existing research often assumes that user-requested content comes from the same content set, and that the satellite can select content from this set for caching. Based on this assumption, all content requests are within the satellite's line of sight, allowing for clear selection of specific content for caching decisions. However, in real-world scenarios, content popular in one region may be considered "cold content" in another, falling into the satellite's caching strategy's blind spot. Before a cross-regional conflict occurs, the satellite has not received any related content requests or information; this region is an information silo for the satellite. Failure to anticipate or include content preferred by users in the new region leads to unmet user requests, degrading STIN's service quality and network efficiency. To improve STIN's network efficiency, this invention models this scenario as a multi-content sub-repository model, facilitating the construction of an optimization problem regarding caching decisions.

[0144] First, given the massive amounts of data generated daily on the internet, it's difficult to completely fill satellite blind spots with limited computing power and storage capacity. Therefore, this invention focuses on making caching strategies more adaptable to changing environments. Since base stations serve local users, user content preferences are relatively stable. The LEO satellite's orbital speed is much faster than the Earth's rotation speed. This invention assumes that during a satellite area handover, satellite S... j+1 The area covered in the current time period and S j The areas covered in the next time period are largely the same. Therefore, user request information collected by the two satellites can be used to measure the similarity of preferences between the two regions, and a set of alternative content indexes from the satellite side can be constructed based on historical request information. This set is defined as a dynamic collection of content, providing only an index corresponding to each piece of content, rather than the content itself. The purpose is to make the satellite aware of the content's existence. This set dynamically adds or removes content based on user requests when satellites perform area handovers. During the time the satellite is cooperating with ground base stations to transmit content to users, the content in the service set is considered unchanged. This will, to some extent, fill the blind spots of caching strategies and reduce computational overhead.

[0145] Subsequently, because the agent's policy model often becomes limited to the specific environment in which it was trained after multiple rounds of training in a single scenario, it suffers from poor generalization and may not be applicable to other tasks or slightly different environments, requiring retraining. Therefore, to ensure a good user experience, LEO satellites should achieve seamless switching, meaning they can quickly adapt to different scenarios.

[0146] (2) MDP problem transformation

[0147] In order to apply reinforcement learning-based algorithms, we first transform the optimization problem into an MDP problem, thereby restating the optimization problem.

[0148] In the satellite-to-ground buffer network, each base station and its cooperating satellites are considered as an agent. Each agent performs a corresponding action based on the currently observed state. Within this sub-network, the environment that each agent can observe is limited; for example, base station B... i The observed state is represented as in It is base station B i A content request from the user is received in time slot t. It is base station B i The buffer state in time slot t; similarly, the state observed by satellite S can be obtained. Let A t ={a1,a2,...a n} represents the action space of the agent, where each action a i This represents a content caching combination where each piece of content has two states under constraints: selected and not selected.

[0149] The agent performs an action a in each round. i To obtain reward R t Based on the objective function proposed in equation (51), the reward function of the agent is designed, and its expression is:

[0150] R t =E agent,t +γ1E local,t +γ2E total,t (64)

[0151] Since our objective is to improve the service quality of all LEO satellites in orbit and the base stations they cover, the reward consists of the agent's local network utility and global network utility, where γ1 and γ2 are weight parameters, E agent,t It is the current utility of the intelligent agent, E local,t It is the utility of the local network to which the current intelligent agent belongs, E total,t This is the utility of all current LEO satellites and coverage base stations.

[0152] (3) Algorithm design of collaborative caching strategy based on meta-reinforcement learning

[0153] The Meta-Reinforcement Learning (Meta-RL) proposed in this invention can obtain the STIN satellite-ground cooperative caching strategy, which requires only a few further training steps to achieve rapid updates. Compared with the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm, the proposed Multi-Agent Dual-Delay Deep Deterministic Policy Gradient (Meta-MATD3) algorithm enables agents to directly obtain the reward for the selected action from multiple task types, i.e., probability distributions, and reduces the possibility of overestimating value. The algorithm proposed in this invention is more suitable for complex scenarios with multiple content sub-databases, thereby achieving higher long-term network utility.

[0154] The meta-learning algorithm proposed in this invention consists of two stages: the first is an inner loop, the purpose of which is to find a good parameter for a specific task; the second is an outer loop, which combines the results of multiple inner loops to find a good initialization state for all tasks. The design of these two loops will be discussed below.

[0155] First, the goal of the inner loop is to solve a standard reinforcement learning problem. Here, we treat each cache node (including ground base stations and LEO satellites) as an agent. LEO satellite area switching is based on the distance between the satellite and the ground station, resulting in a satellite-base station correlation matrix, which is not used as input for reinforcement learning. Therefore, the environmental state observed by each agent includes received user request information and local cache state, which can be represented as S. t ={r t ,x t The agent's actions are combinations of content from the corresponding cell. The size of the action space is related to the number of available and cacheable content, and can be obtained by calculating the number of combinations. During the meta-training phase, the policy focuses on the differences in the probability distributions of the agent's task execution; therefore, it can be assumed that an index of candidate content serving the satellite already exists.

[0156] This invention uses the MATD3 algorithm training caching strategy, which employs a centralized training and distributed execution architecture, where each agent contains an Actor network π. i Two Critic networks and and the corresponding target network π i ′、 and Through communication between agents, the state, actions, and reward information of agents with cooperative relationships in each round are placed into their respective experience replay pools. During the model training phase, the agents randomly sample a batch of data from the experience replay pools and pass it through the target Actor network π. i 'Get action a'i The target Q-value is calculated using two target Critic networks:

[0157]

[0158] To reduce the bias in Q-value estimation, the minimum output from the two target networks is used in the calculation of the target Q-value, where γ is a discount factor. Subsequently, the loss function is calculated and the parameters of the two Critic networks are updated. and The loss function expression is:

[0159]

[0160] Where k = 1, 2. Finally, the Actor network is updated by maximizing the Q-value of the Critic network; here, only the network is used. Update the policy parameters with the value of The resulting gradient expression is:

[0161]

[0162] Wherein, J(π) i ) is the objective function of the Actor network. Finally, a soft update method is used to update the target network π. i ′、 and The parameters are updated.

[0163] The outer loop is designed to combine the results of multiple inner loops and find a good initialization for all these similar tasks. The first step of the outer loop is to start from the task set. A batch of tasks were sampled from the middle. The task is defined as user requests generated based on changes in the popularity of different types of content, simulating the regional popularity differences that the satellite will face. The inner loop uses each sampled task. This is used to train the model. Therefore, the K experience trajectories for each task can be obtained through the inner loop.

[0164] Let it be τ j ={(s j,0 ,a j,0 ,r j,0 ,s j,1 ),(s j,1 ,a j,1 ,r j,1 ,s j,2 ),...,(s j,K-1 ,a j,K-1 ,r j,K-1 ,s j,K )}.

[0165] After collecting experience trajectories, cross-task learning needs to be performed based on these trajectories to update the agent's meta-parameters. and Since the environment in reinforcement learning is often dynamic and unknown, the meta-optimization part uses policy gradients to estimate gradients. Based on MAML, the objective function for the meta-parameters can be obtained as follows:

[0166]

[0167] Furthermore, the gradient update of the meta-parameters can be obtained, expressed as:

[0168]

[0169] Where β a and β c This involves updating the step size. It can be seen that the MAML method requires precise calculation of gradients in the inner and outer loops, which introduces significant computational overhead. Therefore, this invention employs the Reptile (a first-order optimization algorithm) method, an improvement on MAML, to train the network model. This method significantly reduces computational complexity while still achieving rapid task adaptation. The expression is:

[0170]

[0171] Where m is The number of tasks in the learning process, η is the learning rate.

[0172] Since this invention addresses a problem involving multiple agents, for simplicity, θ is used. π The Actor network parameters for all agents are represented by θ. 1 and θ 2 The parameters of the two Critic networks representing all agents are shown below. The pseudocode for the proposed Meta-MATD3 algorithm is as follows.

[0173] The algorithm flow is as follows:

[0174]

[0175] In satellite communication systems, the dynamic nature of satellites causes their coverage areas to constantly change, necessitating the statistical analysis of differences in user preferences across different areas. The base station is responsible for monitoring and recording the request status within each time slot and transmitting this data to the local LEO satellite. The satellite's task is to record the frequency of different content occurrences within these requests, denoted as a vector. in One element represents the number of times a corresponding piece of content has been requested. The records from different time slots are then summed with weights, where fluctuations closer to the current time slot are given higher weights to reflect their impact on the current state. Considering that a single satellite may serve multiple base stations, the LEO satellite also needs to sum the data from different base stations and further calculate its average. The mathematical expression for this process is as follows:

[0176]

[0177] When a satellite needs to switch regions, the satellite in the next region will accumulate the request status. Send to satellite S j Based on this information, the satellite will calculate the Manhattan distance between the cumulative request state vectors of the two regions, using this distance as an indicator of the difference in popularity between the regions. The expression is as follows:

[0178] This reflects the differences in content characteristics between the two regions. The largest difference occurs when the user request states in the two regions are completely different, and the difference value in this case is denoted as W. max .when Greater than υ·W max When υ∈[0,1], the satellite's actions will tend towards exploration, and exploration will gradually decrease as the reward value increases.

[0179] Since the locations of ground base stations are fixed, the user groups they cover are also relatively fixed. Therefore, we can assume that users' content preferences are relatively stable, and obtain a fixed local content library containing the content requested by users during that period. This local content library consists of one or more content sub-libraries.

[0180] However, LEO satellites encounter content blind spots during their orbits, leading to a decrease in cache hit rate. Since the policy model is difficult to adjust quickly, it tends to continue caching content according to the previous strategy. Therefore, this invention proposes a satellite-side alternative content index library, allowing the satellite to continuously have new content available while controlling the size of the action space, keeping computational overhead within a manageable range. Because this library only stores the index corresponding to each piece of content, rather than the content itself, it does not consume a large amount of cache space.

[0181] The data content is unordered and unique, therefore sets are used. This represents a set of candidate content, corresponding to an index stored on the satellite side. The purpose is to inform the satellite of the existence of this content. (Set) The size is finite, and the upper limit is defined as F. altThe LEO satellite stores the corresponding content indexes of all previously received user requests in a collection. In the middle. Due to the set The number of stored indexes far exceeds the amount of content cached by the LEO satellites, and we do not want the update strategy of this set to incur excessive computational overhead. Furthermore, user preferences in the region do not exhibit periodicity. Therefore, this invention employs a First-In-First-Out (FIFO) strategy algorithm as the set. The update algorithm, i.e., set When the size reaches the limit, the earliest content index that entered the collection is discarded first.

[0182] The algorithm flow is as follows:

[0183]

[0184]

[0185] The invention is based on the following: (1) a satellite-base station caching network model is established based on the application background of the satellite-ground fusion caching network, and a satellite-ground caching strategy optimization problem is constructed under this model for scenarios with highly heterogeneous content popularity; (2) a distributed caching strategy based on meta-reinforcement learning for satellite-ground and inter-satellite cooperation is proposed to minimize network service latency and cache update cost.

[0186] The beneficial effects of this invention are as follows: The dynamic cache update method of this invention fully considers the cooperation between satellites and between satellites and ground in scenarios with highly heterogeneous content popularity, uses meta-learning to improve the generalization ability of deep reinforcement learning, uses multi-agent reinforcement learning to improve the cooperation of distributed systems, and makes full use of the local environmental information observed by the agents to construct an adaptive threshold (a method for calculating regional content feature differences), which effectively improves network energy efficiency.

[0187] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A dynamic cache update method for satellite-ground fusion networks, characterized in that, Includes the following steps: Step 1: Establish a time-varying correlation matrix between LEO satellites and base stations based on the location information of LEO satellites, and construct a three-layer caching architecture network model of ground base station-LEO satellite-core network gateway based on the time-varying correlation matrix between LEO satellites and base stations. Step 2: Construct a multi-content sub-library model to represent the popularity of content under different satellite coverage areas; each content sub-library contains multiple contents, representing a certain type of content. Users in different regions will prefer the content in different sub-libraries. The content popularity distribution in each region follows a Zipf distribution. There are differences in preferences between different regions. Some sub-libraries' content is only popular in some regions. Step 3: Construct a communication link model between satellites and base stations in the network, and establish performance indicators to measure the effectiveness of satellite-to-ground caching strategies; Step 4: First, propose a satellite-to-ground and inter-satellite adaptive threshold collaborative caching strategy based on meta-reinforcement learning to quickly learn the changing characteristics of content popularity between different coverage areas. Then, based on the similarity of content features between satellite coverage areas and the set of alternative content indexes at satellite nodes, realize fast cache updates when switching regions. In step four, during the meta-training of meta-reinforcement learning, the outer loop iterates from the task set. A batch of tasks were sampled from the middle. The task is defined as user requests generated based on changes in the popularity of different types of content, simulating the regional popularity differences that the satellite will face, and the inner layer uses each sampled task in a loop. To train the model, the inner loop obtains K experience trajectories for each task, denoted as τ. j ={(s j,0 ,a j,0 ,r j,0 ,s j,1 ),(s j,1 ,a j,1 ,r j,1 ,s j,2 ),...,(s j,K-1 ,a j,K-1 ,r j,K-1 ,s j,K Finally, the outer loop performs cross-task learning based on the experience trajectories collected by the inner loop, and updates the agent's meta-parameters. and 2. The dynamic cache update method according to claim 1, characterized in that, Step one also includes: Step 1: Use a size S J ×B I matrix G t To describe the association status between each time slot satellite and the base station, it is represented as: in Indicates that in time slot t, satellite S j With base station B i The connection relationship is shown in Figure 1, indicating a connection. Satellite S j With base station B i A value of 0 indicates an associated state, while 0 indicates the opposite. Step 2: Deploy caches on both the base station and satellite sides, with a cache capacity of [capacity to be specified] for each base station. The buffer capacity of each satellite is The base station's buffer state is represented as The satellite's cache status is represented as and These represent base station B. i and LEO satellite S j The c-th content is stored in the cache.

3. The dynamic cache update method according to claim 1, characterized in that, Step two specifically involves dividing the content set into M content sub-libraries to represent different types of content, represented as follows: Users in each region may be interested in content from one or more content sub-repositories, a content sub-repository There are multiple types of files, and the popularity ranking of each piece of content changes over time, but the probability of a user requesting content in each sub-library always follows probability P. m , is represented as: Content popularity It follows a Zipf distribution, where i is the popularity ranking of the content; In step two, time is divided into multiple discrete time slots, denoted as t = 1, 2, ..., T. In each time slot t, users within the region generate content requests based on their content preferences. Satellite S j The maximum number of user requests received is the sum of the number of users in each coverage area; In step two, the method of transmitting content to the user uses a layered service architecture. First, the local base station responds. If there is no corresponding content, the request is forwarded to the associated satellite. Finally, the remote gateway provides the corresponding content. The base station and LEO satellite update the cached content through the gateway connected to the remote core network.

4. The dynamic cache update method according to claim 1, characterized in that, Step three includes: Step 1: Model the downlink feeder link, uplink feeder link, and terrestrial wireless transmission link, and obtain the unit content transmission delay τ of the base station. b The unit content transmission delay τ of LEO satellites s The unit content transmission delay τ of the gateway g and uplink power loss P UL and power loss P of terrestrial wireless propagation link BS ; Step 2: Based on the cache architecture network model in Step 1 and the multi-content sub-library model in Step 2, calculate the latency reduction gain for ground base stations and LEO satellites; Step 3: Calculate the symmetric difference between the cache states of the two time slots to obtain the number of updated contents. The expression is: D t =|x t Δx t-1 | (5) Subsequently, LEO satellite S was obtained. j and ground base station set The formula for calculating cache cost is: Then use P respectively local,t and E local,t This represents the caching cost and utility of the local network to which the current agent belongs: Where ξ represents the proportionality coefficient. Indicates the buffer status of the base station. Indicates the satellite's buffer status, G t This represents the revenue generated from traffic obtained through caching.

5. The dynamic cache update method according to claim 4, characterized in that, Step 2 specifically includes: First, defining a function I(x,y) to determine whether element x is in vector y, with the expression: Base Station B i User requests received in time slot t Cache status is Therefore, base station B is obtained. i The delay reduction gain is: according to The result is related to satellite S j Cooperative base station set After a user's request receives a service response from the base station layer, any content requests that do not receive a response are forwarded to the satellite S. j ,Right now From Equation 11, we get: LEO satellite cache status is Therefore, satellite S is obtained. j The delay reduction gain is: Subsequently, the satellite S in time slot t was obtained. j With base station set The delay reduction gain is:

6. The dynamic cache update method according to claim 1, characterized in that, In step four, when a satellite needs to switch areas, the satellite in the next area will accumulate the request status. Send to satellite S j The satellite will calculate the Manhattan distance between the cumulative request state vectors of the two regions, using this distance as a measure of the difference in content features between the regions. The expression is as follows: This reflects the differences in content characteristics between the two regions. The largest difference occurs when the user request states in the two regions are completely different, and the difference value in this case is denoted as W. max ,when When υ∈[0,1], v is a scaling parameter; the satellite's actions tend to explore, decreasing exploration as the reward value increases. f This represents the user's request, and f represents the content. The meaning is related to satellite S j The number of requests for different content received by the associated base station from users during the time interval t0 to t is represented by a vector. This represents a set of base stations.

7. The dynamic cache update method according to claim 1, characterized in that, In step four, sets are used. This represents the set of candidate content, which corresponds to the index stored on the satellite side. The upper limit of the size is defined as F alt The LEO satellite stores the corresponding content indexes of all previously received user requests in a collection. Then, a first-in-first-out (FIFO) algorithm is used as the set. The update algorithm, i.e., set When the size reaches the limit, the earliest content index that entered the collection is discarded first.

8. The dynamic cache update method according to claim 1, characterized in that, The MATD3 algorithm is used in the inner loop of the meta-training to perform few-sample learning on different distributions. The specific steps are as follows: Step S1: Initialize the meta-policy parameters and the Critic network parameters θ π θ 1 and θ 2 Discount factor γ, outer learning rate η, inner learning rate ε, task distribution Step S2: Sample each task using the outer loop. Training the model; Step S3: Use strategy parameters Collect experience trajectory τ j ; Step S4: Through formula R t =E agent,t +γ1E local,t +γ2E total,t Calculate the reward R on the experience trajectory according to the formula Calculate the loss function of the Critic network; Step S5: Using stochastic gradient descent Update parameter θ 1 and θ 2 ε represents the inner learning rate. Represents the loss function. This represents the network parameters within the inner loop; Step S6: According to the formula and Calculate gradient update parameters θ π , Indicates the strategy parameters, ε represents the gradient, and ε represents the inner learning rate. Step S7: Perform soft updates on the target network parameters, where k = 1, 2, These are the meta-parameters of the agent. These are the parameters of the target network.

9. The dynamic cache update method according to claim 1, characterized in that, The network model is trained using the Reptile method, which is an improvement based on meta-learning. The expression is as follows: Where m is The number of tasks in the middle layer, η is the outer learning rate, a is the actor network, and c is the critic network.

Citation Information

Patent Citations

  • Satellite network cache placement method based on regional user interest perception

    CN113472420A

  • Satellite Internet of Things online resource joint allocation method based on meta reinforcement learning

    CN115629540A