Base station clustering and service caching joint distribution method

By adopting the deep reinforcement learning framework of soft actor-criticist in the joint optimization of base station clustering and service cache, the shortcomings of multi-objective optimization problems in the existing technology are solved, efficient base station clustering and service cache allocation are achieved, and network performance and user experience are significantly improved.

CN120050715APending Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510231455.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology has shortcomings in the joint optimization of base station clustering and service cache, and has failed to effectively deal with multi-target needs such as delay, resource allocation and cache efficiency, and lacks a dynamic adjustment mechanism, making it difficult to meet the efficient response needs in complex network environments.

Method used

Adopting a deep reinforcement learning framework based on soft actor-criticist, we build a optimization problem of minimizing the total delay of users, and improve training convergence speed through improved adaptive discount rates and learning rates, design coordinated beamforming vectors to eliminate interference, and realize intelligent and autonomous base station clustering and service cache allocation.

Benefits of technology

It significantly improves network performance, resource utilization and user experience, and can efficiently respond to state changes in complex network environments, adapt to complex dynamic environments, reduce computing overhead, and improve cache hit rate and network energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050715A_ABST
    Figure CN120050715A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of edge service caching, and particularly relates to a base station clustering and service caching joint distribution method. The method comprises: constructing an intensive mobile edge network, the network comprising a plurality of base stations and a plurality of users, the base stations following Poisson distribution; based on the intensive mobile edge network, constructing an optimization problem of minimizing the total time delay of users; constructing a deep reinforcement learning framework based on a soft actor-commentator; solving the minimization user total time delay optimization problem according to a deep reinforcement learning framework based on soft actor-commentator to obtain an optimal base station clustering and service caching scheme and executing the optimal base station clustering and service caching scheme; according to the method, efficient response to state change can be realized in a complex network environment, and the method has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge service cache, and in particular relates to a base station clustering and service cache joint allocation method. Background Art

[0002] With the prosperity of the Internet era, more and more applications have been developed to meet users' various entertainment and functional needs. Various intelligent, big data analysis-based application software will generate computing-intensive tasks during operation. Generally, these tasks are processed on the user's terminal device and respond immediately. However, users have higher and higher requirements for performance experience, and the computing tasks generated are becoming more and more complex. For example, the most popular interactive applications based on virtual reality technology require not only ultra-low latency, but also a large amount of computing and storage to meet users' real-time and realistic interactive experience. It is difficult for users' terminal devices to support such high-resource applications. Therefore, many lightweight applications upload their generated computing tasks to the cloud for processing to reduce the configuration requirements for user devices. However, large data centers are often deployed in centralized geographical locations far away from users. User data needs to be transmitted through the core network before it can reach the cloud data center, and the transmission delay generated cannot be ignored. Fortunately, the emergence of multi-access edge computing (MEC) has brought a solution to this situation of resource and transmission contradiction.

[0003] MEC provides IT service environment and cloud computing capabilities at the edge of the network. It is generally deployed in the Radio Access Network (RAN), which is geographically closer to mobile users and gradually integrated into the evolving and increasingly perfect 5G architecture. Various applications with low-latency performance requirements will no longer be limited by the limited energy and computing power of terminal devices. Users can offload some computing tasks to nodes that support MEC at the edge of the network to obtain ultra-low latency responses. For example, applications based on augmented reality, virtual reality, 3D games, etc. require a lot of computing resources and energy consumption. Terminal devices send these computing tasks as requests to the network, and the network nodes forward the requests to nodes with MEC capabilities for processing, and finally return the results to the users.

[0004] In order to be able to handle the computing tasks generated by the application types required by users, MEC nodes need to deploy services before the tasks arrive. Service deployment refers to the process in which the computing node obtains the dependencies of the application from the cloud or other source and deploys them on the node so that it can run. Specifically, taking the image processing application based on deep learning as an example, processing an image requires the use of a trained neural network. The computing node downloads the neural network from the application provider and configures an executable environment for the program. Due to the prosperous development of the Internet industry, there are countless applications similar to image processing. Each application has its own specific dependencies, which are called services of this type of application. Deploying a service requires caching some dependent databases, function libraries and other data information, which will occupy storage resources; deploying the environment or software required to execute the application on the node will occupy computing resources and other resources that support computing needs. Therefore, deploying services involves the allocation and utilization of multiple resources. Furthermore, whether the deployed service is required by users and the frequency of use will determine the resource utilization efficiency brought about by this deployment decision. The MEC node needs to deploy services first so that the user's offloading task can be successfully executed, otherwise the entire offloading process cannot be started; and the high frequency and high dynamics of user requests require careful decision-making in the key process of service deployment to fully utilize the various resources of the MEC node, thereby providing users with continuous reliable performance guarantees.

[0005] There are still many shortcomings in the existing research on the joint optimization of base station clustering and service cache. These include the lack of research on multi-objective optimization problems, and the lack of a unified framework that can comprehensively handle multi-objective requirements such as latency, resource allocation, and cache efficiency. In the training phase, random experience extraction is generally used, and data with large TD errors are not prioritized, thereby reducing learning efficiency and convergence speed. Most exploration strategies use a fixed exploration probability and lack a dynamic adjustment mechanism, which cannot meet the requirements of neural networks for the degree of exploration at different training stages. In addition, the dynamic characteristics of base stations and service caches are not sufficiently considered, making it difficult to achieve efficient response to state changes in a complex network environment. These problems restrict the actual application effect and performance improvement of existing optimization methods. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention proposes a base station clustering and service cache joint allocation method, the method comprising:

[0007] S1: Build a dense mobile edge network, which includes multiple base stations and multiple users, where the base stations follow a Poisson distribution;

[0008] S2: Based on dense mobile edge networks, we construct an optimization problem to minimize the total user delay.

[0009] S3: Build a soft actor-critic based deep reinforcement learning framework;

[0010] S4: Solve the optimization problem of minimizing the total user delay according to the deep reinforcement learning framework based on soft actor-critic, obtain the optimal base station clustering and service caching scheme and execute it.

[0011] Preferably, the process of constructing the optimization problem of minimizing the total user delay includes:

[0012] S21: Calculate the task processing delay and uplink transmission delay of the user, and take the sum of the task processing delay and uplink transmission delay of the user as the total delay of the user;

[0013] S22: constructing constraint conditions, including base station cluster constraint, average cost constraint of user long-term cost, base station cache resource constraint, base station computing resource constraint, base station clustering strategy constraint and base station cache service type constraint;

[0014] S23: Determine and construct an initial optimization problem of minimizing the total user delay according to the total user delay and the constraint conditions;

[0015] S24: Based on Lyapunov optimization, the initial optimization problem of minimizing the total user delay is transformed into an instantaneous optimization problem with drift penalty as the optimization target, that is, the final optimization problem of minimizing the total user delay.

[0016] Furthermore, the initial optimization problem of minimizing the total user delay is expressed as:

[0017]

[0018] st:

[0019]

[0020] Among them, C(t) represents the clustering matrix of the base station at time t, X(t) represents the service cache matrix of the user at time t, T represents the total delay, represents the total task delay of user u at time t, c u,m (t) indicates whether base station m serves user u at time t, B indicates the total number of base stations in the base station cluster, Indicates whether base station m actually serves user u at time t, Cost m (t) represents the total cost of base station m at time t, Cost th represents the cache cost threshold, x k,m represents the probability that base station m caches service k at time t, s k represents the cache resources required for service k, S m represents the cache resources of base station m, Represents a collection of users, represents the base station set, represents the time slot set, f k represents the computing resources required for service k, C m represents the computing resources of base station m, Represents a collection of service types.

[0021] Furthermore, the optimization problem of minimizing the total user delay is expressed as:

[0022]

[0023] st

[0024]

[0025] Where C(t) represents the clustering matrix of the base station at time t, X(t) represents the service cache matrix of the user at time t, a(t) represents the task arrival rate of the queue at time t, b(t) represents the service rate of the queue at time t, and V represents the weight factor. represents the total task delay of user u at time t, c u,m (t) indicates whether base station m serves user u at time t, B indicates the total number of base stations in the base station cluster, and x k,m represents the probability that base station m caches service type k at time t, s k represents the cache resources required for service k, S m represents the cache resources of base station m, Represents a collection of users, represents the base station set, represents the time slot set, f k represents the computing resources required for service k, C m represents the computing resources of base station m, Represents a collection of service types.

[0026] Preferably, the process of solving the optimization problem of minimizing the total user delay based on the deep reinforcement learning framework of the soft actor-critic includes:

[0027] S41: Formulate the optimization problem of minimizing the total user delay as a Markov decision process, and define the state space, action space and reward function;

[0028] S42: Initialize the strategy network, value network and their corresponding target network;

[0029] S43: Each agent outputs an action from its policy network according to the current state. The agent interacts its action with the environment to obtain a new environment state and corresponding reward. The agent stores the current state, action and reward in the experience replay pool.

[0030] S44: Randomly sample a batch of experiences from the experience replay pool for learning;

[0031] S45: Update the Q value network and calculate the target Q value for each agent;

[0032] S46: Update the policy network parameters by maximizing the target Q value;

[0033] S47: softly updating the target network according to the parameters of the Q value network and the policy network;

[0034] S48: Repeat the training process from S35 to S37 until the agent converges or reaches the maximum number of iterations, and obtain the optimal clustering strategy and caching strategy.

[0035] Furthermore, the state space is expressed as:

[0036]

[0037] Among them, S(t) represents the state space at time t, Q u (t) represents the task queue status of user u, H u,m (t) represents the channel state information between user u and base station m, represents the cache status of base station m, represents the computing resource status of base station m, Represents a collection of users, Represents a set of base stations.

[0038] Furthermore, the action space is expressed as:

[0039]

[0040] Among them, A(t) represents the action space at time t, C(t) represents the clustering matrix at time t, and X(t) represents the service cache matrix at time t. Represents a collection of users, represents the base station set, Represents a collection of service types.

[0041] Furthermore, the reward function is expressed as:

[0042]

[0043] Among them, r(t) represents the reward at time t, C(t) represents the virtual cache cost queue, a(t) represents the task arrival rate of the queue at time t, b(t) represents the task service rate of the queue at time t, and V represents the weight factor. represents the total delay of user u at time t, Represents a collection of users.

[0044] The beneficial effects of the present invention are:

[0045] The present invention uses improved adaptive discount rates and learning rates to improve training convergence speed and network stability; the present invention eliminates interference within the cluster by designing coordinated beamforming vectors to obtain a higher transmission rate. The present invention takes into account factors such as the distribution of users and base stations in actual scenarios, the imperfection of channels, etc., and while realizing intelligent and autonomous base station clustering, it makes the service cache allocation under dense networks more in line with actual conditions; the present invention can achieve efficient response to state changes in complex network environments, and can significantly improve network performance, resource utilization and user experience. Its advantages are adapting to complex dynamic environments, efficiently processing continuous action spaces, reducing computational overhead, improving cache hit rates and network energy efficiency, and supporting large-scale network deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flow chart of the base station clustering and service cache joint allocation method in the present invention;

[0047] Figure 2 A mobile edge computing network scenario diagram in the present invention;

[0048] Figure 3 This is a diagram of the algorithm framework for base station clustering and service cache joint allocation in the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] The present invention proposes a base station clustering and service cache joint allocation method, such as Figure 1 As shown, the method includes the following contents:

[0051] S1: Build a dense mobile edge network, which includes multiple base stations and multiple users, where the base stations follow a Poisson distribution.

[0052] like Figure 2 As shown, a dense mobile edge network is constructed, specifically:

[0053] In this dense mobile edge network, base stations are deployed in the target area according to Poisson distribution, and user u moves in a random walk model. The user's movement characteristics include: the moving speed is satisfy The moving direction is satisfy

[0054] Channel gain is determined by path loss, large-scale shadow fading, small-scale Rayleigh fading, and beamforming gain. Path loss is determined by the distance from the base station to the user, which is manifested as macro signal attenuation; large-scale shadow fading gain is a fixed value β m,u , reflecting the influence of obstacles in the environment; small-scale Rayleigh fading is dynamically modeled using the following recursive formula: h m,u (t) = ρh m,u (t-1)+e m,u (t), where ρ = J 0 (2πf d T s ) is the correlation factor, f d is the maximum Doppler frequency, T s is the sampling time interval, is complex Gaussian noise, and the initial state is

[0055] Through the coordination mechanism within the base station cluster, all base stations share the channel state information (CSI) of the target user u to achieve joint optimization of beamforming.

[0056] Calculate the channel vector g of the target user u u (t) and the channel matrix G for other users in the cluster -u (t), channel vector g u (t) represents the channel between the target user u and the base station in the cluster, and the calculation formula is: Where M is the number of base stations in the cluster, β m,u is the path loss and large-scale shadow fading gain between base station m and user u, h m,u (t) is the small-scale Rayleigh fading coefficient; the channel matrix G for other users in the cluster -u (t) represents the channel set between non-target users and the base station, and its column vector is the channel vector of each non-target user.

[0057] Using g u (t) and G -u (t), design the zero-forcing beamforming vector w of the target user u (t), the zero-forcing beamforming vector calculation formula is: in is the identity matrix with dimension equal to the cluster size, It is a pseudo-reverse.

[0058] According to the beamforming vector w u (t), adjust the transmission power and direction of the base stations within the cluster to effectively suppress the interference within the cluster and enhance the signal quality of the target user. The independent channel gain is expressed as:

[0059] g m,u (t)=|w u H (t)g u (t)| 2 β m,u , where the beamforming gain is given by w u (t) Ensure that the channel characteristics are fully reflected by combining small-scale Rayleigh fading and path loss. The beamforming vector w is periodically recalculated based on the real-time updated channel state information (CSI). u (t) to adapt to user motion and channel changes in dynamic environments and continuously optimize system performance.

[0060] S2: Based on dense mobile edge networks, construct an optimization problem to minimize the total user delay.

[0061] S21: Calculate the task processing delay and uplink transmission delay of the user, and take the sum of the task processing delay and uplink transmission delay of the user as the total delay of the user.

[0062] Assume that there are K types of services stored in the remote cloud, and the service set is recorded as Different services consume different cache resources and computing resources when cached on BSs. Let the cache and computing resources required by service k be {s k ,f k}. We use the continuous variable x k,m ∈[0,1] represents the probability that service k is cached on BS m at time slot t, and the service caching strategy of BS m at time slot t is x m (t) = [x 1,m (t),...,x k,m (t)].

[0063] The user's service request tendency shows a long-tail distribution characteristic. The request probability of high-ranking services (such as k = 1) is high, while the request probability of low-ranking services gradually decreases. The user's service request follows the Zipf distribution, whose distribution parameter is α = 0.5 and the total number of services is K. Specifically, the probability of the kth service being requested (which can be used for simulation calculations) is defined by the following formula:

[0064]

[0065] Where: P k represents the probability of service k being requested; α is the offset parameter of the distribution, which determines the concentration of service popularity. When α is larger, the request probability of high-ranking services is more concentrated; K is the total number of services; is a normalization factor used to ensure that the sum of probabilities is 1.

[0066] The coordinated scheduling / beamforming (CS / CB) mode is used to achieve user wireless cooperative transmission. For a specific user u, all BSs in its cluster will receive the user's offloaded data and decode it together by exchanging channel state information (CSI) through the backbone network. It should be noted that a BS can provide services to different users at the same time, and the BS clusters of different users can intersect. The BS cluster of user u is denoted as Φ u (t), Φ u 9t) The user set served is denoted as Ω u (t). In addition to user u, the BS cluster Φ u (t) Other users served are called intra-cluster users, that is, Users other than user u and users in the cluster are called inter-cluster users.

[0067] Assume that in a time slot, each user generates an offloading task T that only requires one service u (t), and remains unchanged in each time slot. Task T u The data volume and workload of (t) are respectively and (CPU cycles, GHz). Each BS needs to cache services at the beginning of each slot to meet the service requirements of the task. In each time slot, the entire offloading process of each user will go through three steps: task offloading, task processing, and result return. Since the amount of data in the processing result is relatively small, the downlink data transmission delay is ignored.

[0068] According to the channel status setting, the signal to interference and noise ratio is calculated as:

[0069]

[0070] Among them, SINR u (t) is the signal-to-noise ratio of user u at time t, which represents the strength ratio of the received signal to the interference plus noise; p u is the transmission power of user u, which is used to adjust the transmission strength of the signal; w u (t) is the beamforming vector of user u, which enhances the target signal by optimizing the transmission directivity; g u (t) is the channel vector between the base station and user u, reflecting the gain and attenuation during signal transmission; the numerator p u |w u (t) H g u (t)| 2 Represents the useful signal power received by user u; the first part of the denominator represents the interference power caused by other users to user u, where w is the interfering user; the second part of the denominator represents the noise power suffered by user u, where is the noise power spectral density, which describes the intensity of the background noise.

[0071] The uplink transmission rate of target user u is:

[0072] r u (t) = Wlog 2 (1+SINR u (t))

[0073] Where W is the system bandwidth.

[0074] Therefore, the uplink transmission delay of user u is:

[0075]

[0076] in, It is task T u (t) The amount of data.

[0077] Assuming that the offloaded tasks are not decomposable, the task offloaded by each user will be processed by the BS with the highest probability of deploying the service in the BS cluster, i.e. in, represents the service type requested at time t. The probability of BS is The task processing cannot be performed. In this case, the task will be offloaded to the cloud for processing. Let R cloud For the data rate of the backbone network, the expected task processing delay can be obtained:

[0078]

[0079] in, They represent the computation delay at the edge and the transmission delay when offloading to the cloud. Assuming that the remote cloud has sufficient computing power, we ignore the processing delay of offloading tasks to the cloud and focus on the transmission delay of the backbone network. In order to reduce the bandwidth occupation of the backbone network, we hope to process tasks at the edge as much as possible. Therefore, assuming R cloud is a smaller value, offloading to the cloud will result in a larger transmission delay.

[0080] The final total delay is:

[0081]

[0082] S22: Constructing constraint conditions, including base station cluster constraints, average expenditure constraints of user long-term costs, base station cache resource constraints, base station computing resource constraints, base station clustering strategy constraints, and base station cache service type constraints.

[0083] The cache service on the MEC server requires BSs to pay a certain fee to the service provider. Assume that the cache cost of service k is equal to that of service s. k The total cache cost of service m at time slot t is proportional to the data size. Among them, ξ k is the cache cost factor.

[0084] Construct constraints, including base station cluster constraints, average user long-term cost constraints, base station cache resource constraints, base station computing resource constraints, base station clustering strategy constraints, and base station cache service type constraints.

[0085]

[0086] Among them, C1 is the size constraint of the base station cluster; C2 represents the average cost constraint of the user's long-term cost, that is, the average cost of the user's long-term cost does not exceed the threshold Cost th ; C3 represents the cache resource constraint of each base station; C4 represents the computing resource constraint of each base station; C5 represents the base station clustering strategy constraint, that is, whether base station m serves user u in time slot t, 1 represents service, and 0 represents no service; C6 represents the base station cache service type constraint, that is, the probability that base station m caches service type k in time slot t.

[0087] S23: Determine and construct an initial optimization problem for minimizing the total user delay based on the total user delay and constraints.

[0088]

[0089] st:

[0090]

[0091] Where C(t) = [c m (t)|m∈Φ u (t)] represents the clustering matrix at time t, X(t) = [x m (t)|m∈Φ u (t)] represents the service cache matrix at time t, T represents the total delay, represents the total user delay at time t, c u,m (t) indicates whether base station m serves user u at time t, B indicates the total number of base stations in the base station cluster, Indicates whether base station m actually serves user u at time t, Cost m (t) represents the probability of base station m caching service k at time t, Cost th represents the cache cost threshold, x k,m represents the probability that base station m caches service k at time t, s krepresents the cache resources required for service k, S m represents the cache resources of base station m, Represents a collection of users, represents the base station set, represents the time slot set, f k represents the computing resources required for service k, C m represents the computing resources of base station m, Represents a collection of service types.

[0092] S24: Based on Lyapunov optimization, the initial optimization problem of minimizing the total user delay is transformed into an instantaneous optimization problem with drift penalty as the optimization target, that is, the final optimization problem of minimizing the total user delay.

[0093] Lyapunov optimization is used to ensure the stability of the system state, so that the drift of the Lyapunov function approaches zero from the positive direction. Therefore, the constraints in the original optimization problem can be transformed into minimizing the Lyapunov drift. The steps to transform the long-term optimization problem into an instantaneous optimization problem with drift penalty as the optimization target are:

[0094] Based on Lyapunov optimization, a virtual cache cost queue C(t) is constructed, which represents the cache cost backlog of the current time slot. Assuming the initial state is C(0) = 0, the state transition of the queue can be written as: C(t+1) = [C(t) + a(t) - b(t)] +

[0095] in, b(t)=Cost th represent the task arrival rate and service rate of the queue respectively, and [·]+ represents max{·,0}.

[0096] The Lyapunov function is L(t)=1 / 2C(t) 2 , Lyapunov drift is Δ(t)=L(t+1)-L(t). According to the above definition, Lyapunov drift can be rewritten as:

[0097]

[0098] in The upper bound can be derived as:

[0099]

[0100] Based on the Lyapunov optimization framework, the long-term cache cost constraint can be transformed into the drift minimization problem of each slot. Therefore, the long-term optimization problem can be transformed into an instantaneous optimization problem with drift plus penalty as the optimization target, that is, the final optimization problem of minimizing the total user delay:

[0101]

[0102] st

[0103]

[0104] Where C(t) represents the virtual cache cost queue, C(t) represents the clustering matrix of the base station at time t, X(t) represents the service cache matrix of the user at time t, a(t) represents the task arrival rate of the queue at time t, and b(t) represents the service rate of the queue at time t; V represents a non-negative weight factor, which is selected based on the trade-off between cache cost queue drift and offloading delay.

[0105] S3: Building a soft actor-critic based deep reinforcement learning framework.

[0106] S31: Figure 3 As shown in the figure, two agents are set up in the system, where agent 1 processes the clustering strategy of the base station and agent 2 processes the cache strategy of the base station. Each agent contains a policy network and a value network. The policy network is used to generate the probability distribution of actions, and the value network is used to estimate the value function V(s) of each state. The agent obtains the state s through interaction with the environment. i (t), and output action a through the policy network i (t). After executing the action, the environment rewards r based on the action feedback i (t), and returns the new state s i (t+1). During each interaction, the state, action, reward, and new state will be stored in the experience replay pool Provide sample data for subsequent training.

[0107] S32: Each agent generates the probability distribution of actions through the strategy network according to the current environment state Sample the specific action a from it i (t). This action is aggregated and executed by the central controller, and the agent will receive a reward r i (t), state s i (t) and the next state s i (t+1) is uploaded to the central controller as part of the global experience data. This process will continue to ensure that each agent can continue to interact with the environment and accumulate enough experience data for updating the strategy and value network.

[0108] S33: The central controller receives the action, status and reward information of all agents and builds a global experience replay pool And randomly sample small batches of data (s, a, r, s′) from it. Based on the sampled data, the value network is trained using the policy-value fusion soft actor critic (SAC) algorithm. First, calculate the target Q value:

[0109]

[0110] Among them, γ is the discount factor, α is the entropy adjustment coefficient, logπ θ (a′|s′) represents the randomness constraint of the strategy. Then, the target value y is compared with the output value of the value network For mean square error optimization, the loss function of the value network is used to update the Q-network parameters based on the mean square error calculation between the target Q-value and the current Q-network output value:

[0111]

[0112] S34: Based on the value network update, optimize the policy network according to the policy gradient By maximizing the expected policy score (i.e. the expected reward including the entropy term), the loss function of the policy network is based on the policy gradient method:

[0113]

[0114] The optimizer uses the Adam optimizer and performs gradient descent based on batch data to update the parameters of the policy network, thereby improving the quality of the agent's action selection in each state.

[0115] S35: In order to improve the stability of training, the target network is introduced to perform soft updates on the policy network and value network. The parameter φ of the target network targ and θ targ Update according to the following formula:

[0116] φ targ ←τφ+(1-τ)φ targ ,θ targ ←τθ+(1-τ)θ targ ,

[0117] Among them, τ is the soft update coefficient. The soft update of the target network can alleviate the instability problem caused by the drastic fluctuation of the value network during training, making the optimization of the policy network more stable.

[0118] S36: Each agent interacts with the environment independently and updates its state based on the actions output by the policy network. The environment rewards r based on the agent's actions. i (t) and the next state s i(t+1), the agent updates its strategy through the target network based on this information. The strategy updated through the target network has high robustness, can adapt to complex and dynamic changes in environmental states, and improve the efficiency of base station clustering and cache allocation.

[0119] S37: The interaction process between the agent and the environment will be repeated continuously, and the policy network and the value network will be gradually optimized through multiple iterations, and finally stable policy parameters will be obtained. Through the collaborative optimization of multiple agents, the base station clustering and service cache joint allocation strategy can maximize the system performance, reduce latency, cost and improve resource utilization. The final optimized strategy can be widely applied to a variety of dynamic network environments, significantly improving the overall operation efficiency and service quality of the system.

[0120] S4: Solve the optimization problem of minimizing the total user delay according to the deep reinforcement learning framework based on soft actor-critic, obtain the optimal base station clustering and service caching scheme and execute it.

[0121] S41: Formulate the optimization problem of minimizing the total user delay as a Markov decision process and define the state space, action space and reward function.

[0122] Since the joint optimization problem of base station clustering and service caching is a mixed integer nonlinear programming (MINLP) problem, which contains continuous and discrete variables and involves nonlinear objective functions and constraints, traditional optimization methods are difficult to effectively solve the complexity and computational scale of the problem. To overcome this challenge, it is modeled as a Markov decision process (MDP) and the base station clustering and caching strategy is dynamically optimized by introducing a reinforcement learning framework. In MDP, the design of the state space needs to fully characterize the dynamic behavior and resource characteristics of the system.

[0123] The user's task queue status Q u (t) reflects the service demand of the current user; the channel state information H between the user and the base station u,m (t) determines the reliability and rate of data transmission; the cache resource status of the base station

[0124] and computing resource status It reflects the current resource availability of the base station and provides a decision basis for cache and task calculation. Therefore, the expression of the state space is defined as:

[0125]

[0126] Among them, S(t) represents the state space at time t, Q u (t) represents the task queue status of user u, H u,m (t) represents the channel state information between user u and base station m, represents the cache status of base station m,

[0127] represents the computing resource status of base station m, Represents a collection of users, Represents a set of base stations.

[0128] In the multi-agent based soft actor-critic (SAC) structure, the design of the action space fully reflects the collaboration and temporal correlation between agents. According to the characteristics of this optimization problem, the action space consists of two parts, corresponding to the clustering strategy and service cache strategy of the base station. The two strategies interact in the time dimension. This design ensures that the system can capture the dependencies between strategies in a dynamic network environment and achieve efficient joint optimization.

[0129] The action variable of the clustering strategy is defined as c u,m (t)∈{0,1}, indicating whether base station m serves user u. When making decisions, clustering agents need to consider not only the user’s task queue status Q u (t), channel state information H u,m (t) and the computing resource status of the base station It is also necessary to comprehensively analyze the service cache strategy X at the previous moment k,m (t-1). By referring to the cache status, the clustering strategy can better balance resource allocation and user needs and improve the rationality of task allocation.

[0130] The action variable of the service cache policy is defined as x k,m (t)∈[0,1], represents the probability of base station m caching service type k. When making decisions, the cache agent needs to consider the popularity distribution of the service type and the cache resource status of the base station. Combined with the clustering strategy c at the previous moment u,m (t-1). By utilizing historical clustering information, the service cache strategy can dynamically adjust the cache configuration so that cache resources can more efficiently match user needs and improve the efficiency of task completion.

[0131] To sum up, the action space can be expressed as:

[0132]

[0133] Among them, A(t) represents the action space at time t.

[0134] In the multi-agent framework, clustering strategy and caching strategy complement each other and jointly drive the dynamic evolution of the resource optimization process, thereby achieving global optimization of system performance in a complex network environment.

[0135] Since the dimensions and numerical ranges of indicators such as task delay, cost queue and policy matrix are different, in order to eliminate their influence on the optimization of reward function, the maximum-minimum normalization method is adopted to uniformly map all indicators to the interval [0,1]. This method is intuitive and suitable for situations where the range of eigenvalues ​​is known. It can effectively preserve the relative relationship between indicators and avoid numerical imbalance.

[0136] Normalization of task delay: Total task delay of user u Using maximum-minimum normalization, the formula is as follows:

[0137]

[0138] in, and Respectively represent the maximum and minimum values ​​of task delay.

[0139] C(t)(a(t)-b(t)) is treated as a whole, and its normalization formula is:

[0140]

[0141] Among them, [C(ab)] max and [C(ab)] min Indicates the maximum and minimum values ​​of the item.

[0142] After normalization, the final reward function is expressed as:

[0143]

[0144] Wherein, V is a non-negative weight factor used to balance the delay term and the cost term.

[0145] S42: Initialize the policy network, value network and their corresponding target networks.

[0146] In the multi-agent SAC framework, the network structure of two agents is initialized. For each agent, the main strategy network and the main value network (Q main ,π main ), and the corresponding target network (Q target ,π target ). Among them, the clustering strategy is adapted to discrete actions, and the service cache strategy is adapted to continuous actions.

[0147] The state space is designed as [S(t),A cache (t-1),A cluster(t-1)] to capture the temporal dependency and interaction between the two agents. In addition, a shared experience replay pool D is created to store training data. The CTDE framework (Centralized Training, Decentralized Execution) is used during training to achieve global information sharing to improve training efficiency while ensuring distributed decision-making capabilities in the execution phase. Training hyperparameters are set, such as the discount factor γ, the soft update coefficient τ, the initial entropy coefficient α, and the target entropy

[0148] S43: Each agent outputs actions from its policy network according to the current state. The agent interacts its actions with the environment to obtain a new environmental state and corresponding rewards. The agent stores the current state, actions and rewards in the experience replay pool.

[0149] At each time step t, based on the current state S(t) and the service cache strategy A at the previous moment cache (t-1), using the main strategy network π main,cluster Generate current clustering action A cluster (t). Discrete actions are sampled using the Gumbel-Softmax method; clustering strategy A based on the same state and the previous moment cluster (t-1), using the main strategy network π main,cache Generate cache action A cache (t); continuous actions are directly sampled from Gaussian distribution; joint actions (A cluster (t),A cache (t)) is input into the environment, and the environment returns the immediate reward r(t), the next state S(t+1) and the termination flag done; the interaction data (S(t), A(t), r(t), S(t+1)) is stored in the experience replay pool

[0150] S44: Randomly sample a batch of experiences from the experience replay pool for learning.

[0151] Randomly sample a small batch of data (S, A, r, S') from the experience replay pool.

[0152] S45: Update the Q value network and calculate the target Q value for each agent.

[0153] Calculate the target Q value:

[0154]

[0155] in: is the Q-value network, A′(t) is the action of the agent in the next state S′(t); α is the automatically adjusted entropy weight that controls the balance between exploration and utilization.φ′ (A′(t)|S′(t)) is the output of the current policy network, which represents the probability distribution of the base station service strategy or cache strategy. The parameters θ of the Q value network are updated by minimizing the Bellman error. i .

[0156] According to the target Q value and the predicted value of the main value network, the loss function is calculated:

[0157]

[0158] Optimize the main value network parameters by minimizing the loss function Enhance its ability to evaluate the value of joint actions of intelligent agents.

[0159] S46: Update the policy network parameters by maximizing the target Q value.

[0160] The policy network is optimized by maximizing the action value and action entropy, and its loss function is:

[0161]

[0162] The discrete part of the clustering strategy uses the Gumbel-Softmax method to implement gradient transfer, and the continuous part of the caching strategy is optimized by the standard SAC method; this objective function is used to improve the long-term reward of the strategy and adjust the selection of the caching strategy and the clustering strategy to improve the performance of the system.

[0163] In addition, in order to dynamically adjust the balance between exploration and exploitation, a learnable entropy coefficient α is introduced, and its update follows the following objectives:

[0164]

[0165] By optimizing the policy network parameters and entropy coefficient, the exploration efficiency and training stability of the model can be improved.

[0166] S47: Softly update the target network according to the parameters of the Q value network and the policy network.

[0167] After each time step, the parameters θ and φ of the Q-value network and the policy network are soft-updated to the target network θ′ and φ′:

[0168] Soft update of value network:

[0169] Soft update of policy network:

[0170] Among them, τ is the soft update parameter, which is usually set to a small value to smooth the update process and avoid too fast parameter changes.

[0171] S48: Repeat the training process from S35 to S37 until the agent converges and obtains the optimal clustering strategy and caching strategy.

[0172] The strategies of the two agents collaborate at the state space level, and the state design explicitly includes the other party's actions at the previous moment, thereby modeling the temporal dependency between the agents. Joint optimization is achieved through a shared replay pool and reward function, using the CTDE framework to share global state information to improve training efficiency, while the execution phase relies entirely on their respective local states to ensure the independence of distributed decision-making.

[0173] The reward function comprehensively considers the joint optimization objectives of the clustering strategy and the caching strategy:

[0174]

[0175] Where C(t) represents the system's virtual cache cost queue, Represents the total delay of the task; by normalizing various indicators, the impact of numerical scale differences on the optimization process can be avoided.

[0176] The training process terminates when the agent converges or reaches the set maximum number of iterations E. In each iteration, the trained clustering strategy and caching strategy will be used as the initial input for the next round of training, so that the strategy is continuously improved in the loop optimization, and finally the optimal base station clustering and service caching solution is obtained. The execution of this solution maximizes the resource utilization of the base station while minimizing task delays and system costs.

[0177] In this embodiment, s and s(t) both represent an action in a state set, and s(.) emphasizes that it is a state at a certain time step. Similarly, a and a(t) both represent an action in an action set, and a(.) emphasizes that it is an action selected at a certain time step. Q(s,a) and Q(s,a;θ) both represent the Q value function obtained by selecting a specific action in a specific state, where Q(s,a;θ) emphasizes that the network parameter when the current action is acquired is θ.

[0178] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A base station clustering and service cache joint allocation method, characterized in that: The following steps are involved: S1: Build a dense mobile edge network, which includes multiple base stations and multiple users, where the base stations follow a Poisson distribution; S2: Based on dense mobile edge networks, we construct an optimization problem to minimize the total user delay. S3: Build a soft actor-critic based deep reinforcement learning framework; S4: Solve the optimization problem of minimizing the total user delay according to the deep reinforcement learning framework based on soft actor-critic, obtain the optimal base station clustering and service caching scheme and execute it.

2. A base station clustering and service cache joint allocation method according to claim 1, characterized in that: The process of constructing the optimization problem of minimizing the total user delay includes: S21: Calculate the task processing delay and uplink transmission delay of the user, and take the sum of the task processing delay and uplink transmission delay of the user as the total delay of the user; S22: constructing constraint conditions, including base station cluster constraint, average cost constraint of user long-term cost, base station cache resource constraint, base station computing resource constraint, base station clustering strategy constraint and base station cache service type constraint; S23: Determine and construct an initial optimization problem of minimizing the total user delay according to the total user delay and the constraint conditions; S24: Based on Lyapunov optimization, the initial optimization problem of minimizing the total user delay is transformed into an instantaneous optimization problem with drift penalty as the optimization target, that is, the final optimization problem of minimizing the total user delay.

3. A base station clustering and service cache joint allocation method according to claim 2, characterized in that: The initial minimization of the total user delay optimization problem is expressed as: Among them, C(t) represents the clustering matrix of the base station at time t, X(t) represents the service cache matrix of the user at time t, T represents the total delay, represents the total task delay of user u at time t, c u,m (t) indicates whether base station m serves user u at time t, B indicates the total number of base stations in the base station cluster, Indicates whether base station m actually serves user u at time t, Cost m (t) represents the total cost of base station m at time t, Cost th represents the cache cost threshold, x k,m represents the probability that base station m caches service k at time t, s k represents the cache resources required for service k, S m represents the cache resources of base station m, Represents a collection of users, represents the base station set, represents the time slot set, f k represents the computing resources required for service k, C m represents the computing resources of base station m, Represents a collection of service types.

4. A method for joint allocation of base station clustering and service cache according to claim 2, characterized in that: The optimization problem of minimizing the total user delay is expressed as: Where C(t) represents the clustering matrix of the base station at time t, X(t) represents the service cache matrix of the user at time t, a(t) represents the task arrival rate of the queue at time t, b(t) represents the service rate of the queue at time t, and V represents the weight factor. represents the total task delay of user u at time t, c u,m (t) indicates whether base station m serves user u at time t, B indicates the total number of base stations in the base station cluster, and x k,m represents the probability that base station m caches service type k at time t, s k represents the cache resources required for service k, S m represents the cache resources of base station m, Represents a collection of users, represents the base station set, represents the time slot set, f k represents the computing resources required for service k, C m represents the computing resources of base station m, Represents a collection of service types.

5. A base station clustering and service cache joint allocation method according to claim 1, characterized in that: The process of solving the optimization problem of minimizing the total user delay based on the deep reinforcement learning framework of soft actor-critic includes: S41: Formulate the optimization problem of minimizing the total user delay as a Markov decision process, and define the state space, action space and reward function; S42: Initialize the strategy network, value network and their corresponding target network; S43: Each agent outputs an action from its policy network according to the current state. The agent interacts its action with the environment to obtain a new environment state and corresponding reward. The agent stores the current state, action and reward in the experience replay pool. S44: Randomly sample a batch of experiences from the experience replay pool for learning; S45: Update the Q value network and calculate the target Q value for each agent; S46: Update the policy network parameters by maximizing the target Q value; S47: softly updating the target network according to the parameters of the Q value network and the policy network; S48: Repeat the training process from S35 to S37 until the agent converges or reaches the maximum number of iterations, and obtain the optimal clustering strategy and caching strategy.

6. A base station clustering and service cache joint allocation method according to claim 5, characterized in that: The state space is represented as: Among them, S(t) represents the state space at time t, Q u (t) represents the task queue status of user u, H u,m (t) represents the channel state information between user u and base station m, represents the cache status of base station m, represents the computing resource status of base station m, Represents a collection of users, Represents a set of base stations.

7. A base station clustering and service cache joint allocation method according to claim 5, characterized in that: The action space is represented as: Among them, A(t) represents the action space at time t, C(t) represents the clustering matrix at time t, and X(t) represents the service cache matrix at time t. Represents a collection of users, represents the base station set, Represents a collection of service types.

8. A base station clustering and service cache joint allocation method according to claim 5, characterized in that: The reward function is expressed as: Among them, r(t) represents the reward at time t, C(t) represents the virtual cache cost queue, a(t) represents the task arrival rate of the queue at time t, b(t) represents the task service rate of the queue at time t, and V represents the weight factor. represents the total delay of user u at time t, Represents a collection of users.

Citation Information

Cited By

  • Industrial service cache optimization method and system based on intelligent prediction

    CN122437887A