A method and system for minimizing latency through joint caching and transmission in dense small cells

CN117062149BActive Publication Date: 2026-09-18XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311039609.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2026-09-18
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

这些算法可以接近最佳性能,但需要全局信道状态信息,并需要多次迭代计算最优解,因此在实际执行中的高开销和高计算复杂度是难以接受的

Benefits of technology

[0076] In dense small-cell scenarios, unlike traditional cellular base stations, different base stations have different probabilities of being requested. Therefore, when modeling delivery latency, the different request probabilities between base stations should be considered. For transmission decisions, if caching is not considered, the goal is to maximize the transmission rate, which requires balancing interference and useful signals to obtain the scheduling decision with the highest transmission rate. When caching is considered, scheduling more base stations will put more pressure on the cache, resulting in greater delivery latency. Therefore, a balance between transmission rate and delivery latency is needed. This invention considers the request probabilities of different base stations in caching decisions to minimize delivery latency. In terms of transmission decisions, a multi-agent reinforcement learning scheme is proposed. The transmission policy optimization problem is to minimize the queue length given a known caching policy. Each base station trains its own reinforcement learning network, with local information as input and the number of transmitted files as reward. The output can obtain the optimal scheduling user index and power index. This invention simplifies the computation process and achieves low information overhead and high performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117062149B_ABST
    Figure CN117062149B_ABST
Patent Text Reader

Abstract

A method and system for minimizing latency through joint caching and transmission in dense small cells is disclosed. The method includes: calculating a caching decision based on the request probability of each base station, with the goal of minimizing delivery latency; given the caching decision, each base station trains its own reinforcement learning network to calculate a transmission strategy through multi-agent reinforcement learning, with the goal of minimizing queue length; the input of the multi-agent reinforcement learning is local information, using the number of transmitted files as a reward, and outputting the optimal scheduling user index and power index as the transmission strategy. For the caching strategy calculation, this invention proposes an iterative convex optimization solution to obtain optimal performance, and provides a low-complexity suboptimal closed-form solution considering network training. Regarding the transmission decision in each time slot, the local decision of each base station can be completed independently, requiring only a small amount of scalar information exchange between base stations. This simplifies the computation process and achieves low information overhead and high performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication network technology, specifically relating to a method and system for minimizing latency through joint buffering and transmission in dense small cells. Background Technology

[0002] With the increasing prevalence of smart devices and the exponential growth in wireless data demand, current wireless communication networks are struggling to keep up with the rapidly increasing service demands. To improve network coverage and increase network capacity, dense small cell networks have been introduced as a new and efficient architecture technology. By deploying a large number of short-range, low-power, and low-cost small cell base stations, dense small cell networks can not only supplement coverage in areas not covered by macro base stations, but also provide redundant coverage in high-user-density hotspots, ensuring data transmission rates in these areas and thus improving the quality of service for users in those areas.

[0003] Clearly, densely deployed small cell base stations cause severe inter-cell interference, and spectrum sharing significantly reduces system capacity. Coordinated Multipoint Transmission (CoMP) is a promising inter-cell cooperation technique that can significantly mitigate inter-cell interference and improve system capacity. CoMP has various cooperative transmission schemes, such as cooperative beamforming (CB), cooperative scheduling, and joint transmission (JT). Coherent joint transmission is the most advanced CoMP transmission scheme, enabling complete sharing of data and control information among cooperating base stations. However, strict inter-base station synchronization is a major limiting factor in its practical application. Incoherent JT is a more attractive technique, where the user equipment (UE) signal is transmitted jointly by multiple cooperating base stations without prior phase mismatch correction or strict synchronization between base stations. At the UE end, power gain is achieved through incoherent combination of received signals. Due to the large number of small base stations and wireless devices in dense small cell networks, the centralized computation and information overhead of implementing incoherent JT becomes cumbersome. In practical deployments, reducing the information overhead and computational complexity of incoherent JT in dense small cells is crucial.

[0004] The dense deployment of small cells can effectively meet the ever-increasing demand for mobile data traffic. However, due to the installation process and cost considerations, the backhaul link capacity of small cells is limited. When the backhaul link capacity is limited, it may lead to a decrease in user data transmission rate, thereby prolonging the latency for users to obtain requested content and thus affecting the quality of service. When caching is deployed on small cell base stations, if the user's requested content is already stored in the cache, that is, when the user's request hits the cache, the small cell does not need to download the user's requested content from the core network through the backhaul link and can directly send the cached user's requested content to the user. Compared with the small cell needing to obtain the user's requested content from the core network through the backhaul link when the user's request does not hit the cache, the user's request hitting the cache not only alleviates the pressure on the backhaul link but also reduces the latency for users to obtain content. Therefore, deploying caching on small cell base stations is particularly important, especially when active caching is performed on small cell base stations. That is, popular files are cached in advance during off-peak hours to directly serve users. This not only reduces the backhaul link pressure during peak hours but also makes full use of bandwidth resources during off-peak hours.

[0005] In latency-sensitive scenarios, such as vehicle-to-vehicle communication and virtual reality, ensuring timely data delivery takes precedence over other performance improvements. Unlike traditional cellular base stations, the probability of different base stations being requested varies in dense small cell scenarios, and the impact of base station request probability on latency has not yet been studied. Furthermore, minimizing latency using incoherent joint transmission in dense small cells remains an unresolved issue. Current research largely focuses on optimization and rate, and some optimization techniques have been proposed, such as fractional programming algorithms, weighted least mean square error algorithms, and branch and bound algorithms. These algorithms can approach optimal performance, but require global channel state information and multiple iterations to calculate the optimal solution, making their high overhead and computational complexity unacceptable in practical implementation. Summary of the Invention

[0006] The purpose of this invention is to address the problems in the prior art by providing a method and system for minimizing latency through joint caching and transmission in dense small cells, simplifying the calculation process, and achieving low information overhead and high performance.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A method for minimizing latency through joint caching and transmission in dense small cells includes the following steps:

[0009] Based on the request probability of each base station, a caching decision is calculated with the goal of minimizing delivery latency;

[0010] Given the cache decision, each base station trains its own reinforcement learning network with the goal of minimizing the queue length, and calculates the transmission strategy through multi-agent reinforcement learning; wherein the input of the multi-agent reinforcement learning is local information, the number of transmitted files is used as the reward, and the output is the optimal scheduling user index and power index as the transmission strategy.

[0011] As a preferred embodiment, the dense small cell has K single-antenna small base stations and N single-antenna users, where the base station set and user set are respectively represented as... and The base station uses non-coherent joint transmission for interference coordination, and the signal-to-interference-plus-noise ratio (SIR) received by user i is:

[0012]

[0013] In the formula, g k,i P is the channel gain from base station k to user i. k,i σ is the transmit power allocated by base station k to user i. 2 The noise power is: The rate of user i is:

[0014] R i =Wlog(1+SINR) i )

[0015] In the formula, W is the channel bandwidth.

[0016] As a preferred approach, user-requested files follow a Zipf distribution, and the probability of requesting the m-th popular file is:

[0017]

[0018] In the formula, N represents the total number of contents, and γ is a shape parameter representing the skewness of the document's popularity, when m << N and γ < 1:

[0019]

[0020] As a preferred embodiment, the base station cache is divided into two parts: one part uses the MPC scheme to store the most popular files of the same type, and the other part uses the LCD scheme to store different files to increase file diversity; the base station can cache M files, and the scaling factor of base station k is ρ. k , is the ratio of the number of most popular files cached by base station k to the total number of files that the base station can cache. In other words, it represents the number of files with the same rank stored by each base station. The most popular files are stored in different rankings in the remaining space. The file, N k Store the starting sequence number of different files for base station k. The probability of a local cache hit. for:

[0021]

[0022] The probability of a local cache miss for base station k is The overall system hit probability is:

[0023]

[0024] The probability of missing is The content delivery latency for base station k is:

[0025]

[0026] In the formula, S f This is the file size, in bits, r b and r c These are the transmission rates between base stations and the transmission rate from the base station to the core network, respectively; the long-term average content delivery latency of the entire system is:

[0027]

[0028] In the formula, Let be the request probability of base station k.

[0029] As a preferred option, with the goal of minimizing delivery latency, the expression for calculating the caching decision is as follows:

[0030]

[0031]

[0032] The pseudocode for iterative convex optimization to minimize delivery latency includes the following steps:

[0033] Sort all base station request probabilities in descending order;

[0034] Optimize the base station buffer coefficient sequentially using CVX.

[0035] After convergence, the buffer coefficient of each base station is obtained;

[0036] With a consistent buffer coefficient across all base stations, an approximate closed-form solution can be obtained using the following formula:

[0037]

[0038] In the formula,

[0039]

[0040]

[0041] As a preferred embodiment, the base station makes a transmission decision in each time slot. After making the transmission decision, it delivers content according to the requested content of the user it serves: if the requested content of the user served by the base station is local to the base station, the base station does not deliver the content; if the requested content of the user served by the base station is not local to the base station but is cached in other base stations, the corresponding base station obtains the requested content from other base stations through the inter-base station link; if the requested content of the user served by the base station is neither local to the base station nor in other base stations, the corresponding base station obtains the requested content from the central server through the link between the base station and the central server; after completing the content delivery, the base station transmits signals according to the transmission decision.

[0042] As a preferred approach, when a user is served by multiple base stations, the delivery delay for that user is the maximum delivery delay among all the base stations serving the user, calculated as follows:

[0043]

[0044] In the formula, Let be the set of all serving base stations for user i; therefore, the number of files transmitted by user i at time (t+1) is:

[0045]

[0046] In the formula, τ is the time slot length, and the queue state of user i at time (t+1) is:

[0047] q i (t+1)=q i (t)-T i (t)+A i (t)

[0048] In the formula, A i (t) represents the number of services received by user i at time (t+1), which follows a Poisson distribution with strength λ.

[0049] The latency minimization problem is transformed into a queue length minimization problem, and the calculation expression is:

[0050]

[0051]

[0052]

[0053]

[0054] As a preferred approach, with the goal of minimizing the queue length, the expression for the transmission strategy computed through multi-agent reinforcement learning is as follows:

[0055]

[0056]

[0057]

[0058] The process of multi-agent reinforcement learning includes:

[0059] Action settings: The actions of a small cell base station include indexes for user scheduling and transmission power;

[0060] definition Let the set of available transmit power at base station j in the small cell be represented as:

[0061]

[0062] In the formula, N p It is the number of available transmit power levels, P max This is the maximum available transmit power; small cell base stations are based on... and action a j Transmit power index in (t) Select the transmit power;

[0063] In the t-th time slot, the available actions of base station j are:

[0064]

[0065] In the formula, I j (t) represents the user index scheduled by base station j in time slot t;

[0066] State settings: The state includes channel state information and queue length information for all users; the useful signal power, interference power, and queue length of all users at the previous time step are used as the state input to the network. The state of base station j is:

[0067]

[0068] In the formula, This represents the useful signal power of the i-th user in the t-th time slot. This represents the interference signal power of the i-th user in the t-th time slot;

[0069] Reward settings: Similar to the reward settings in a single-user network, the reward for the k-th DQN network of base station j is:

[0070]

[0071] A latency-minimizing system for joint caching and transmission in dense small cells includes:

[0072] The caching decision calculation module is used to calculate caching decisions based on the request probability of each base station, with the goal of minimizing delivery latency.

[0073] The transmission strategy calculation module is used to train each base station's own reinforcement learning network with the goal of minimizing the queue length, given a known cache decision. The transmission strategy is calculated through multi-agent reinforcement learning. The input of the multi-agent reinforcement learning is local information, the number of transmitted files is used as a reward, and the output is the optimal scheduling user index and power index as the transmission strategy.

[0074] A computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for minimizing latency of joint caching and transmission in dense small cells.

[0075] Compared with the prior art, the present invention has at least the following beneficial effects:

[0076] In dense small-cell scenarios, unlike traditional cellular base stations, different base stations have different probabilities of being requested. Therefore, when modeling delivery latency, the different request probabilities between base stations should be considered. For transmission decisions, if caching is not considered, the goal is to maximize the transmission rate, which requires balancing interference and useful signals to obtain the scheduling decision with the highest transmission rate. When caching is considered, scheduling more base stations will put more pressure on the cache, resulting in greater delivery latency. Therefore, a balance between transmission rate and delivery latency is needed. This invention considers the request probabilities of different base stations in caching decisions to minimize delivery latency. In terms of transmission decisions, a multi-agent reinforcement learning scheme is proposed. The transmission policy optimization problem is to minimize the queue length given a known caching policy. Each base station trains its own reinforcement learning network, with local information as input and the number of transmitted files as reward. The output can obtain the optimal scheduling user index and power index. This invention simplifies the computation process and achieves low information overhead and high performance.

[0077] Furthermore, for caching strategy calculation, this invention proposes an iterative convex optimization solution to obtain optimal performance, and provides a low-complexity suboptimal closed-form solution considering network training. Regarding transmission decisions in each time slot, the base station's local decision-making can be completed independently; only a small amount of scalar information exchange is needed between base stations, without requiring feedback channel state information between the base station and the user. The overall system's interaction feedback is independent of antenna configuration. Attached Figure Description

[0078] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0079] Figure 1 This is a schematic diagram of the dense small cell system model structure according to an embodiment of the present invention;

[0080] Figure 2 This is a schematic diagram of the base station caching method according to an embodiment of the present invention;

[0081] Figure 3 This is a schematic diagram of the base station's operation in each time slot according to an embodiment of the present invention;

[0082] Figure 4 The delivery latency performance diagrams are for the embodiment of the present invention and the comparative scheme under the same caching coefficient.

[0083] Figure 5 This is a graph showing how delivery delay changes with the number of iterations.

[0084] Figure 6 A graph showing the change in delivery delay as a function of shape parameters;

[0085] Figure 7 A graph showing how the probability of successful transmission changes over time. Detailed Implementation

[0086] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, those skilled in the art can obtain other embodiments without creative effort.

[0087] like Figure 1 As shown, consider a dense small cell system with K single-antenna small base stations and N single-antenna users. The set of base stations and the set of users are represented as follows: and

[0088] The base station uses non-coherent joint transmission for interference coordination, and the signal-to-interference-plus-noise ratio (SIR) received by user i is:

[0089]

[0090] Among them, g k,i P is the channel gain from base station k to user i. k,i σ is the transmit power allocated by base station k to user i. 2Let be the noise power. The rate of user i is:

[0091] R i =Wlog(1+SINR) i (2)

[0092] Where W is the channel bandwidth.

[0093] User-requested files follow a Zipf distribution, and the probability of requesting the m-th popular file is:

[0094]

[0095] Where N represents the total number of contents, and γ is a shape parameter representing the skewness of the file's popularity, when m << N and γ < 1:

[0096]

[0097] The base station cache is divided into two parts: one part uses the MPC (most popular content) scheme to store the same most popular files, and the other part uses the LCD (largest content diversity) scheme to store different files to increase file diversity, such as... Figure 2 As shown. The base station can cache M files, and the scaling factor for base station k is ρ. k , is the ratio of the number of most popular files cached by base station k to the total number of files that the base station can cache. In other words, it represents the number of files with the same rank stored by each base station. The most popular files are stored in different rankings in the remaining space. The file, N k Store the starting sequence number of different files for base station k. The probability of a local cache hit. for:

[0098]

[0099] The probability of a local cache miss for base station k is The overall system hit probability is:

[0100]

[0101] The probability of missing is The content delivery latency for base station k is:

[0102]

[0103] Among them, S f This is the file size, in bits, r b and r cThese are the transmission rates between base stations and from the base station to the core network, respectively. The long-term average content delivery latency of the entire system is:

[0104]

[0105] in, Let be the request probability of base station k.

[0106] The base station operates in each time slot as follows Figure 3 As shown. First, the base station needs to make a transmission decision in each time slot, that is, how much power the base station should serve which users. After making the transmission decision, it delivers content according to the requested content of the users it serves. If the requested content of the user served by the base station is local to the base station, the base station does not need to deliver the content; if the requested content of the user served by the base station is not local to the base station, but cached in other base stations, the base station needs to obtain the requested content from other base stations through the inter-base station link; if the requested content of the user served by the base station is neither local to the base station nor in other base stations, the base station needs to obtain the requested content from the central server through the link between the base station and the central server. After completing the content delivery, the base station will transmit signals according to the transmission decision. The length of each time slot is fixed. The longer the content delivery time, the shorter the available transmission time. When a user is served by multiple base stations, the delivery delay of the user is the maximum value of the delivery delays of all the base stations serving the user, that is:

[0107]

[0108] in, Let be the set of all serving base stations for user i. Therefore, the number of files transmitted by user i at time (t+1) is expressed as:

[0109]

[0110] Where τ is the time slot length, the queue state of user i at time (t+1) is represented as:

[0111] q i (t+1)=q i (t)-T i (t)+A i (t) (11)

[0112] Among them, A i (t) represents the number of services received by user i at time (t+1), which follows a Poisson distribution with strength λ.

[0113] The goal is to minimize packet-level latency in dense small cells by utilizing buffering and joint transmission. According to Little's law, minimizing packet-level latency is equivalent to minimizing queue length. Therefore, the latency minimization problem can be transformed into a queue length minimization problem, expressed as:

[0114]

[0115] To minimize the long-term average queue length, the queue update formula shows that this metric is related to delivery latency and transmission rate. This invention aims to balance maximizing transmission rate and minimizing delivery latency. Caching inherently involves a trade-off: caching popular files at each base station to increase local hit rate versus caching non-popular files to increase overall system file diversity. In dense small-cell scenarios, unlike traditional cellular base stations, different base stations have different probabilities of being requested. Therefore, when modeling delivery latency, the different request probabilities between base stations should be considered. For transmission decisions, without considering caching, maximizing transmission rate is the goal, which requires balancing interference and useful signals to achieve the scheduling decision with the highest transmission rate. Considering caching, scheduling more base stations puts greater pressure on the cache, resulting in higher delivery latency; therefore, a balance between transmission rate and delivery latency is necessary.

[0116] This invention decouples the problem into two parts: a caching strategy optimization problem and a transmission strategy optimization problem. The caching strategy optimization problem aims to minimize the delivery latency given a known request probability for each base station, i.e.:

[0117]

[0118] The transmission strategy optimization problem is to minimize the queue length given a known caching strategy, i.e.:

[0119]

[0120] The existing technology has not yet studied the delivery delay minimization caching algorithm that considers the base station request probability under dense small cells. The present invention proposes the following method to solve the optimization problem of equation (13).

[0121] In other ρ k′ When k′≠k is considered a constant, the system delivery delay d div Regarding ρ k It is a convex function.

[0122] The proof is as follows:

[0123]

[0124] d div For ρk Taking the first derivative (if k=1, then removing the two terms of the summation that do not include the first base station), we get:

[0125]

[0126] in,

[0127]

[0128]

[0129]

[0130] d div For ρ k Taking the second derivative (if k=1, then removing the two terms of the summation that do not include the first base station), we get:

[0131]

[0132] in,

[0133]

[0134]

[0135]

[0136] because so That is, other ρ k′ When k′≠k is considered a constant, the system delivery delay d div Regarding ρ k It is a convex function.

[0137] When other base stations are considered constant, the system delivery delay can be optimized using CVX. Based on this, this invention proposes an iterative convex optimization solution, namely, iterative convex optimization pseudocode that minimizes the delivery delay.

[0138] 1. Sort all base station request probabilities in descending order;

[0139] 2. Optimize the base station buffer coefficient sequentially using CVX according to the order;

[0140] 3. After convergence, the buffer coefficient of each base station is obtained.

[0141] To reduce the complexity of the caching scheme, this embodiment of the invention obtains an approximately closed-form solution with consistent caching coefficients for each base station, namely:

[0142]

[0143] in,

[0144]

[0145]

[0146] The derivation process is as follows:

[0147] when At that time, the first derivative of the system delivery delay can be simplified to:

[0148]

[0149] in,

[0150]

[0151]

[0152]

[0153] Therefore, we can conclude that:

[0154]

[0155] make,

[0156]

[0157]

[0158] but,

[0159]

[0160] make have to:

[0161]

[0162] For the transmission strategy, this embodiment of the invention solves the optimization problem of equation (14) through multi-agent reinforcement learning. Each base station sets up a pair of DQN networks and trains and executes them through reinforcement learning.

[0163] Action Setup: For small base stations, user and transmit power need to be obtained based on actions. Therefore, the actions of a small cell base station include indexes for scheduling users and transmit power. However, transmit power is a continuous value, and reinforcement learning struggles to handle continuous action spaces. Therefore, a simple and effective discrete method is proposed to handle continuous value problems. Definition Let the set of available transmit power at base station j in the small cell be represented as:

[0164]

[0165] Where, N p It is the number of available transmit power levels, P max This is the maximum available transmit power. Small cell base stations are based on... and action a j Transmit power index in (t) Choose the transmit power. Therefore, in the t-th time slot, the available actions of small base station j are:

[0166]

[0167] Among them, I j (t) is the user index of base station j in time slot t.

[0168] State Setup: The state should include channel state information and queue length information for all users. Directly inputting all user channel state information and queue length information results in an excessively large input vector, making training difficult to converge. Here, the useful signal power, interference power, and queue length of all users at the previous time step are used as the state input to the network. The state of base station j is...

[0169]

[0170] in, This represents the useful signal power of the i-th user in the t-th time slot. This represents the interference signal power of the i-th user in the t-th time slot.

[0171] Reward settings: Similar to the reward settings in a single-user network, the reward for the k-th DQN network of base station j is:

[0172]

[0173] To verify the performance of the latency minimization method for joint caching and transmission in dense small cells according to the embodiments of the present invention, the following simulations were performed:

[0174] To verify the effectiveness of the method of the present invention in terms of caching strategy, the embodiments of the present invention select two classic caching algorithms, MPC and LCD, as comparison algorithms. The embodiments of the present invention were simulated in a scenario where K=16 and N=8.

[0175] Scenario 1: To verify the effectiveness of the proposed closed-form solution in caching, Figure 4 The performance of the proposed scheme in terms of delivery latency was compared with other comparable schemes under the same caching factor. From... Figure 4It can be observed that the performance of the proposed closed-form solution is comparable to that of the traversal solution, indicating that the proposed closed-form solution can achieve optimal performance. At the same time, the proposed solution significantly outperforms the two heuristic solutions, namely the MPC and LCD solutions.

[0176] Case 2: To observe the convergence of the proposed iterative convex optimization scheme, Figure 5 The curve showing the delivery delay as a function of the number of iterations was plotted. Figure 5 As can be observed, after K iterations, the delivery latency of the iterative convex optimization scheme tends to stabilize, and subsequent iterations offer little performance improvement. This indicates that the iterative convex optimization scheme has converged after K iterations, and further optimization brings very limited performance gains.

[0177] Scenario 3: To observe the impact of shape parameters on delivery delay, in Figure 6 The curves showing delivery delay as a function of different shape parameters were plotted. Figure 6 As can be seen, the suboptimal closed-form solution performs close to the iterative convex optimization solution under different shape parameters, and is significantly better than the LCD and MPC solutions.

[0178] Scenario 4: To observe the impact of the proposed caching and transmission joint optimization scheme on latency, the success transmission probabilities of different schemes were compared. Here, files successfully transmitted are defined as those transmitted within the specified time slot, while those transmitted after the specified time slot are defined as transmission failures. The success transmission probability is the ratio between the number of successfully transmitted files and the number of files requested by the user. A higher success transmission probability means lower latency for the user. Specific schemes compared include:

[0179] DRL CF scheme: The proposed closed debuffering scheme and reinforcement learning transport scheme.

[0180] DRL MPC schemes: MPC caching scheme and reinforcement learning transport scheme.

[0181] DRL LCD Solution: LCD caching solution and reinforcement learning transmission solution.

[0182] Greedy BS CF scheme: a closed debuffering scheme and a base station greedy scheduling scheme, in which each base station selects the user with the best channel quality to serve.

[0183] Greedy UE CF scheme: closed debuffering scheme and user greedy scheduling scheme, in which each user selects the base station with the best channel quality for service.

[0184] from Figure 7It can be observed that the DRL CF scheme performs best, outperforming both the DRL MPC and DRL LCD schemes. This further proves the effectiveness of the closed-loop debugging scheme. Simultaneously, the DRL CF scheme also outperforms the other two classic scheduling schemes, namely those based on Greedy UE CF and Greedy BS CF. This demonstrates the effectiveness of the joint optimization scheme for caching and transmission based on multi-agent reinforcement learning.

[0185] Therefore, as can be seen from the above, the performance of the caching and transmission joint optimization method proposed in this invention is far superior to existing solutions.

[0186] This invention also proposes a latency minimization system for joint caching and transmission in dense small cells, comprising:

[0187] The caching decision calculation module is used to calculate caching decisions based on the request probability of each base station, with the goal of minimizing delivery latency.

[0188] The transmission strategy calculation module is used to train each base station's own reinforcement learning network with the goal of minimizing the queue length, given a known cache decision. The transmission strategy is calculated through multi-agent reinforcement learning. The input of the multi-agent reinforcement learning is local information, the number of transmitted files is used as a reward, and the output is the optimal scheduling user index and power index as the transmission strategy.

[0189] This invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for minimizing latency of joint caching and transmission in dense small cells.

[0190] For example, the instructions stored in the memory can be divided into one or more modules / units. These modules / units are stored in a computer-readable storage medium and executed by the processor to complete the latency minimization method for joint caching and transmission in dense small cells according to the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the server.

[0191] The electronic device may be a smartphone, laptop, PDA, or cloud server, among other computing devices. It may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the electronic device may also include more or fewer components, or combinations of certain components, or different components; for example, it may also include input / output devices, network access devices, buses, etc.

[0192] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0193] The memory can be an internal storage unit of the server, such as a hard drive or RAM. Alternatively, it can be an external storage device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units. The memory is used to store computer-readable instructions and other programs and data required by the server. It can also be used to temporarily store data that has been output or will be output.

[0194] It should be noted that the information interaction and execution process between the above-mentioned module units are based on the same concept as the method embodiment. For details on their specific functions and technical effects, please refer to the method embodiment section. They will not be repeated here.

[0195] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0196] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0197] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0198] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for minimizing latency through joint caching and transmission in dense small cells, characterized in that, Includes the following steps: Based on the request probability of each base station, a caching decision is calculated with the goal of minimizing delivery latency; Given the cache decision, each base station trains its own reinforcement learning network with the goal of minimizing the queue length, and calculates the transmission strategy through multi-agent reinforcement learning; wherein, the input of the multi-agent reinforcement learning is local information, the number of transmitted files is used as the reward, and the output of the optimal scheduling user index and power index is the transmission strategy. When a user is served by multiple base stations, the delivery delay for that user is the maximum delivery delay among all the base stations serving that user. The calculation expression is as follows: In the formula, For users The set of all serving base stations; therefore, the user exist The number of files transferred at any given time is: In the formula, For users rate, For time slot length, user exist The queue state at time t is: In the formula, User exist The number of transactions arriving at a given time, with a compliance strength of [value missing]. The Poisson distribution; The latency minimization problem is transformed into a queue length minimization problem, and the calculation expression is: In the formula, Represents a set of base stations. Represents a set of users; With the goal of minimizing the queue length, the expression for the transmission policy computed through multi-agent reinforcement learning is as follows: The process of multi-agent reinforcement learning includes: Action settings: The actions of a small cell base station include indexes for user scheduling and transmission power; definition For small cell base stations The set of available transmission powers is represented as: In the formula, It is the number of available transmission power levels. This is the maximum available transmit power; small cell base stations are based on... and actions Transmit power index in Select the transmit power; In the In each time slot, the base station The available actions are: In the formula, For base stations exist User index for time slot scheduling; State settings: The state includes channel state information and queue length information for all users; the useful signal power, interference power, and queue length of all users at the previous moment are used as state inputs to the network and base station. The status is: In the formula, Indicates the first The first time slot Useful signal power for each user Indicates the first The first time slot Interference signal power for each user; Reward settings: Same as reward settings in a single-user network, base station The The reward for each DQN network is: 。 2. The latency minimization method for joint caching and transmission in dense small cells according to claim 1, characterized in that, The densely populated small communities have Single-antenna small base station Single-antenna users; the base station uses incoherent joint transmission for interference coordination. The received signal-to-interference-plus-noise ratio is: In the formula, It is a base station To users Channel gain, It is a base station Assigned to user The transmission power, Noise power; User The rate is: In the formula, It is the channel bandwidth.

3. The latency minimization method for joint caching and transmission in dense small cells according to claim 2, characterized in that, The user-requested file follows a Zipf distribution, the nth The probability of requesting a popular file is: In the formula, Indicates the total number of contents. It is a shape parameter representing the skewness of the file. and hour: 。 4. The latency minimization method for joint caching and transmission in dense small cells according to claim 3, characterized in that, The base station cache is divided into two parts: one part uses the MPC scheme to store the most popular files, and the other part uses the LCD scheme to store different files to increase file diversity; the base station can cache a maximum of [number missing]. Base station The scaling factor is It is a base station The ratio of the number of most popular cached files to the number of files that a base station can cache; that is, the ratio of the number of files with the same rank stored by each base station. The most popular files are stored in different rankings in the remaining space. The file, For base stations Store the starting sequence number of different files. The probability of a local cache hit is... for: base station The probability of a local cache miss is The overall system hit probability is: The probability of missing is Base station Content delivery latency is: In the formula, This is the file size, in bits. and These are the transmission rates between base stations and the transmission rate from the base station to the core network, respectively; the long-term average content delivery latency of the entire system is: In the formula, For base stations The probability of a request.

5. The latency minimization method for joint caching and transmission in dense small cells according to claim 4, characterized in that, With the goal of minimizing delivery latency, the expression for calculating the caching decision is as follows: The pseudocode for iterative convex optimization to minimize delivery latency includes the following steps: Sort all base station request probabilities in descending order; Optimize the base station buffer coefficient sequentially using CVX. After convergence, the buffer coefficient of each base station is obtained; With a consistent buffer coefficient across all base stations, an approximate closed-form solution can be obtained using the following formula: In the formula, 。 6. The latency minimization method for joint caching and transmission in dense small cells according to claim 4, characterized in that, The base station makes a transmission decision in each time slot. After making the transmission decision, it delivers content according to the content requested by the user it serves: if the content requested by the user served by the base station is local to the base station, the base station does not deliver the content. If the requested content of the user served by the base station is not local to the base station, but is cached in other base stations, the corresponding base station obtains the requested content from other base stations through the inter-base station link; If the requested content of the user served by the base station is neither on the base station itself nor on other base stations, the corresponding base station obtains the requested content from the central server through the link between the base station and the central server. After content delivery is completed, the base station transmits signals based on transmission decisions.

7. A latency minimization system for joint caching and transmission in dense small cells, characterized in that, include: The caching decision calculation module is used to calculate caching decisions based on the request probability of each base station, with the goal of minimizing delivery latency. The transmission strategy calculation module is used to train each base station's own reinforcement learning network with the goal of minimizing the queue length, under the known cache decision. The transmission strategy is calculated through multi-agent reinforcement learning. The input of the multi-agent reinforcement learning is local information, the number of transmitted files is used as the reward, and the output of the optimal scheduling user index and power index is the transmission strategy. When a user is served by multiple base stations, the delivery delay for that user is the maximum delivery delay among all the base stations serving that user. The calculation expression is as follows: In the formula, For users The set of all serving base stations; therefore, the user exist The number of files transferred at any given time is: In the formula, For users rate, For time slot length, user exist The queue state at time t is: In the formula, User exist The number of transactions arriving at a given time, with a compliance strength of [value missing]. The Poisson distribution; The latency minimization problem is transformed into a queue length minimization problem, and the calculation expression is: In the formula, Represents a set of base stations. Represents a set of users; With the goal of minimizing the queue length, the expression for the transmission policy computed through multi-agent reinforcement learning is as follows: The process of multi-agent reinforcement learning includes: Action settings: The actions of a small cell base station include indexes for user scheduling and transmission power; definition For small cell base stations The set of available transmission powers is represented as: In the formula, It is the number of available transmission power levels. This is the maximum available transmit power; small cell base stations are based on... and actions Transmit power index in Select the transmit power; In the In each time slot, the base station The available actions are: In the formula, For base stations exist User index for time slot scheduling; State settings: The state includes channel state information and queue length information for all users; the useful signal power, interference power, and queue length of all users at the previous moment are used as state inputs to the network and base station. The status is: In the formula, Indicates the first The first time slot Useful signal power for each user Indicates the first The first time slot Interference signal power for each user; Reward settings: Same as reward settings in a single-user network, base station The The reward for each DQN network is: 。 8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the latency minimization method for joint caching and transmission in dense small cells as described in any one of claims 1 to 6.