Edge caching method and device based on multi-agent reinforcement learning model

By optimizing the edge caching method through a multi-agent reinforcement learning model, the problem of underutilized edge server partnerships is solved, the content hit rate and acquisition probability are improved, the latency is reduced, and the user experience is improved.

CN114185677BActive Publication Date: 2025-10-03HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111523410.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-10-03
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

Existing edge caching methods do not fully consider the cooperative relationship between the local edge server and the neighboring edge servers, resulting in increased latency in obtaining the requested content when the content requested by the terminal device is not cached in the local server, which reduces the user experience.

Method used

A multi-agent reinforcement learning model is used to determine the content cached by the local server at the next moment. By obtaining the popularity and storage status of the content, the storage strategy of the content on the local and neighboring servers is optimized to improve the content hit rate and acquisition probability.

Benefits of technology

By optimizing the storage strategy of content on local and neighboring servers, the latency of terminal devices requesting content is reduced, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114185677B_ABST
    Figure CN114185677B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an edge caching method and device based on a multi-agent reinforcement learning model. The method obtains information about multiple currently cached high-popularity content and medium-popularity content, including a content identifier, a first storage state, and a first popularity of the content. The first popularity indicates the probability of the content being requested, and the medium-popularity content can be used to cooperate with neighboring servers and be acquired by neighboring servers. The method processes the content identifier, the first storage state, and the first popularity through a multi-agent reinforcement learning model to obtain the target content identifier and target storage state of the target content to be cached at the next moment. The method also updates the currently cached content. The technical solution provided by the present application improves the hit rate of content requested by terminal devices in local servers and neighboring servers, thereby reducing the latency of content requests by terminal devices and effectively improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to an edge caching method and device based on a multi-agent reinforcement learning model. Background Art

[0002] As the number of devices connected to the Internet continues to increase, users are increasingly demanding bandwidth and latency for network communications, especially for emerging services such as autonomous driving and virtual reality. Edge caching, a technology that caches requested content on edge servers closer to end devices, allows users to retrieve cached content directly from the edge server without having to send it through the backbone network to a remote data center. Therefore, edge caching reduces latency and alleviates pressure on backbone networks and data centers.

[0003] Currently, edge caching methods randomly store portions of content on local edge servers and neighboring edge servers, while storing the entire content on a remote central server. Upon receiving a request from a terminal device, the local edge server first retrieves the requested content and sends it to the terminal device. If the local edge server doesn't cache the requested content, it retrieves the requested content from a neighboring edge server or a remote central server, ensuring that the terminal device receives the requested content.

[0004] However, the current edge caching method stores the content stored in the local edge server and the neighboring edge server according to the probability of user requests, and does not take into account the cooperative relationship between the local edge server and the neighboring edge server. It may happen that when the content requested by the terminal device does not exist in the local edge server, the content cannot be obtained in the neighboring edge server either, so that the content requested by the terminal device can only be obtained from the central server. When the content requested by the terminal device does not exist in the local edge server, the probability of obtaining the content in the neighboring edge server is small, which increases the delay in obtaining the requested content, thereby reducing the user experience. Summary of the Invention

[0005] An embodiment of the present application provides an edge caching method and device based on a multi-agent reinforcement learning model. By determining the content cached by the local server at the next moment through the multi-agent reinforcement learning model, the hit rate of the content can be improved, thereby enhancing the user experience.

[0006] In a first aspect, an embodiment of the present application provides an edge caching method based on a multi-agent reinforcement learning model, which is applied to a local server. The edge caching method based on a multi-agent reinforcement learning model includes:

[0007] Information about multiple contents currently cached is obtained, the information including a content identifier, a first storage state, and a first popularity of the contents, the multiple contents including high-popularity content having a first popularity greater than a first popularity threshold, and medium-popularity content having a first popularity less than the first popularity threshold and greater than a second popularity threshold, the first popularity threshold being greater than the second popularity threshold, the first popularity indicating a probability of the contents being requested, and the medium-popularity content being used to be requested by a terminal device or obtained by the neighboring server in cooperation with the neighboring server.

[0008] The content identifier, the first storage state, and the first popularity are processed through a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of the target content to be cached at the next moment.

[0009] The currently cached content is updated according to the target content identifier, the target storage state, and the target popularity corresponding to the target content.

[0010] In a second aspect, an embodiment of the present application provides an edge caching device based on a multi-agent reinforcement learning model, wherein the edge caching device based on the multi-agent reinforcement learning model includes:

[0011] an acquisition module, configured to acquire information of a plurality of currently cached contents, the information including a content identifier, a first storage state, and a first popularity of the contents, the plurality of contents including high-popularity contents having a first popularity greater than a first popularity threshold, and medium-popularity contents having a first popularity less than the first popularity threshold and greater than a second popularity threshold, the first popularity threshold being greater than the second popularity threshold, the first popularity indicating a probability of the contents being requested, the medium-popularity contents being intended to be requested by a terminal device or acquired by the proximity server in cooperation with the proximity server;

[0012] a processing module, configured to process the content identifier, the first storage state, and the first popularity through a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of target content to be cached at a next moment;

[0013] An updating module is configured to update the currently cached content according to the target content identifier, the target storage state, and the target popularity corresponding to the target content.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device, the electronic device comprising: a processor, and a memory communicatively connected to the processor;

[0015] The memory stores computer-executable instructions;

[0016] The processor executes the computer-executable instructions stored in the memory to implement the edge caching method based on the multi-agent reinforcement learning model described in any possible implementation of the first aspect above.

[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the edge caching method based on the multi-agent reinforcement learning model described in any possible implementation method of the first aspect above is implemented.

[0018] In the fifth aspect, an embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the edge caching method based on the multi-agent reinforcement learning model described in any possible implementation method of the first aspect above.

[0019] It can be seen that the embodiment of the present application provides an edge caching method and device based on a multi-agent reinforcement learning model, by obtaining information of multiple contents currently cached, the information includes a content identifier, a first storage state and a first popularity of the content, the multiple contents include high-popularity content whose first popularity is greater than a first popularity threshold, and medium-popularity content whose first popularity is less than the first popularity threshold and greater than the second popularity threshold, the first popularity threshold is greater than the second popularity threshold, the first popularity represents the probability of the content being requested, and the medium-popularity content is used to be requested by a terminal device or to be obtained by a neighboring server in cooperation with a neighboring server; the content identifier, the first storage state and the first popularity are processed by a multi-agent reinforcement learning model to obtain the target content identifier and target storage state of the target content cached at the next moment; according to the target content identifier, the target storage state and the target popularity of the target content, the currently cached content is updated. The technical solution provided in the embodiment of the present application obtains the target content to be cached at the next moment through a multi-agent reinforcement learning model and information about multiple contents currently cached by the local server, and takes into account the popularity of the content. At the same time, medium-popularity content is used to cooperate with neighboring servers. In addition to improving the hit rate of the content requested by the terminal device in the local server, it can also increase the probability of obtaining it in the neighboring servers, thereby reducing the delay of the terminal device requesting content and effectively improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram of an application scenario of an edge caching method based on a multi-agent reinforcement learning model provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of an edge caching method based on a multi-agent reinforcement learning model provided in an embodiment of the present application;

[0022] Figure 3 A flowchart of a method for determining a target content identifier and a target storage state of target content provided in an embodiment of the present application;

[0023] Figure 4 A schematic diagram of an architecture for storing content in a server provided in an embodiment of the present application;

[0024] Figure 5 A line diagram illustrating the total revenue generated by all contents cached by a server according to an embodiment of the present application;

[0025] Figure 6 A broken line diagram of a hit rate corresponding to content cached by a server provided in an embodiment of the present application;

[0026] Figure 7 A schematic diagram of the delay corresponding to the content cached by a server provided in an embodiment of the present application;

[0027] Figure 8 A schematic diagram of the structure of an edge cache device based on a multi-agent reinforcement learning model provided in an embodiment of the present application;

[0028] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this application.

[0029] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0031] In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural. In the text description of this application, the character " / " generally indicates that the associated objects are in an "or" relationship.

[0032] In recent years, with the continued increase in the number of connected devices and the reduction in data costs, the promotion of 5G has provided users with higher bandwidth and lower latency, leading to explosive traffic growth. Furthermore, the number of hyperscale data centers and traffic volume are rapidly increasing worldwide. At the same time, emerging services such as autonomous driving and virtual reality are placing higher demands on latency. Edge caching is a method that caches content requested by terminal devices at an edge server, closer to the terminal device. This allows users to directly access cached services requested by their terminal devices from the edge server, without having to transmit them through the backbone network to the data center. Edge caching reduces latency and reduces pressure on the backbone network and data centers.

[0033] When caching at the edge, the cached content is based on popularity. Content of varying popularity has different probability of being requested. Two key areas of research are pre-caching and collaborative caching. Pre-caching primarily focuses on determining what content the edge server should cache next based on historically known or learned information to maximize benefits. Edge servers have limited space, making it impossible to cache all content. Collaborative caching, on the other hand, focuses on how to enable edge servers to collaborate and efficiently utilize space to maximize overall system benefits.

[0034] Most current research on collaborative caching assumes that end devices can obtain requested content from their local edge servers. If the local edge server is not caching the content, the requested content is retrieved from neighboring edge servers or remote central servers. However, existing approaches fail to fully consider the impact of content placement on collaboration. Specifically, they fail to account for the probability that the requested content will be available from neighboring servers if the local server is not available. This increases latency for end devices to retrieve the requested content, degrading user experience.

[0035] In order to solve the problem of poor user experience caused by not considering the impact of content placement on latency in the server, it is possible to store popular content and medium-popular content in the local server based on the popularity of the content, that is, the probability of the content being requested. The medium-popular content can be used to cooperate with neighboring servers, and the content to be cached by the local server at the next moment is determined through a multi-agent reinforcement learning model. Since the result determined by the multi-agent reinforcement learning model is the one with the greatest benefit, that is, the content cached in the local server at the next moment is the one with the greatest system benefit, that is, the content with the greatest reduced latency, when the content requested by the terminal device is not cached in the local server, it can be obtained from the neighboring server, which can reduce the latency when the terminal device requests content from the local server, thereby improving the user experience. Among them, the multi-agent reinforcement learning model is a multi-agent multi-armed bandit model.

[0036] The technical solution provided by the embodiment of the present application can be applied to scenarios where edge server cache content is determined, especially content cached by a local server. Figure 1 A schematic diagram of an application scenario of an edge caching method based on a multi-agent reinforcement learning model provided in an embodiment of the present application. Figure 1 It includes multiple base stations, multiple users in each base station, and a remote central server. The central server caches all content, and a storage server is deployed at each base station to cache content. And the storage capacity of each storage server is limited, so the content cached in the storage server is part of the total content stored in the central server. The corresponding circular area under each base station is the coverage area of ​​the base station signal, that is, the wireless transmission range. In order to ensure the service guarantee of obtaining content from neighbors, base stations with overlapping wireless transmission ranges are regarded as adjacent base stations, and there are wired connections between them, so that content can be transmitted through the wired connection between the base stations. According to Figure 1 As shown, all base stations can be connected to the central server via a backhaul link, so that any desired content can be obtained from the central server at any time.

[0037] For example, Figure 1 Among the servers corresponding to the base stations of neighboring servers, when any server acts as a local server, the content cached in the local server is determined by a multi-agent reinforcement learning model. That is to say, when the neighboring server acts as a local server, the cached content is also determined by a multi-agent reinforcement learning model. It is understandable that when determining the content cached in the local server through the multi-agent reinforcement learning model, it is necessary to consider the popularity of the content in the local server, as well as the influence of the neighboring server on its cached content, and determine the content with the greatest benefit, that is, the greatest reduction in delay, as the content cached by the local server at the next moment. This includes determining the content identifier and content storage status of the cached content at the next moment.

[0038] For example, according to Figure 1As shown, when the local server receives a request instruction sent by a user within the wireless transmission range of its corresponding base station through a terminal device, it can determine the corresponding content in the cache of the local server based on the content requested by the request instruction, and then send the content to the terminal device. If the requested content does not exist in the local server, the local server sends a request to the neighboring server for the content. If it does not exist in the neighboring server, it requests the content from the central server. Since the content cached in the local server is the content with the greatest benefit determined after considering the influence of the neighboring servers, the probability of caching the content of the terminal device's request instruction in the local server is relatively high. Even when the content of the terminal device's request instruction is not cached from the local server, the corresponding content can be sent to the terminal device while ensuring low latency, thereby effectively improving the user experience.

[0039] The neighboring servers are connected for transmission via a wired connection, and when the central server sends the requested content to the local server, it is transmitted via a backhaul link between the local server and the central server. The embodiment of the present application does not impose any limitation on the specific transmission process.

[0040] Below, the edge caching method based on the multi-agent reinforcement learning model provided by this application will be described in detail through specific embodiments. It is understandable that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0041] Figure 2 This is a flow chart of an edge caching method based on a multi-agent reinforcement learning model provided in an embodiment of the present application. The edge caching method based on a multi-agent reinforcement learning model can be executed by software and / or hardware devices. For example, the hardware device can be an edge caching device based on a multi-agent reinforcement learning model, and the edge caching device based on a multi-agent reinforcement learning model can be a local server or a processing chip in a local server. For an example, see Figure 2 As shown, the edge caching method based on the multi-agent reinforcement learning model may include:

[0042] S201. Obtain information about multiple contents currently cached, where the information includes a content identifier, a first storage state, and a first popularity of the contents, the multiple contents including high-popularity contents having a first popularity greater than a first popularity threshold, and medium-popularity contents having a first popularity less than the first popularity threshold and greater than a second popularity threshold.

[0043] For example, the first popularity threshold is greater than the second popularity threshold. The first popularity indicates the probability of the content being requested, and the medium-popularity content is used to be requested by a terminal device or obtained by a neighboring server through cooperation with the neighboring server. It is understood that the medium-popularity content cached in the two servers serving as neighboring servers can be as different as possible, so that if the content requested by the terminal device is not cached in the local server, it can be obtained from the neighboring server.

[0044] For example, the first popularity threshold and the second popularity threshold can be automatically determined by an algorithm, and the embodiments of the present application do not impose any restrictions on this. It is understandable that when the algorithm determines the first popularity threshold and the second popularity threshold, it may be determined based on the size of the storage space of the local server or the type of cached content, or based on other parameters, and the embodiments of the present application do not impose any specific restrictions on this. The first popularity threshold and the second popularity threshold of each local server may be the same or different, and the embodiments of the present application do not impose any restrictions on this.

[0045] For example, the method provided in the embodiment of the present application can be used for simulation research. When simulating the method provided in the embodiment of the present application, the popularity of each to-be-cached content among all to-be-cached content can be determined according to the Zipf distribution of the following formula (1):

[0046]

[0047] In formula (1), p i represents the popularity of the content i to be cached, α k represents the parameter corresponding to the local server k, N represents the number of contents i to be cached, and N>1.

[0048] Furthermore, the content to be cached whose popularity is greater than a first popularity threshold can be determined as high-popularity content, and the content whose popularity is less than the first popularity threshold and greater than a second popularity threshold can be determined as medium-popularity content, so that the high-popularity content and the medium-popularity content are cached in the local server.

[0049] In the embodiment of the present application, a set of multiple base stations is established, and all base stations are represented as k is the base station numbered k, that is, the local server corresponding to the base station numbered k. The content identifier can be represented by f, and the status of all content cached in the local server is represented by X={x kf}, where x kf represents the state of content f in base station k, x kf =1 means that content f is cached in base station k, x kf= 0 means that the content is not cached in base station k. The popularity of the content is expressed as P = {p kf}, where p kf represents the popularity of content f at base station k, that is, the frequency with which content f is requested. It is understandable that during the simulation, the content cached by the local server can be initialized, that is, all the above information is kept in the most initial state, that is, all x kf =0, no content is cached in the local server, all p kf =0, that is, all base stations do not know any popularity-related content in advance, and thus determine the content cached in the local server according to the above method.

[0050] It is understandable that the content requested by the user's terminal device is first obtained from the local server, that is, the server corresponding to the nearest base station. If the local server has the content cached, the content is transmitted to the terminal device. When the local server does not cache the content, it asks for help from the neighboring server. The neighboring server corresponding to the neighboring base station that caches the requested content is transmitted to the destination base station through the backhaul link between the base stations, that is, the wired connection, and then transmitted to the user. If the content is not cached in the neighboring server, the content is requested from the central server. According to the above, content can be divided into three categories according to its popularity: high-popularity content, that is, content with high popularity; medium-popularity content, that is, content with average popularity; and low-popularity content, that is, content with low popularity. In order to obtain highly popular content as quickly as possible, most of the highly popular content can be cached in the server corresponding to each base station. For content with average popularity, in order to reduce the number of times this content is obtained from the central server, the remaining space of the servers corresponding to all base stations can be reasonably utilized. It is hoped that the remaining storage space of the server corresponding to each base station can cache different content with average popularity. In this way, this content can serve the terminal devices of users of servers corresponding to as many neighboring base stations as possible, thereby increasing the possibility of obtaining content with average popularity on the server corresponding to the local base station or the server corresponding to the neighboring base station.

[0051] During the simulation, the content cached on the local server is initialized so that no content is cached on the local server. The popularity of the content cached on the local server is determined by Zipf distribution, so that the content that needs to be cached on the local server can be determined based on the popularity of the content. This can increase the probability that the terminal device obtains the requested content on the local server, effectively reduce latency, and thus improve the user experience.

[0052] S202: Process the content identifier, the first storage state, and the first popularity through a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of the target content to be cached at the next moment.

[0053] For example, in this application, the multi-agent reinforcement learning model is a multi-agent multi-armed bandit model. When the content identifier, the first storage state, and the first popularity are processed by the multi-agent reinforcement learning model, the action of executing one arm of the multi-armed bandit can be regarded as determining the content corresponding to the cache. The multiple arms selected at one time constitute a super arm in an action, that is, a joint action, which represents the content state of the base station cache at the next moment. For example, the joint action can be expressed as X k ={x k1 ,x k2 ...x kN}.

[0054] For example, the content cached in the local server at the next moment, that is, the target content identifier and target storage status of the target content cached at the next moment, can be determined through the first storage status and first popularity of each content in the local server at the current moment, the second storage status and second popularity in the neighboring server, the first delay for the neighboring server to send the content to the local server, and the second delay for the central server to send the content to the local server.

[0055] S203: Update the currently cached content according to the target content identifier, the target storage state, and the target popularity corresponding to the target content.

[0056] For example, when the currently cached content is updated based on the target content identifier, target storage status and target popularity corresponding to the target content, all the currently cached content can be deleted, and the target content, i.e., the target content identifier, target storage status and target popularity of the target content, can be cached in the local server. Alternatively, the currently cached content can be compared with the target content to be cached at the next moment, and the first storage status and first popularity of the content whose storage status or popularity has changed can be updated to the target storage status and target popularity.

[0057] It can be seen that the edge caching method based on the multi-agent reinforcement learning model provided by the embodiment of the present application obtains information of multiple contents currently cached, the information including content identification, first storage state and first popularity of the content, the multiple contents including high popularity content whose first popularity is greater than the first popularity threshold, and medium popularity content whose first popularity is less than the first popularity threshold and greater than the second popularity threshold, the first popularity threshold is greater than the second popularity threshold, the first popularity represents the probability of the content being requested, and the medium popularity content is used to be requested by the terminal device or obtained by the neighboring server in cooperation with the neighboring server; the content identification, the first storage state and the first popularity are processed by the multi-agent reinforcement learning model to obtain the target content identification and target storage state of the target content cached at the next moment; according to the target content identification, the target storage state and the target popularity of the target content, the currently cached content is updated. The technical solution provided by the embodiment of the present application obtains the target content cached at the next moment through the multi-agent reinforcement learning model and the information of multiple contents currently cached by the local server. Since the multi-agent reinforcement learning model can determine the result with the greatest benefit, the target content cached at the next moment is determined to be the content with the greatest delay reduction. This approach also takes content popularity into consideration, caching as many different moderately popular content as possible on neighboring servers, expanding the content capacity of the entire edge system. Furthermore, in addition to improving the hit rate of content requested by terminal devices on the local server, it also increases the probability of obtaining it on neighboring servers, thereby reducing the latency of content requests from terminal devices and effectively improving the user experience.

[0058] This application embodiment will provide a detailed description of the process of determining target content through a multi-agent reinforcement learning model. For details, please refer to Figure 3 As shown, Figure 3 A flowchart of a method for determining the target content identification and target storage status of target content provided in an embodiment of the present application. The method for determining the target content identification, target storage status, and target popularity of target content can be executed by software and / or hardware devices. For example, the hardware device can be an edge cache device based on a multi-agent reinforcement learning model, and the edge cache device based on a multi-agent reinforcement learning model can be a local server or a processing chip in a local server. For example, the method for determining the target content identification and target storage status of target content can include:

[0059] S301. For each content, based on the content identifier, obtain a first delay for the neighboring server to send the content to the local server, and a second delay for the central server to send the content to the local server, and obtain a second storage status and a second popularity of the content in the neighboring server.

[0060] According to the above embodiment, the content identifier is represented by f. Therefore, the storage status and popularity of each content in the neighboring server, ie, the second storage status and the second popularity, can be obtained in the neighboring server according to the content identifier of the content.

[0061] It is understandable that the first delay for the neighboring server to send content to the local server and the second delay for the central server to send content to the local server can be directly obtained in the local server or obtained through other methods. The embodiments of the present application do not impose any restrictions on the specific acquisition method. In addition, the first delay and the second delay are related to various factors, such as the distance between the servers, or other factors, which are not limited in the embodiments of the present application. Since the speed at which the neighboring server transmits content to the local server is greater than the speed at which the central server transmits content to the local server, the first delay is less than the second delay.

[0062] S302. Calculate the instantaneous benefit, average benefit, and benefit estimate corresponding to the cached content based on the first storage state, the first popularity, the second storage state, the second popularity, the first delay, and the second delay. The instantaneous benefit represents the delay reduction corresponding to the content.

[0063] For example, when calculating the instantaneous benefit, average benefit, and estimated benefit corresponding to the cached content, the instantaneous benefit corresponding to the cached content can be calculated according to the following formula (2):

[0064]

[0065] In formula (2), represents the instantaneous benefit corresponding to content f, x kf represents the first storage state of content f in the local server k, p kf represents the first popularity of content f in local server k, d s represents the second delay, p k'f represents the second popularity of content f in neighboring server k', x k'f represents the second storage state of content f in the neighboring server k', d n represents the first delay, represents the set of all servers, k represents the local server, N represents the number of contents f, and N>1.

[0066] Furthermore, the average revenue corresponding to the cached content is calculated according to the following formula (3):

[0067]

[0068] In formula (3), represents the average revenue corresponding to caching content f on local server k at time t, represents the average revenue corresponding to caching content f on local server k at time t-1, Represents the number of times content f is cached on the local server k until time t-1.

[0069] Furthermore, according to the following formula (4), the estimated revenue value corresponding to the cached content is calculated:

[0070]

[0071] In formula (2), represents the estimated revenue corresponding to caching content f on local server k at the current time t, It represents the average revenue corresponding to caching content f on the local server k at time t-1.

[0072] It is understandable that when the multi-agent reinforcement learning model is a multi-agent multi-armed bandit model, there is an exploration and exploitation (EE) problem in the multi-armed bandit model, and the estimated value of the reward of the arm executing the multi-armed bandit can be calculated by the UCB (Upper Confidence Bound) algorithm.

[0073] For example, there are multiple users, that is, multiple terminal devices, within the range of each base station, represented by U = {u}, and each user's terminal device u corresponds to a local base station, that is, the nearest base station, represented by represents the user's terminal device d with base station k as the local base station uk' is the distance between the user's terminal device u and the base station k'.

[0074] In the embodiment of the present application, the formula can accurately calculate the instantaneous benefit, average benefit and benefit estimate corresponding to the cached content, which can improve the accuracy of the determined target content, thereby increasing the reduction in latency and improving user experience.

[0075] S303: Determine the temporary content identifier and temporary storage state of the temporary content to be cached at the next moment according to the instantaneous revenue, the average revenue and the estimated revenue.

[0076] For example, when determining the temporary content identifier and temporary storage state of the temporary content cached at the next moment based on the instantaneous revenue, the average revenue and the revenue estimation value, a preset number of contents with the largest revenue estimation value of the contents currently cached by the local server can be determined as the temporary content of the local server, that is, the contents with the largest revenue estimation value can be selected. The largest C kArms form a super arm, and the content determined by the multi-agent reinforcement learning model cannot exceed the cache storage space.

[0077] It is understandable that in the embodiment of the present application, the determined temporary content needs to meet the following constraints: The rest x=0. Among them, f1, f2, f3……f Ck Both are content identifiers. This significantly reduces the latency of content cached in the local server, effectively reducing latency.

[0078] S304: Repeat the above steps for the temporary content until the target content identifier and target storage state of the target content that meets the preset conditions are obtained.

[0079] It is understandable that the temporary content determined in the above step 303 is not necessarily the temporary content with the greatest instantaneous benefit. Therefore, in order to obtain the temporary content with the greatest instantaneous benefit, the obtained temporary content can be used as the content currently cached by the server, and the above steps S301-S303 can be repeated to determine the temporary content with the greatest instantaneous benefit, and the final determined temporary content can be used as the target content to update the content currently cached by the local server.

[0080] For example, when the above steps are repeatedly performed on temporary content until the target content identifier and target storage state of the target content that meets the preset conditions are obtained, the total revenue estimate corresponding to the temporary content can be calculated during each execution; when the difference between the total revenue estimate of the current temporary content and the total revenue estimate of the last temporary content is less than a preset threshold, or when the number of times the above steps are repeated reaches the loop count threshold, a preset number of temporary contents are determined as target content in descending order of the revenue estimates corresponding to the temporary content, and the target content identifier and target storage state are determined. It is understandable that when the above steps are repeatedly performed on the temporary content, the temporary popularity corresponding to the temporary content also needs to be determined. The loop count threshold can be set according to the specific situation and is not limited in this embodiment of the present application. In order to ensure that the obtained temporary content is convergent, the preset threshold is generally a very small value, that is, the difference between the total revenue estimate of the current temporary content and the total revenue estimate of the last temporary content is small. For example, the system can be updated with the assistance of gradient ascent (Coordinate Ascent).

[0081] It is understandable that a preset number of temporary contents are determined as target contents in descending order of the estimated revenue values ​​corresponding to the obtained temporary contents. Please refer to the above steps, and the embodiments of the present application will not be repeated here.

[0082] It is understandable that when determining the temporary content, it is not necessary to update the content currently cached by the local server, but to update it when determining the final temporary content.

[0083] In an embodiment of the present application, when the difference between the total revenue estimate of the currently obtained temporary content and the total revenue estimate of the previously obtained temporary content is less than a preset threshold, the target content is determined, so that the total revenue estimate of the target content converges, and temporary content with a larger total revenue estimate is obtained. In addition, when the number of iterations of the above steps reaches a loop count threshold, the temporary content is determined, ensuring that if a convergent solution cannot be obtained, the loop can be stopped and the final temporary content can be determined.

[0084] For example, when calculating the estimated total revenue corresponding to the temporary content, the estimated total revenue corresponding to the temporary content can be calculated according to the following formula (5):

[0085]

[0086] In formula (5), The total estimated revenue corresponding to all temporary contents cached by the local server, represents the estimated value of revenue corresponding to content f, represents the set of all servers, k represents the local server, N represents the number of contents f, and N>1.

[0087] For example, the content stored in the local server satisfies the constraint price adjustment described in the following formula (6):

[0088]

[0089] In formula (6), x kf represents the storage status of content f in the local server k, c k Represents the storage space of the local server k.

[0090] In an embodiment of the present application, by calculating the total revenue estimate of all temporary content cached by the local server and constraining the content stored in the local server through the storage space of the local server, the calculated total revenue estimate is made more accurate.

[0091] It can be seen that the method for determining the target content identifier, target storage state and target popularity of the target content provided by the embodiment of the present application is, for each content, based on the content identifier, respectively obtaining the first delay for the neighboring server to send the content to the local server, and the second delay for the central server to send the content to the local server, and obtaining the second storage state and second popularity of the content in the neighboring server; based on the first storage state, first popularity, second storage state, second popularity, first delay and second delay, calculating the instantaneous benefit, average benefit and benefit estimate corresponding to the cached content, where the instantaneous benefit represents the amount of delay reduction corresponding to the content; based on the instantaneous benefit, average benefit and benefit estimate, determining the temporary content identifier and temporary storage state of the temporary content cached at the next moment; repeating the above steps for the temporary content until the target content identifier and target storage state of the target content that meets the preset conditions are obtained. The multi-agent reinforcement learning model can accurately determine the target content, so that the terminal device has a high probability of directly obtaining the requested content in the local server, and can effectively provide user experience.

[0092] In order to facilitate the understanding of the edge caching method based on the multi-agent reinforcement learning model provided by the embodiment of the present application, the edge server can be regarded as multiple agents, considering the status of the neighboring servers. In order to reduce the spatial size of the joint action, the decision problem of cache placement is modeled as a multi-agent multi-armed bandit problem. In order to optimize the joint decision, the coordinate ascent method can be used for auxiliary update. Below, the edge caching method based on the multi-agent reinforcement learning model provided by the embodiment of the present application will be described in detail through specific steps. For example, the edge caching method based on the multi-agent reinforcement learning model provided by the embodiment of the present application may include the following steps:

[0093] Step 1: Initialize the status of the base station and content.

[0094] In this step, initialization is performed so that the base station initially does not store any information related to popularity or collaboration with neighboring base stations. In other words, the base station's server does not cache any content or store any popularity information. As described in the above embodiment, this step can be performed during simulation, but the method described in this step may not be used during the use of the local server.

[0095] Step 2: Set the base station cooperation strategy.

[0096] Since only local methods cannot fully utilize the cooperative relationship between base stations, and only consider the hit rate, that is, the probability that the server stores the content requested by the terminal device, the method treats the space of all base stations as a whole and fully utilizes it, but sacrifices latency. This application comprehensively considers the above two aspects. For content with high popularity, each base station prioritizes caching content with high popularity, and the remaining small amount of space is used for cooperation, that is, for providing services to neighboring base stations. At the same time, medium-popular content is cached to provide services to all neighboring servers that have not cached a certain content.

[0097] Step 3: Establish all base stations in the scenario into a multi-agent multi-armed bandit system.

[0098] In this step, each base station is treated as an agent, and a multi-agent multi-armed bandit system is established for all base stations. In this system, each content corresponds to an arm, and executing an arm represents caching that content. Assuming that all content is the same size, each base station selects a super arm that combines multiple ordinary arms for execution each time. That is, the instantaneous profit, average profit, and profit estimate of the content are determined using the above formulas (2), (3), and (4). The content cached in the base station's server must meet the constraints of formula (6).

[0099] Step 4: Select the arm solution with the largest overall benefit based on the learned instantaneous benefit of each arm.

[0100] For example, selecting the arm plan with the largest overall benefit is to determine the temporary content with the largest estimated benefit value.

[0101] Step 5: Based on the historical information learned, use the CA algorithm to assist in optimizing the decision-making of the multi-armed bandit to determine the content to be cached at the next moment and update the placement of the content.

[0102] For example, the CA algorithm can be used to determine the cache content at the next moment. For details, see the following steps:

[0103] Step 51: Save a temporary cache status table for each base station and start a loop.

[0104] Step 52: Generate a copy X' of the cache state at the current moment.

[0105] Step 53: In each loop, traverse each base station, and when it is the turn of base station k, keep the temporary cache status of other base stations The estimated benefit of each arm of base station k is updated according to formula (4), and the constraints need to be met according to the temporary content: For the rest x=0, the super arm is selected to update the cache state copy X'. The result of the execution is not used to update the cache state of the base station at the current moment, but to update the temporary cache state. Therefore, in the loop process constant.

[0106] Step 54: Repeat step 52 until the total revenue estimate does not change or exceeds the number of cycles. The final temporary content is the target content, that is, the final temporary cache state is the cache state at the next moment.

[0107] Step 6: The user's request content and the base station receives and services the request.

[0108] In this step, a time period can be set during which multiple terminal devices initiate multiple requests. The multi-agent system is only responsible for processing the requests and does not update the cache. Assuming that each terminal device only requests one piece of content at a time, each base station records the number of times the updated content appears and the total number of requests upon receiving the content to estimate its popularity. This embodiment of the application does not impose any restrictions on the specific estimation method.

[0109] Step 7: Repeat steps 4 to 6 until the entire system converges.

[0110] It can be seen that the edge caching method based on the multi-agent reinforcement learning model provided by the present application, by modeling each base station into a multi-agent system, multiple base stations cooperate to maximize the global benefit of the system, and the strategy adopted for the use of the storage space of each base station is to cache the most popular part of each space, and cache the content with average popularity in the other part of the space, and as many neighboring base stations as possible that have not cached the content. If the content requested by the terminal device is not stored in the local base station, the content can be requested from the neighboring base station. The above problem is modeled as an integer programming problem. In order to maximize the global benefit of the system, the integer programming problem is converted into a multi-agent multi-armed bandit problem. The coordinate ascent algorithm is used to assist in optimizing the benefit of each arm of the multi-armed bandit, thereby reducing the delay of the edge cache and improving the hit rate, that is, the probability that the terminal device can directly obtain the requested content in the local base station.

[0111] According to the method described in the above embodiment, the content cached by the local server and the neighboring server in the embodiment of the present application will be described below through a specific embodiment. For details, please refer to Figure 4 As shown, Figure 4 A schematic diagram of an architecture for storing content in a server provided in an embodiment of the present application.

[0112] according to Figure 4As shown, the central server stores all contents 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10; the server of base station 1 stores contents 1, 2, 3, and 5; the server of base station 2 stores contents 1, 2, 3, and 4; and the server of base station 3 stores contents 1, 2, 5, and 6.

[0113] For example, terminal device 1 within the effective radio transmission range of base station 2 requests content 6 from base station 2, but content 6 is not stored in the server of base station 2. Then base station 2 requests content 6 from its neighboring base station, namely base station 3, and upon receiving content 6 returned by base station 3, base station 2 sends content 6 to terminal device 1. Terminal device 2 within the effective radio transmission range of base station 2 requests content 1 from base station 2. Since content 1 is cached in the server of base station 2, base station 2 can directly send content 1 to terminal device 2. When terminal device 3 within the effective radio transmission range of base station 2 requests content 9 from base station 2, since content 9 is not cached in base station 2 and its neighboring base stations, namely base station 1 and base station 3, base station 2 requests content 9 from the central server, and upon receiving content 9 returned by the central server, base station 2 sends content 9 to terminal device 3.

[0114] Furthermore, in order to demonstrate the practicality of the edge caching method based on the multi-agent reinforcement learning model provided in the embodiment of the present application, the method provided in the present application, the method that only considers the hit rate of the content, and the method that only considers the local server without cooperating with neighboring servers are compared. Specifically, the instantaneous benefit, hit rate, and latency generated by the cached content in the server determined by the three methods are compared. In this application, method 1 represents the method that only considers the local server without cooperating with neighboring servers, method 2 represents the method that only considers the hit rate of the content, and method 3 represents the method provided in the present application.

[0115] Figure 5 A line diagram of the total revenue generated by all content cached by a server according to an embodiment of the present application is provided. Instantaneous revenue is the reduction in latency. For example, the total revenue generated by all content cached by the server can be determined using the following formula (7):

[0116]

[0117] In formula (7), D reduce (X) represents the total revenue corresponding to all the content cached by the server, represents the instantaneous benefit corresponding to content f, represents the set of all servers, k represents the server numbered k, N represents the number of contents f, and N>1.

[0118] according to Figure 5 It can be seen that Figure 5The vertical axis represents the total revenue of all the contents cached in the server, and the horizontal axis represents time. Alternatively, the horizontal axis can also be used to represent the number of updates of the cached contents in the server. Figure 5 As can be seen, as the number of cached content updates on the server increases, the total revenue corresponding to the content cached by the three methods continues to increase until it reaches a certain threshold, at which point it no longer increases. Ultimately, the total revenue of the cached content determined by the edge caching method based on the multi-agent reinforcement learning model of this application is the highest, followed by method 2, and the lowest total revenue is obtained by method 1.

[0119] Figure 6 A line diagram of the hit rate corresponding to the content cached by a server provided in an embodiment of the present application. Figure 6 The vertical axis represents the hit rate of all the contents cached in the server, and the horizontal axis represents time, that is, the horizontal axis represents the number of times the cached contents in the server are updated. Figure 6 As can be seen, as the number of updates to the cached content on the server increases, the hit rates of the three methods continue to increase, until they reach a certain threshold, where they no longer increase. Ultimately, Method 2 achieves the highest hit rate, followed by the cached content hit rate determined by the edge caching method based on the multi-agent reinforcement learning model of this application, considering only the local server. Method 1 achieves the lowest hit rate.

[0120] Figure 7 A broken line diagram of the delay corresponding to the content cached by a server provided in an embodiment of the present application. Figure 7 The vertical axis represents the corresponding delay of all the contents cached in the server, and the horizontal axis represents time, that is, the horizontal axis represents the number of updates of the cached contents in the server. Figure 7 As can be seen, as the number of updates to the cached content in the server increases, the latency corresponding to the content cached by the three methods continues to decrease until it reaches a certain threshold and then stops decreasing. Ultimately, Method 2 has the highest latency, considering only the local server, followed by Method 1. The edge caching method based on the multi-agent reinforcement learning model of this application determines the latency corresponding to the cached content. In other words, the method provided by this application determines the latency corresponding to the cached content.

[0121] In summary, the edge caching method based on a multi-agent reinforcement learning model provided in the embodiments of the present application determines the content to be cached by the server at the next moment, that is, determines the storage status of the content, by comprehensively considering the hit rate, i.e., popularity, and latency. This ensures that the latency and instantaneous benefit of the content cached by the server at the next moment are both optimal, while ensuring a high hit rate, effectively improving the user experience.

[0122] Figure 8This is a schematic diagram of the structure of an edge cache device 80 based on a multi-agent reinforcement learning model provided in an embodiment of the present application. For example, see Figure 8 As shown, the edge caching device 80 based on the multi-agent reinforcement learning model may include:

[0123] Acquisition module 801 is used to obtain information of multiple contents currently cached, the information including content identification, first storage status and first popularity of the contents, the multiple contents including high-popularity content whose first popularity is greater than a first popularity threshold, and medium-popularity content whose first popularity is less than the first popularity threshold and greater than a second popularity threshold, the first popularity threshold is greater than the second popularity threshold, the first popularity indicates the probability of the content being requested, and the medium-popularity content is used to be requested by a terminal device or obtained by a neighboring server in cooperation with a neighboring server.

[0124] The processing module 802 is used to process the content identifier, the first storage state and the first popularity through a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of the target content to be cached at the next moment.

[0125] The updating module 803 is configured to update the currently cached content according to the target content identifier, the target storage state and the target popularity corresponding to the target content.

[0126] In one possible implementation, the processing module 802 is specifically configured to obtain, for each content, a first delay for the neighboring server to send the content to the local server and a second delay for the central server to send the content to the local server, respectively, based on the content identifier, and obtain a second storage state and a second popularity of the content in the neighboring server; calculate the instantaneous benefit, average benefit, and estimated benefit corresponding to the cached content based on the first storage state, the first popularity, the second storage state, the second popularity, the first delay, and the second delay, where the instantaneous benefit represents the amount of delay reduction corresponding to the content; determine the temporary content identifier and temporary storage state of the temporary content to be cached at the next moment based on the instantaneous benefit, the average benefit, and the estimated benefit, and determine the temporary popularity corresponding to the temporary content based on the temporary content identifier; and repeat the above steps for the temporary content until the target content identifier and target storage state of the target content that meet the preset conditions are obtained.

[0127] In a possible implementation, the processing module 802 is specifically configured to calculate the output according to the formula: Calculate the instantaneous benefit corresponding to the cached content; where, represents the instantaneous benefit corresponding to content f, x kf represents the first storage state of content f in the local server k, p kf represents the second popularity of content f in local server k, ds represents the second delay, p k'f represents the first popularity of content f in neighboring server k', x k'f represents the second storage state of content f in the neighboring server k', d n represents the first delay, represents the set of all servers, k represents the local server, N represents the number of contents f, and N>1.

[0128] According to the formula: Calculate the average revenue corresponding to the cached content; where, represents the average revenue corresponding to caching content f on local server k at time t, represents the average revenue corresponding to caching content f on local server k at time t-1, Represents the number of times content f is cached on the local server k until time t-1.

[0129] According to the formula: Calculate the estimated benefit corresponding to the cached content; where, represents the estimated revenue corresponding to caching content f on local server k at the current time t, It represents the average revenue corresponding to caching content f on the local server k at time t-1.

[0130] In one possible implementation, the processing module 802 is specifically configured to calculate the total revenue estimate corresponding to the temporary content each time it is executed; when the difference between the total revenue estimate of the currently obtained temporary content and the total revenue estimate of the last obtained temporary content is less than a preset threshold, or when the number of times the above steps are repeated reaches a loop count threshold, a preset number of temporary contents are determined as target contents in descending order of the revenue estimate corresponding to the obtained temporary contents, and the target content identifier and target storage status are determined.

[0131] In a possible implementation, the processing module 802 is specifically configured to calculate the output according to the formula: and formula Calculate the estimated total revenue corresponding to the temporary content; where, The total estimated revenue corresponding to all temporary contents cached by the local server, represents the estimated value of revenue corresponding to content f, represents the set of all servers, k represents the local server, N represents the number of content f, N>1, x kf represents the storage status of content f in the local server k, c k Represents the storage space of the local server k.

[0132] In one possible implementation, the neighboring servers meet the following constraints: Where k' represents the neighboring server, represents the set of all neighboring servers, d kk' represents the straight-line distance between the local server k and the neighboring server k', and r represents the effective radio transmission range of the base station corresponding to the server.

[0133] The edge caching device based on the multi-agent reinforcement learning model provided in the embodiment of the present application can execute the technical solution of the edge caching method based on the multi-agent reinforcement learning model in any of the above embodiments. Its implementation principle and beneficial effects are similar to the implementation principle and beneficial effects of the edge caching method based on the multi-agent reinforcement learning model. Please refer to the implementation principle and beneficial effects of the edge caching method based on the multi-agent reinforcement learning model, which will not be repeated here.

[0134] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 9 As shown, the electronic device 900 may include: at least one processor 901 and a memory 902 .

[0135] The memory 902 is used to store programs. Specifically, the programs may include program codes, and the program codes include computer operation instructions.

[0136] The memory 902 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0137] The processor 901 is used to execute the computer-executable instructions stored in the memory 902 to implement the edge caching method based on the multi-agent reinforcement learning model described in the aforementioned method embodiment. Among them, the processor 901 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. Specifically, when implementing the edge caching method based on the multi-agent reinforcement learning model described in the aforementioned method embodiment, the electronic device may be, for example, an electronic device with processing functions such as a terminal and a server. When implementing the edge caching method based on the multi-agent reinforcement learning model described in the aforementioned method embodiment, the electronic device may be, for example, an electronic control unit on a vehicle.

[0138] Optionally, the electronic device 900 may further include a communication interface 903. In a specific implementation, if the communication interface 903, the memory 902, and the processor 901 are implemented independently, the communication interface 903, the memory 902, and the processor 901 may be interconnected via a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified as address buses, data buses, control buses, etc., but this does not mean that there is only one bus or only one type of bus.

[0139] Optionally, in a specific implementation, if the communication interface 903, the memory 902 and the processor 901 are integrated on a chip, the communication interface 903, the memory 902 and the processor 901 can complete communication through an internal interface.

[0140] The present application also provides a computer-readable storage medium, which may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes. Specifically, the computer-readable storage medium stores program instructions, and the program instructions are used for the methods in the above embodiments.

[0141] The present application also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of an electronic device can read the execution instructions from the readable storage medium, and at least one processor executes the execution instructions so that the electronic device implements the edge caching method based on the multi-agent reinforcement learning model provided in various embodiments described above.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An edge caching method based on a multi-agent reinforcement learning model, characterized in that: Applicable to local servers, including: Obtaining information about multiple contents currently cached, the information including content identifiers, first storage states, and first popularity of the contents, the multiple contents including highly popular contents having a first popularity greater than a first popularity threshold, and medium popular contents having a first popularity less than the first popularity threshold and greater than a second popularity threshold, the first popularity threshold being greater than the second popularity threshold, the first popularity indicating a probability of the contents being requested, the medium popular contents being requested by a terminal device or obtained by the neighboring server in cooperation with the neighboring server; Processing the content identifier, the first storage state, and the first popularity through a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of target content to be cached at a next moment; updating the currently cached content according to the target content identifier, the target storage state, and the target popularity corresponding to the target content; The processing of the content identifier, the first storage state, and the first popularity by a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of target content to be cached at a next moment includes: Step 1: For each content, based on the content identifier, obtain a first time delay for the neighboring server to send the content to the local server, and a second time delay for the central server to send the content to the local server, and obtain a second storage status and a second popularity of the content in the neighboring server; Step 2: Calculate an instantaneous benefit, an average benefit, and an estimated benefit corresponding to caching the content based on the first storage state, the first popularity, the second storage state, the second popularity, the first delay, and the second delay, where the instantaneous benefit represents a reduction in delay corresponding to the content. Step 3: determining a temporary content identifier and a temporary storage state of the temporary content to be cached at the next moment based on the instantaneous revenue, the average revenue, and the estimated revenue value, and determining a temporary popularity corresponding to the temporary content based on the temporary content identifier; Step 4: Repeat steps 1-3 for the temporary content until a target content identifier and a target storage state of the target content that meets preset conditions are obtained.

2. The method according to claim 1, characterized in that The calculating, based on the first storage state, the first popularity, the second storage state, the second popularity, the first delay, and the second delay, of an instantaneous benefit, an average benefit, and a benefit estimate corresponding to caching the content includes: According to the formula: , calculating the instantaneous benefit corresponding to caching the content; in, represents the instantaneous benefit corresponding to content f, represents the first storage state of content f in the local server k, represents the first popularity of content f in local server k, represents the second delay, Indicates that content f is on a nearby server The second most popular Indicates that content f is on a nearby server The second storage state in represents the first delay, ={k}, represents the set of all servers, k represents the local server, N represents the number of content f, N>1; According to the formula: , calculate the average revenue corresponding to caching the content; in, represents the average revenue corresponding to caching the content f on the local server k at time t, represents the average revenue corresponding to caching the content f on the local server k at time t-1, represents the number of times content f is cached on the local server k until time t-1; According to the formula: , calculating an estimated benefit corresponding to caching the content; in, represents the estimated revenue value corresponding to caching the content f on the local server k at the current time t, represents the average revenue corresponding to caching the content f in the local server k at time t-1.

3. The method according to claim 2, characterized in that Repeating the above steps on the temporary content until a target content identifier and a target storage state of the target content meeting preset conditions are obtained includes: During each execution, calculating an estimated total revenue corresponding to the temporary content; When the difference between the total revenue estimate value of the temporary content currently obtained and the total revenue estimate value of the temporary content obtained last time is less than a preset threshold, or when the number of times the above steps are repeated reaches the loop number threshold, a preset number of the temporary contents are determined as target contents in descending order of the revenue estimate values ​​corresponding to the temporary contents, and the target content identifier and target storage status are determined.

4. The method according to claim 3, characterized in that The calculating the estimated total revenue corresponding to the temporary content includes: According to the formula: and formula , calculating an estimated total revenue corresponding to the temporary content; in, The total estimated revenue corresponding to all temporary contents cached by the local server, represents the estimated value of revenue corresponding to content f, ={k}, represents the set of all servers, k represents the local server, N represents the number of content f, N>1, represents the storage status of content f in the local server k, Represents the storage space of the local server k.

5. The method according to claim 1, wherein The neighboring servers meet the constraints: ; in, represents the neighboring server, represents the set of all said neighboring servers, represents the local server k and the neighboring server The straight-line distance between the servers is r, and r represents the effective radio transmission range of the base station corresponding to the server.

6. An edge caching device based on a multi-agent reinforcement learning model, characterized in that: Applicable to local servers, including: an acquisition module, configured to acquire information of a plurality of currently cached contents, the information including a content identifier, a first storage state, and a first popularity of the contents, the plurality of contents including highly popular contents having a first popularity greater than a first popularity threshold, and medium popular contents having a first popularity less than the first popularity threshold and greater than a second popularity threshold, the first popularity threshold being greater than the second popularity threshold, the first popularity indicating a probability of the contents being requested, the medium popular contents being intended to be requested by a terminal device or acquired by the neighboring server in cooperation with the neighboring server; a processing module, configured to process the content identifier, the first storage state, and the first popularity through a multi-agent reinforcement learning model to obtain a target content identifier and a target storage state of target content to be cached at a next moment; An updating module, configured to update the currently cached content according to the target content identifier, the target storage state, and the target popularity corresponding to the target content; The processing module is specifically used to: Step 1: For each content, based on the content identifier, obtain a first time delay for the neighboring server to send the content to the local server, and a second time delay for the central server to send the content to the local server, and obtain a second storage status and a second popularity of the content in the neighboring server; Step 2: Calculate an instantaneous benefit, an average benefit, and an estimated benefit corresponding to caching the content based on the first storage state, the first popularity, the second storage state, the second popularity, the first delay, and the second delay, where the instantaneous benefit represents a reduction in delay corresponding to the content. Step 3: determining a temporary content identifier and a temporary storage state of the temporary content to be cached at the next moment based on the instantaneous revenue, the average revenue, and the estimated revenue value, and determining a temporary popularity corresponding to the temporary content based on the temporary content identifier; Step 4: Repeat steps 1-3 for the temporary content until a target content identifier and a target storage state of the target content that meets preset conditions are obtained.

7. An electronic device comprising: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 5 when executed by a processor.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Edge computing and caching method and system in heterogeneous IoT network

    CN112689296A

  • Edge computing intelligent caching method based on deep reinforcement learning

    CN113687960A