Content caching methods, apparatus, devices and media
By dividing the edge caching environment into sub-environments with short time scales and using meta-reinforcement learning to adjust the caching strategy, the problem of low caching accuracy is solved, achieving fast and accurate content response and reduced latency, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing content caching solutions have low accuracy and cannot effectively cope with the high time-varying nature of user-requested content, resulting in prolonged user retrieval time and inaccurate content retrieval.
By dividing the long-term mobile edge caching environment into several sub-environments with shorter time scales, defining the state space, action space, and reward rules, and using meta-reinforcement learning methods, the base station caching strategy is adjusted to improve the accuracy of content caching and achieve rapid response to user requests.
It improves the accuracy of user content caching at the edge, reduces the latency of user content retrieval, and enhances the user experience.
Smart Images

Figure CN119629682B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mobile edge caching, and more particularly to a content caching method, apparatus, device, and medium. Background Technology
[0002] With the development of mobile internet, edge caching technology has become a key means to improve user experience quality. By placing servers with storage capabilities at the network edge to cache frequently requested content, the content delivery pressure on base stations is reduced, and the latency for users to retrieve requests is decreased.
[0003] In current content caching solutions, the cache is located at edge base stations or D2D user terminals. These caching devices have limited storage capacity and cannot cache all possible content. They can only select content that is likely to be requested by users in the future for caching.
[0004] However, existing content caching solutions suffer from low accuracy in caching content. Summary of the Invention
[0005] This application provides a content caching method, apparatus, device, and medium to address the problem of low accuracy of cached content in existing content caching schemes.
[0006] Firstly, this application provides a content caching method, the method comprising:
[0007] Define the preset state space, identifier space, initial policy, and reward rules. The state space includes historical user content. The identifier space is used to store the storage identifier of historical user content in the storage space. The initial policy is used to determine the storage identifier corresponding to the next state. The reward rules are used to determine the reward value of the state.
[0008] Based on the initial strategy, the corresponding storage identifier is determined from the identifier space, and the new state is determined based on the storage identifier and the state space;
[0009] Based on the new state and reward rules, determine the total reward value of the initial strategy within the first preset time period;
[0010] Determine the target reward expectation, and adjust the parameters of the initial policy based on the total reward value and the target reward expectation to obtain the updated policy;
[0011] Based on the update strategy and reward rules, determine the total update reward value within the second preset time period, and compare the total update reward value with the preset reward threshold range to obtain the comparison result;
[0012] Based on the comparison results, the parameters of the update strategy are adjusted to obtain the target strategy, and the target user content is determined based on the target strategy.
[0013] Cache content for the target user.
[0014] In this embodiment, according to the initial strategy, the corresponding storage identifier is determined from the identifier space, and a new state is determined based on the storage identifier and the state space, including:
[0015] Determine the historical user content corresponding to the storage identifier, and determine the corresponding state of the historical user content in the state space;
[0016] The corresponding state is designated as the new state.
[0017] In this embodiment, the total reward value of the initial strategy within a first preset time period is determined according to the new state and reward rules, including:
[0018] A first preset time is determined, and within the first preset time, the steps of determining the target storage identifier from the identifier space according to the initial strategy and determining the new state according to the target storage identifier and the state space are repeated to obtain all the new states within the first preset time.
[0019] Based on the reward rules, determine the reward value corresponding to all new states, and then determine the total reward value based on the reward value.
[0020] In this embodiment, the target reward expectation is determined, and the parameters of the initial policy are adjusted based on the total reward value and the target reward expectation to obtain an updated policy, including:
[0021] Based on the initial policy and the identifier space, the states in the state space are traversed to obtain all states, and the reward value corresponding to each state is determined according to the reward rule.
[0022] Based on all reward values, determine the total reward value, and then determine the expected value corresponding to the total reward value;
[0023] Determine the expected reward based on the expected value;
[0024] Based on the total reward value and the expected target reward, the parameters of the initial strategy are adjusted to obtain the updated strategy.
[0025] In this embodiment, the parameters of the initial policy are adjusted based on the total reward value and the expected target reward to obtain an updated policy, including:
[0026] Determine the loss function for the total reward value and the expected target reward, and determine the loss value corresponding to the loss function;
[0027] Determine the gradient value of the loss function, and adjust the parameters of the initial policy based on the gradient value to obtain the adjusted policy parameters;
[0028] Determine the update strategy based on the adjusted strategy parameters.
[0029] In this embodiment, the parameters of the update strategy are adjusted based on the comparison results to obtain the target strategy, and the target user content is determined based on the target strategy, including:
[0030] Determine the comparison results between the updated total reward value and the reward threshold range;
[0031] If the comparison result shows that the total updated reward value does not exceed the reward threshold range, then the parameters of the update strategy will not be adjusted, and the update strategy will be determined as the target strategy.
[0032] If the comparison result shows that the total updated reward exceeds the reward threshold, then the loss value between the total updated reward and the expected target reward is determined, and the gradient value of the loss value is determined.
[0033] Based on the gradient value, the parameters of the update strategy are adjusted to obtain the adjusted update strategy. Based on the adjusted update strategy and the reward rule, the total update reward value for the next second preset time period is determined.
[0034] Repeat the process of determining the comparison results of the total update reward value and the reward threshold range until the parameters of the update strategy are adjusted according to the gradient value to obtain the adjusted update strategy. Then, based on the adjusted update strategy and the reward rule, determine the total update reward value for the next second preset time period until the adjusted update strategy is determined as the target strategy.
[0035] Based on the target strategy, identify the target user content.
[0036] In this embodiment, determining the target user content according to the target strategy includes:
[0037] Determine the target strategy, identify the target storage identifier from the identifier space, and determine the target state corresponding to the target storage identifier;
[0038] Determine the historical user content included in the target state, and determine the target user content based on the historical user content.
[0039] Secondly, this application provides a content caching device, the device comprising:
[0040] The preset determination module is used to determine the preset state space, identifier space, initial strategy, and reward rules. The state space includes historical user content, the identifier space is used to store the storage identifier of historical user content in the storage space, the initial strategy is used to determine the storage identifier corresponding to the next state, and the reward rules are used to determine the reward value of the state.
[0041] The state determination module is used to determine the corresponding storage identifier from the identifier space according to the initial strategy, and to determine the new state based on the storage identifier and the state space.
[0042] The total value determination module is used to determine the total reward value of the initial strategy within the first preset time period based on the new state and reward rules.
[0043] The strategy determination module is used to determine the target reward expectation and adjust the parameters of the initial strategy based on the total reward value and the target reward expectation to obtain the updated strategy;
[0044] The result determination module is used to determine the total update reward value within the second preset time period based on the update strategy and reward rules, and compare the total update reward value with the preset reward threshold range to obtain the comparison result.
[0045] The content determination module is used to adjust the parameters of the update strategy based on the comparison results, obtain the target strategy, and determine the target user content based on the target strategy.
[0046] The caching module is used to cache content for the target user.
[0047] Thirdly, this application provides an apparatus, including: a processor, and a memory communicatively connected to the processor;
[0048] The memory stores the instructions that the computer executes;
[0049] The processor executes computer execution instructions stored in memory to implement the method of this application.
[0050] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method of this application.
[0051] The content caching method, apparatus, device, and medium provided in this application determine a preset state space, an identifier space, an initial strategy, and a reward rule. The state space includes historical user content, the identifier space stores the storage identifiers of historical user content, the initial strategy determines the storage identifier corresponding to the next state, and the reward rule determines the reward value for that state. Based on the initial strategy, the corresponding storage identifier is determined from the identifier space, and a new state is determined based on the storage identifier and the state space. Based on the new state and the reward rule, the total reward value of the initial strategy within a first preset time period is determined. A target reward expectation is determined, and the parameters of the initial strategy are adjusted based on the total reward value and the target reward expectation to obtain an updated strategy. Based on the updated strategy and the reward rule, the total updated reward value within a second preset time period is determined, and the total updated reward value is compared with a preset reward threshold range to obtain a comparison result. Based on the comparison result, the parameters of the updated strategy are adjusted to obtain a target strategy, and the target user content is determined based on the target strategy. Finally, the target user content is cached.
[0052] Thus, by using a mobile edge caching algorithm based on meta-reinforcement learning, the time-varying characteristics of the popularity of user-requested content are effectively addressed, improving the accuracy of user content cached at the edge. This enables a fast and accurate response when users request content, reducing the latency of content retrieval, improving content accuracy, and better meeting user experience needs. Attached Figure Description
[0053] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0054] Figure 1 A flowchart illustrating a content caching method provided in an embodiment of this application;
[0055] Figure 2 A schematic diagram illustrating the framework of a content caching method provided in an embodiment of this application;
[0056] Figure 3 A diagram illustrating the model training process of a content caching method provided in this application embodiment;
[0057] Figure 4 A model testing process diagram of a content caching method provided in this application embodiment;
[0058] Figure 5 A flowchart illustrating another content caching method provided in this application embodiment;
[0059] Figure 6 This is a schematic diagram of the structure of a content caching device provided in an embodiment of this application;
[0060] Figure 7 This is a structural block diagram of a device for performing a content caching method according to an embodiment of this application.
[0061] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed Implementation
[0062] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0063] Edge caching is a caching technology deployed at the edge of the network. It stores frequently used data and content closer to the user to improve access speed and reduce network traffic. When a user requests data or content, the edge cache can respond quickly and return it to the user without having to access a remote server or download the content from the network again. While the implementation of edge caching may vary, its basic principle is to improve access speed and reduce network traffic by storing frequently used data and content closer to the user.
[0064] A Markov Decision Process (MDP) is a mathematical model of sequential decision-making, used to simulate stochastic policies and rewards achievable by an agent in an environment where the system state exhibits Markov properties. An MDP is constructed based on a set of interacting objects: the agent and the environment. Its elements include state, action, policy, and reward. In an MDP simulation, the agent perceives the current system state, performs actions on the environment according to the policy, thereby changing the state of the environment and receiving a reward. The reward, accumulated over time, is called the payoff. Future states depend only on the current state and are independent of past states.
[0065] Current research on edge caching places the cache at edge base stations or D2D user terminals. These caching devices have limited storage capacity and cannot cache all possible content. They can only select content that is likely to be requested by users in the future for caching. Content popularity reflects the popularity of content in a certain area over a period of time. However, with the continuous enrichment and diversification of network content and the increasing mobility of users, user request preferences exhibit high time-varying characteristics. Prediction-based methods are difficult to cope with highly time-varying content changes. Although deep reinforcement learning methods do not require prediction and can directly change caching strategies through environmental feedback, in dynamically changing environments, before environmental attributes change, the agent can only interact with the environment a limited number of times. Deep reinforcement learning methods have low sample efficiency and cannot adjust strategies to adapt to environmental changes based on samples generated from a limited number of interactions.
[0066] To address the aforementioned issues, this application proposes a content caching method. By dividing a long-term mobile edge caching environment into several sub-environments with shorter time scales, the Markov decision process problem within each sub-environment is defined as a task, along with corresponding state space, action space, and reward rules. Based on the interaction between the base station and the state environment, the new state of the base station in the state space is determined, and the reward value corresponding to the new state is determined according to the reward rules. The policy parameters of the initial policy are then updated based on the reward value to obtain the target policy. This allows the next state of the base station in the state space to be determined based on the target policy, and the user content corresponding to the new state is identified. This user content is then edge-cached, enabling rapid response to user requests and improving the accuracy of edge-cached content.
[0067] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0068] Figure 1 This is a flowchart illustrating a content caching method provided in an embodiment of this application. Figure 1 As shown, this content caching method may include the following steps:
[0069] S110. Determine the preset state space, identifier space, initial policy, and reward rules. The state space includes historical user content. The identifier space is used to store the storage identifier of historical user content in the storage space. The initial policy is used to determine the storage identifier corresponding to the next state. The reward rules are used to determine the reward value of the state.
[0070] In this context, the state space can be understood as the set of all possible states; that is, the state space is a set that contains all possible states in the environment. A state can be represented as... in, It is an estimated content popularity information. This refers to the user's historical request content; that is, a state can be understood as a piece of historical request content. All historical request content acquired by a base station constitutes the state space corresponding to that base station, which can be represented as O. i The total historical user request content acquired by all base stations constitutes the global state space corresponding to all base stations, denoted as S, where S = O1 × O2 × … × O N .
[0071] The identifier space can also be understood as the action space. An identifier can be understood as an index of historical request content in the cache space, used to determine the corresponding historical request content in the space. The action space is a set that includes all actions that can be executed in a specific state. An action can be replacing the user's historical request content stored in the current edge cache with other historical request content. For example, each base station caches W pieces of content, whose indices in the cache space are {1,…,W}. If the action of each base station is to replace the historical request content in the storage space with the newly determined historical request content, then the action space of the base station can be represented as A. i =[0,W], Base station cache update action when At that time, the base station's action is to determine the corresponding historical request content based on the index in the cache space and replace the historical request content in the edge cache storage space; When the time is not specified, it indicates that no replacement is performed; the global action space consisting of all actions of all base stations can be represented as A = A1 × A2 × … × A N .
[0072] The strategy is used by the base station to determine the corresponding action from the action space based on the current state, execute the action, and obtain a new state. The initial strategy can be set by the user or randomly generated by the base station. The target strategy can be obtained by adjusting the strategy parameters of the initial strategy. The base station can then determine the new state and its corresponding target historical request content based on the target strategy, and replace the currently cached historical request content in the edge cache storage space with the target historical request content. This improves the prediction accuracy of content caching, enabling the base station to quickly and accurately provide the requested content to the user when the user requests it, thereby improving user satisfaction.
[0073] The reward rule can be understood as a reward function. Minimizing the global average user acquisition time latency of the edge network system is a common goal pursued by all base stations. Therefore, all base stations share a single reward function, which can be expressed as follows: Where NU represents the total number of users, N represents the number of base stations, and U represents the number of users per base station. Let $\frac{u}{t+1}$ represent the acquisition time delay of the $u$-th user in base station $i$ at time $t+1, i.e., the time spent from initiating a request to obtaining data. This is measured by averaging the negative values of the acquisition time delays of all users across all base stations, and is used to measure the global average user acquisition time delay. The goal of a base station is to minimize this global average user acquisition time delay; therefore, a larger value of the reward function indicates better system performance.
[0074] Based on this, by determining the preset state space, identifier space, initial strategy, and reward rules, the edge caching environment, which has a long time scale and exhibits high time-varying characteristics, is divided into several short-term and relatively stable sub-environments. The Markov decision process problem within each sub-environment is defined as a task, resulting in different tasks. This leads to the implementation of a mobile edge caching algorithm based on meta-reinforcement learning. The target historical request content to be replaced is determined based on the interaction between the base station and the preset task.
[0075] S120. Based on the initial strategy, determine the corresponding storage identifier from the identifier space, and determine the new state based on the storage identifier and the state space.
[0076] The storage identifier is the identifier of the historical request content in the storage space, which can also be understood as the content storage index, used to determine the corresponding historical request content.
[0077] The new state is the new state corresponding to the action determined from the action space according to the initial policy, after the action is executed.
[0078] Based on this, an initial strategy is used to determine the action corresponding to the current state from the preset action space, that is, to determine the corresponding new storage identifier. The action is then executed to obtain a new state corresponding to the new storage identifier. Subsequently, the reward value corresponding to the new state is determined according to the preset reward rules. Based on the reward value, the strategy parameters of the initial strategy are adjusted to determine the target strategy.
[0079] S130. Based on the new state and reward rules, determine the total reward value of the initial strategy within the first preset time period.
[0080] The first preset time is a pre-set time period, which can be understood as the support phase of the task.
[0081] The total reward value is the total reward value determined by determining and executing actions from the action space through an initial strategy within the first preset time period, thereby continuously changing the current state, and determining the reward value corresponding to each new state according to the reward rules, and thus determining the total reward value based on the reward values corresponding to all new states.
[0082] Based on this, by continuously determining and executing actions according to the initial strategy within the first preset time period, the state is continuously changed, and the total reward value of the new state is determined according to the reward rules, so that the initial strategy can be adjusted according to the total reward value to obtain the target strategy.
[0083] S140. Determine the target reward expectation, and adjust the parameters of the initial policy according to the total reward value and the target reward expectation to obtain the updated policy.
[0084] Here, the target reward expectation is a predetermined reward expectation used to adjust the strategy parameters, which can also be understood as the target reward value.
[0085] Based on this, by determining the target reward expectation, the strategy parameters of the initial strategy are adjusted according to the total reward value of the initial strategy and the target reward expectation within the first preset time period, thus obtaining the adjusted updated strategy.
[0086] S150. Based on the update strategy and reward rules, determine the total update reward value within the second preset time period, and compare the total update reward value with the preset reward threshold range to obtain the comparison result.
[0087] The second preset time is a pre-set time period, which can be understood as the task query phase.
[0088] The total reward value is determined by updating the action space to determine and execute actions through the update strategy within the second preset time period, thereby continuously changing the current state. Based on the reward rules, the reward value corresponding to each new state is determined, and the total reward value is determined based on the reward values corresponding to all new states.
[0089] The preset reward threshold range is a pre-defined threshold range corresponding to the total reward value. By comparing the updated total reward value with the reward threshold range, it is determined whether the updated total reward value exceeds the reward threshold range, and the corresponding parameter adjustment method is determined based on the comparison result.
[0090] Based on this, by continuously determining and executing actions according to the update strategy within the second preset time period, the new state is continuously changed, and the total update reward value of the new state is determined according to the reward rules. The total update reward value is then compared with the preset reward threshold range so that the corresponding update strategy parameter adjustment method can be determined based on the comparison results.
[0091] S160. Based on the comparison results, adjust the parameters of the update strategy to obtain the target strategy, and determine the target user content based on the target strategy.
[0092] The target strategy can be understood as the final determined user history request content that needs to be stored in the edge cache storage space; the target user content is the history request content that needs to be stored in the edge cache storage space.
[0093] Based on this, the strategy parameters of the update strategy are adjusted by comparing the results to obtain the target strategy, so as to determine the target user content that needs to be stored according to the target strategy.
[0094] S170. Cache the content for the target user.
[0095] Based on this, by storing the target user content in the edge cache storage space, the accuracy of predicting the content that the user will request is improved. When the user requests the content of the previous request, the content stored in the edge cache storage space can be used to accurately meet the user's needs, thereby improving the speed and accuracy of the user response.
[0096] Based on the feasible implementation of S120 described above, this application further provides a process for determining the corresponding new state based on the storage identifier determined according to the initial strategy:
[0097] Determine the historical user content corresponding to the storage identifier, and determine the corresponding state of the historical user content in the state space;
[0098] The corresponding state is designated as the new state.
[0099] Based on this, since the state in the state space includes historical request content, and the storage identifier represents the storage index of the historical request content in the storage space, the corresponding historical request content can be determined through the storage identifier, and then the corresponding state can be determined, and this state can be determined as the new state obtained after the action is executed.
[0100] Based on the feasible implementation of S130 described above, this application further provides a process for determining the reward value of a new state through reward rules, and determining the corresponding total reward value based on all states determined within a first preset time period:
[0101] A first preset time is determined, and within the first preset time, the steps of determining the target storage identifier from the identifier space according to the initial strategy and determining the new state according to the target storage identifier and the state space are repeated to obtain all the new states within the first preset time.
[0102] Based on the reward rules, determine the reward value corresponding to all new states, and then determine the total reward value based on the reward value.
[0103] Based on this, by continuously determining and executing actions according to the initial strategy within a first preset time period, corresponding new states are obtained, and the reward value corresponding to each new state is determined according to the reward rules, thereby determining the total reward value.
[0104] Based on the feasible implementation of S140 described above, this application further provides a process for determining the reward value corresponding to all states by traversing all states in the state space, and further determining the expected target reward:
[0105] Based on the initial policy and the identifier space, the states in the state space are traversed to obtain all states, and the reward value corresponding to each state is determined according to the reward rule.
[0106] Based on all reward values, determine the total reward value, and then determine the expected value corresponding to the total reward value;
[0107] Determine the expected reward based on the expected value;
[0108] Based on the total reward value and the expected target reward, the parameters of the initial strategy are adjusted to obtain the updated strategy.
[0109] The traversal process involves continuously determining the stored identifiers and their corresponding historical request content from the identifier space according to the initial strategy, thereby determining the new state corresponding to the historical request content, and continuously realizing the state transition until all states in the state space are determined.
[0110] The expected value is the mathematical expectation, which is the sum of the probabilities of each possible outcome in the experiment multiplied by the result. It is one of the most basic mathematical characteristics, reflecting the average value of a random variable. By determining the expected value corresponding to the total reward obtained by traversing all states as the target reward expectation, the policy parameters can be adjusted subsequently based on the total reward value and the target reward expectation.
[0111] Based on this, by traversing all states in the state space, the reward value corresponding to each state is determined according to the reward rule, and the expected target reward is determined according to the expected total reward value, so as to adjust the parameters of the initial policy in the future.
[0112] Based on the feasible implementation of S140 described above, this application further provides a process for adjusting policy parameters using the gradient value of the loss function:
[0113] Determine the loss function for the total reward value and the expected target reward, and determine the loss value corresponding to the loss function;
[0114] Determine the gradient value of the loss function, and adjust the parameters of the initial policy based on the gradient value to obtain the adjusted policy parameters;
[0115] Determine the update strategy based on the adjusted strategy parameters.
[0116] The loss function is a function that maps the values of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of that random event. In applications, the loss function is often used as a learning criterion in relation to optimization problems; that is, the model is solved and evaluated by minimizing the loss function. Specifically, the smaller the loss value between the total reward and the expected target reward, the closer the total reward is to the expected target reward. For example, the loss function can be expressed as:
[0117] The gradient value is the numerical value obtained by taking the gradient of the loss value. For example, the gradient value can be...
[0118]
[0119] The adjusted strategy parameters can be expressed as
[0120] Based on this, by determining the loss value between the total reward and the expected target reward, and calculating the gradient of the loss value, the adjusted policy parameters are determined according to the policy parameters and the gradient value, thus obtaining the updated policy.
[0121] Based on the feasible implementation of S160 described above, this application further provides a process for determining the corresponding target strategy by determining whether the total updated reward value is within the reward threshold range:
[0122] Determine the comparison results between the updated total reward value and the reward threshold range;
[0123] If the comparison result shows that the total updated reward value does not exceed the reward threshold range, then the parameters of the update strategy will not be adjusted, and the update strategy will be determined as the target strategy.
[0124] If the comparison result shows that the total updated reward exceeds the reward threshold, then the loss value between the total updated reward and the expected target reward is determined, and the gradient value of the loss value is determined.
[0125] Based on the gradient value, the parameters of the update strategy are adjusted to obtain the adjusted update strategy. Based on the adjusted update strategy and the reward rule, the total update reward value for the next second preset time period is determined.
[0126] Repeat the process of determining the comparison results of the total update reward value and the reward threshold range until the parameters of the update strategy are adjusted according to the gradient value to obtain the adjusted update strategy. Then, based on the adjusted update strategy and the reward rule, determine the total update reward value for the next second preset time period until the adjusted update strategy is determined as the target strategy.
[0127] Based on the target strategy, identify the target user content.
[0128] Wherein, the loss value is determined based on a loss function derived from the total updated reward and the expected target reward. This loss function can be expressed as: The gradient value is the numerical value obtained by taking the gradient of the loss value.
[0129] The adjusted update strategy can be represented as
[0130] Based on this, by comparing the total updated reward value with the reward threshold range, it is determined whether the total updated reward value is within the reward threshold range. If it is within the range, it indicates that the policy parameters of the updated policy meet the preset parameter conditions, and the updated policy can be used as the target policy. If it is not within the range, it indicates that the policy parameters of the updated policy need to be adjusted again within a second preset time. That is, by continuously determining the gradient value corresponding to the loss value of the total updated reward value and the expected target reward, and adjusting the policy parameters of the updated policy according to the gradient value, until an updated policy with a total updated reward value within the reward threshold range is determined and is determined as the target policy.
[0131] Based on the feasible implementation of S160 described above, this application further provides a process for determining a target storage identifier through a target strategy, and then determining the corresponding target user content based on the target storage identifier:
[0132] Determine the target strategy, identify the target storage identifier from the identifier space, and determine the target state corresponding to the target storage identifier;
[0133] Determine the historical user content included in the target state, and determine the target user content based on the historical user content.
[0134] Based on this, by determining the target strategy, the target storage identifier is determined from the identifier space according to the target strategy, thereby determining the state corresponding to the target storage identifier and setting this state as the new state. The state transition is realized according to the target strategy, and the content stored in the edge cache storage space is replaced according to the historical user content corresponding to the new state, so as to achieve accurate prediction of user request content.
[0135] Please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the framework of a content caching method provided in an embodiment of this application, such as... Figure 2 As shown, meta-reinforcement learning can be used to determine the moving edge cache in a non-stationary environment. Here, S is the first preset time of the task, i.e. the support phase, and Q is the second preset time, i.e. the query phase. S and the Q connected after it constitute the task time of a task.
[0136] Please refer to Figure 3 , Figure 3 A diagram illustrating the model training process of a content caching method provided in this application embodiment is shown below. Figure 3As shown, in the meta-training stage, M tasks T_i (i = 1, 2, ..., M) are randomly selected. In the support stage of each task, the initial strategy and the environment interact to generate a set of K trajectories. The proxy loss function of the support stage is calculated. Then, the initial strategy is fine-tuned based on the VPG learner, as shown in Equation (1). The updated strategy calculates the loss in the query stage of the task. The losses of the query stages of the M tasks are summed. The initial strategy is updated based on the PPO meta-learner. The above operation is repeated to continuously update the initial strategy of multiple tasks until the cumulative profit of the strategy on the query data tends to stabilize after fine-tuning on all tasks. Training is stopped, and the optimized initial trading strategy is used in the testing stage.
[0137] Please refer to Figure 4 , Figure 4 A model testing process diagram for a content caching method provided in this application embodiment is shown below. Figure 4 As shown, the trained meta-reinforcement learning mobile edge caching model can be used for testing in the meta-testing phase.
[0138] In this embodiment, to address the issue of highly time-varying user request preferences and the difficulty of predictive methods in handling such content changes, the long-term mobile edge caching environment can be divided into several sub-environments with shorter time scales. The Markov decision process problem within each sub-environment is defined as a task. By pre-setting a state space, an identifier space (action space), an initial policy, and reward rules, actions are determined and executed according to the initial policy to obtain a new state. The reward value corresponding to the new state is determined according to the reward rules, thus determining the total reward value for all states within a first preset time period. By traversing all states, the target reward expectation corresponding to the total reward value is determined, thereby determining the gradient value corresponding to the loss value between the total reward value and the target reward expectation. The policy parameters are adjusted according to the gradient value to determine the update policy. By comparing the updated total reward value corresponding to the update policy with a preset reward threshold range, the parameters of the update policy are adjusted based on the comparison results to determine the target policy.
[0139] In this way, by determining the target strategy, the target storage identifier is determined from the identifier space according to the target strategy, and the corresponding status and its historical request content are determined according to the target storage identifier. Thus, the historical request content is determined as the target historical request content and stored in the edge cache storage space. This improves the accuracy of predicting the content that the user will request. When the user requests the historical request content, the user's needs can be accurately met according to the content pre-stored in the edge cache storage space, thereby improving the speed and accuracy of the user response.
[0140] Figure 5This is a flowchart illustrating another content caching method provided in an embodiment of this application. Figure 5 As shown, this content caching method may include the following steps:
[0141] S510. Divide the long-term mobile edge caching environment into sub-environments with shorter time scales. Due to the short time scale of each sub-environment, the Markov decision process problem within each sub-environment is defined as a task.
[0142] Based on this, different tasks are obtained, and each task is divided into a support phase and a query phase according to the time sequence.
[0143] S520. Construct a meta-reinforcement learning model, which consists of a learner and a meta-learner.
[0144] The meta-reinforcement learning model can be understood as an agent in the Markov decision process, which uses the PPO meta-learner to learn the common initial policy for multiple tasks.
[0145] The learner is the VPG (Vanilla Policy Gradients) learner, and the meta-learner is the PPO (Proximal Policy Optimization) meta-learner.
[0146] S530. The trained meta-learner is used in the support phase of the testing phase to optimize and update the learner parameters of the initialized model, and then the performance of the updated learner parameters is evaluated on the query set of the test set.
[0147] Based on this, the optimization update is to determine the total reward value and the loss value of the expected target reward by identifying the multiple states obtained by the model in the support phase through interaction with the environment and their corresponding reward values, and then calculate the gradient, so as to fine-tune the initial policy based on the support data of each task.
[0148] Performance evaluation involves comparing the total reward value of the fine-tuned strategy during the query phase with a preset reward range to determine whether the total reward value of the fine-tuned strategy is stable. In this embodiment, to accommodate the time-varying nature of content popularity, a mobile edge caching model based on meta-reinforcement learning can be constructed. This model consists of a VPG learner and a PPO meta-learner. During the training phase, multiple tasks are randomly selected. For each task, the total reward value and the expected loss value of the target reward are determined by identifying multiple states obtained by the model in the support phase through interaction with the environment and their corresponding reward values. The gradient is then calculated, and the initial strategy is fine-tuned using the support data for each task. The fine-tuned strategy is then used to calculate the loss on the query data of the task. Finally, the average proxy loss is calculated on the query data, thus updating the initial strategy parameters based on the PPO meta-learner to obtain the target strategy.
[0149] Figure 6 This is a schematic diagram of the structure of a content caching device 600 provided in an embodiment of this application, as shown below. Figure 6 As shown, the content caching device 600 includes: a preset determination module 610, a status determination module 620, a total value determination module 630, a strategy determination module 640, a result determination module 650, a content determination module 660, and a caching module 670.
[0150] The preset determination module 610 is used to determine the preset state space, identifier space, initial strategy, and reward rules. The state space includes historical user content, the identifier space is used to store the storage identifier of historical user content in the storage space, the initial strategy is used to determine the storage identifier corresponding to the next state, and the reward rules are used to determine the reward value of the state.
[0151] The state determination module 620 is used to determine the corresponding storage identifier from the identifier space according to the initial strategy, and to determine the new state according to the storage identifier and the state space;
[0152] The total value determination module 630 is used to determine the total reward value of the initial strategy within the first preset time period based on the new state and reward rules.
[0153] The strategy determination module 640 is used to determine the target reward expectation and adjust the parameters of the initial strategy according to the total reward value and the target reward expectation to obtain the updated strategy;
[0154] The result determination module 650 is used to determine the total update reward value within the second preset time period according to the update strategy and reward rules, and compare the total update reward value with the preset reward threshold range to obtain the comparison result.
[0155] The content determination module 660 is used to adjust the parameters of the update strategy based on the comparison results, obtain the target strategy, and determine the target user content based on the target strategy.
[0156] The caching module 670 is used to cache content for the target user.
[0157] In this embodiment of the application, the state determination module 620 can also be specifically used for:
[0158] Determine the historical user content corresponding to the storage identifier, and determine the corresponding state of the historical user content in the state space;
[0159] The corresponding state is designated as the new state.
[0160] In this embodiment of the application, the total value determination module 620 can also be specifically used for:
[0161] A first preset time is determined, and within the first preset time, the steps of determining the target storage identifier from the identifier space according to the initial strategy and determining the new state according to the target storage identifier and the state space are repeated to obtain all the new states within the first preset time.
[0162] Based on the reward rules, determine the reward value corresponding to all new states, and then determine the total reward value based on the reward value.
[0163] In this embodiment of the application, the strategy determination module 640 can also be specifically used for:
[0164] Based on the initial policy and the identifier space, the states in the state space are traversed to obtain all states, and the reward value corresponding to each state is determined according to the reward rule.
[0165] Based on all reward values, determine the total reward value, and then determine the expected value corresponding to the total reward value;
[0166] Determine the expected reward based on the expected value;
[0167] Based on the total reward value and the expected target reward, the parameters of the initial strategy are adjusted to obtain the updated strategy.
[0168] In this embodiment of the application, the strategy determination module 640 can also be specifically used for:
[0169] Determine the loss function for the total reward value and the expected target reward, and determine the loss value corresponding to the loss function;
[0170] Determine the gradient value of the loss function, and adjust the parameters of the initial policy based on the gradient value to obtain the adjusted policy parameters;
[0171] Determine the update strategy based on the adjusted strategy parameters.
[0172] In this embodiment of the application, the content determination module 660 can also be specifically used for:
[0173] Determine the comparison results between the updated total reward value and the reward threshold range;
[0174] If the comparison result shows that the total updated reward value does not exceed the reward threshold range, then the parameters of the update strategy will not be adjusted, and the update strategy will be determined as the target strategy.
[0175] If the comparison result shows that the total updated reward exceeds the reward threshold, then the loss value between the total updated reward and the expected target reward is determined, and the gradient value of the loss value is determined.
[0176] Based on the gradient value, the parameters of the update strategy are adjusted to obtain the adjusted update strategy. Based on the adjusted update strategy and the reward rule, the total update reward value for the next second preset time period is determined.
[0177] Repeat the process of determining the comparison results of the total update reward value and the reward threshold range until the parameters of the update strategy are adjusted according to the gradient value to obtain the adjusted update strategy. Then, based on the adjusted update strategy and the reward rule, determine the total update reward value for the next second preset time period until the adjusted update strategy is determined as the target strategy.
[0178] Based on the target strategy, identify the target user content.
[0179] In this embodiment of the application, the content determination module 660 can also be specifically used for:
[0180] Determine the target strategy, identify the target storage identifier from the identifier space, and determine the target state corresponding to the target storage identifier;
[0181] Determine the historical user content included in the target state, and determine the target user content based on the historical user content.
[0182] Figure 7 This is a schematic diagram of the device provided in an embodiment of this application. Figure 7 As shown, the device 700 includes:
[0183] The device 700 may include a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a communication component 703, and other components. The processor 701, memory 702, and communication component 703 are connected via a bus 704.
[0184] In the specific implementation process, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to execute the message processing method described above.
[0185] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0186] In the above Figure 7In the illustrated embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0187] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0188] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0189] In some embodiments, a computer program product is also provided, comprising a computer program or instructions that, when executed by a processor, implement the steps in any of the above-described content caching methods.
[0190] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0191] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0192] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of computer-executable instructions, which can be loaded by a processor to execute the steps in any of the content caching methods provided in embodiments of this application.
[0193] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0194] According to one aspect of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium.
[0195] Since the instructions stored in the storage medium can execute the steps in any of the content caching methods provided in the embodiments of this application, the beneficial effects that any of the content caching methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0196] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0197] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A content caching method, characterized in that, The method includes: A preset state space, identifier space, initial strategy, and reward rule are determined. The state space includes historical user content. The identifier space is used to store the storage identifier of the historical user content in the storage space. The initial strategy is used to determine the storage identifier corresponding to the next state. The reward rule is used to determine the reward value of the state. According to the initial strategy, the corresponding storage identifier is determined from the identifier space, and a new state is determined based on the storage identifier and the state space; Based on the new state and the reward rules, determine the total reward value of the initial strategy within a first preset time period; Determine the target reward expectation, and adjust the parameters of the initial strategy based on the total reward value and the target reward expectation to obtain an updated strategy; Based on the update strategy and the reward rules, the total update reward value within the second preset time period is determined, and the total update reward value is compared with the preset reward threshold range to obtain the comparison result; Based on the comparison results, the parameters of the update strategy are adjusted to obtain a target strategy. Then, based on the target strategy, the target user content is determined, including: determining the comparison results of the total update reward value and the reward threshold range; if the comparison result indicates that the total update reward value does not exceed the reward threshold range, the parameters of the update strategy are not adjusted, and the update strategy is determined as the target strategy; if the comparison result indicates that the total update reward value exceeds the reward threshold range, the loss value between the total update reward value and the expected target reward is determined, and the gradient value of the loss value is determined; based on the gradient value, the update strategy is adjusted... The parameters of the update strategy are adjusted to obtain the adjusted update strategy. Based on the adjusted update strategy and the reward rule, the total update reward value for the next second preset time period is determined. The process of determining the comparison result of the total update reward value and the reward threshold range is repeated until the steps of adjusting the parameters of the update strategy based on the gradient value to obtain the adjusted update strategy and determining the total update reward value for the next second preset time period are completed, until the adjusted update strategy is determined as the target strategy. Based on the target strategy, the target user content is determined. The content for the target user is cached.
2. The method according to claim 1, characterized in that, The step of determining the corresponding storage identifier from the identifier space according to the initial strategy, and determining the new state according to the storage identifier and the state space, includes: Determine the historical user content corresponding to the storage identifier, and determine the corresponding state of the historical user content in the state space; The corresponding state is determined as the new state.
3. The method according to claim 1, characterized in that, The step of determining the total reward value of the initial strategy within a first preset time period based on the new state and the reward rule includes: Determine the first preset time, and within the first preset time, repeatedly execute the steps of determining the target storage identifier from the identifier space according to the initial strategy, and determining the new state according to the target storage identifier and the state space, to obtain all the new states within the first preset time. Based on the reward rules, determine the reward value corresponding to all the new states, and based on the reward value, determine the total reward value.
4. The method according to claim 1, characterized in that, The step of determining the target reward expectation and adjusting the parameters of the initial strategy based on the total reward value and the target reward expectation to obtain an updated strategy includes: Based on the initial strategy and the identifier space, the states in the state space are traversed to obtain all the states, and the reward value corresponding to all the states is determined according to the reward rule. Based on all reward values, determine the total reward value, and then determine the expected value corresponding to the total reward value; Based on the expected value, determine the expected target reward; Based on the total reward value and the expected target reward, the parameters of the initial strategy are adjusted to obtain the updated strategy.
5. The method according to claim 4, characterized in that, The step of adjusting the parameters of the initial strategy based on the total reward value and the expected target reward to obtain an updated strategy includes: Determine the loss function for the total reward value and the expected target reward, and determine the loss value corresponding to the loss function; Determine the gradient value of the loss function, and adjust the parameters of the initial policy according to the gradient value to obtain the adjusted policy parameters; The update strategy is determined based on the adjusted strategy parameters.
6. The method according to claim 1, characterized in that, The step of determining the target user content according to the target strategy includes: Determine the target strategy, identify the target storage identifier from the identifier space, and determine the target state corresponding to the target storage identifier; The target state includes historical user content, and the target user content is determined based on the historical user content.
7. A content caching device, characterized in that, The device includes: The preset determination module is used to determine a preset state space, an identifier space, an initial strategy, and a reward rule. The state space includes historical user content, the identifier space is used to store the storage identifier of the historical user content in the storage space, the initial strategy is used to determine the storage identifier corresponding to the next state, and the reward rule is used to determine the reward value of the state. The state determination module is used to determine the corresponding storage identifier from the identifier space according to the initial strategy, and to determine the new state according to the storage identifier and the state space; The total value determination module is used to determine the total reward value of the initial strategy within a first preset time period based on the new state and the reward rule; The strategy determination module is used to determine the target reward expectation and adjust the parameters of the initial strategy according to the total reward value and the target reward expectation to obtain an updated strategy; The result determination module is used to determine the total update reward value within a second preset time period according to the update strategy and the reward rule, and compare the total update reward value with a preset reward threshold range to obtain a comparison result; The content determination module is used to adjust the parameters of the update strategy based on the comparison results to obtain a target strategy, and to determine the target user content based on the target strategy, including: determining the comparison results of the total update reward value and the reward threshold range; if the comparison result shows that the total update reward value does not exceed the reward threshold range, then the parameters of the update strategy are not adjusted, and the update strategy is determined as the target strategy; if the comparison result shows that the total update reward value exceeds the reward threshold range, then the loss value between the total update reward value and the expected target reward is determined, and the gradient value of the loss value is determined; based on the gradient value, the content is adjusted accordingly. The parameters of the update strategy are adjusted to obtain the adjusted update strategy. Based on the adjusted update strategy and the reward rule, the total update reward value for the next second preset time period is determined. The process of determining the comparison result between the total update reward value and the reward threshold range is repeated until the steps of adjusting the parameters of the update strategy based on the gradient value to obtain the adjusted update strategy and determining the total update reward value for the next second preset time period are completed, until the adjusted update strategy is determined as the target strategy. Based on the target strategy, the target user content is determined. The caching module is used to cache the target user content.
8. A device, characterized in that, include: One or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that can be invoked by a processor to perform the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Edge region collaborative cache updating method based on multi-agent deep reinforcement learning
CN117675918A
Conversation-depth social engineering attack detection using attributes from automated dialog engagement
US20230179628A1