Cache optimization method for cell-free multi-input multi-output environment based on VPPO algorithm

By employing the VPPO algorithm for multi-environment parallel interaction and the Retrace algorithm in a CF-MIMO environment, an Actor-Critic network is constructed, which solves the problem of the adaptability of caching strategies in dynamically changing networks and achieves efficient cache optimization and service quality assurance.

CN120957192APending Publication Date: 2025-11-14NORTHEASTERN UNIV AT QINHUANGDAO
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511083686.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional caching strategies are ill-suited to the dynamic changes in device demands and the uncertainty of network conditions in modern urban communication networks. In particular, under CF-MIMO conditions, caching strategies cannot respond to changes in user device demands in a timely manner, resulting in low content distribution efficiency and drastic latency fluctuations.

Method used

A cache optimization method based on the VPPO algorithm for cellless multi-input/output environments is adopted. By generating multiple independent environment instances to run in parallel, an Actor policy network and a Critic value network are constructed. The Retrace algorithm is used to calculate the advantage function and collaboratively update the policy network parameters to achieve dynamic adjustment of the cache ratio.

Benefits of technology

It improves the adaptability and decision-making accuracy of caching strategies in dynamically changing scenarios, reduces latency and increases cache hit rate, and ensures the stability of service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120957192A_ABST
    Figure CN120957192A_ABST
Patent Text Reader

Abstract

The invention discloses a cache optimization method for a cell-free multiple-input-output environment based on a VPPO algorithm, and relates to the technical field of edge cache optimization. Multiple independent environment instances are generated and run in parallel; training the strategy network based on the plurality of independent environment instances and the value network to obtain a trained strategy network; and inputting the cache proportion of the current edge server into the trained strategy network to obtain the adjustment of the cache proportion of the edge server. A complex dynamic scene is simulated through multi-environment parallel interaction, multi-scene data are collected in parallel, and the adaptability and decision-making ability of the model to dynamic changes are improved. According to the method, a non-Markov environment is adapted by using Retract advantage estimation, and the problem of deviation of traditional advantage estimation in such scenes is solved by truncation weight and recursive calculation and processing of the condition that state transition depends on history, so that the advantage function calculation is more accurate, a foundation is laid for strategy optimization, and the performance of a model in a complex dependency relationship scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of edge cache optimization technology, and particularly relates to a cache optimization method for cellless multi-input / output environments based on the VPPO algorithm. Background Technology

[0002] In the context of the 5G digital era, the explosive growth of mobile smart terminals and internet applications has driven the rapid development of IoT services, leading to unprecedented data access pressure on modern urban communication networks. This pressure is particularly pronounced during peak business hours, highlighting the limitations of traditional centralized network architectures. Edge caching technology, by pre-positioning content at network edge nodes (i.e., the network access layer close to end users), can effectively reduce data backhaul frequency, alleviate core network congestion, and significantly improve Quality of Service (QoS). However, the intelligent decision-making mechanism for cached content remains a critical issue that urgently needs to be addressed. Notably, the unique network characteristics of cellless multiple-input multiple-output (CF-MIMO) systems deployed in modern urban communication networks present new challenges to the design of edge caching strategies. These challenges include: 1) Dynamism: the user-access point relationship changes in real time; 2) Complexity: collaborative management of large-scale distributed access points; 3) Non-Markovian characteristics: long-term dependency in system state transitions. These characteristics mean that edge caching optimization in CF-MIMO environments not only needs to address latency and bandwidth constraints in traditional networks but also must overcome new technical challenges brought about by distributed architectures, especially the problem of collaborative caching optimization among multiple access points.

[0003] Currently, the main methods for solving the above problems include caching strategies based on traditional algorithms and edge caching strategies based on deep reinforcement learning (DRL). The caching strategies based on traditional algorithms specifically include predictive caching strategies based on popularity and learning-based predictive caching strategies. The core idea of ​​predictive caching strategies based on popularity is to assume that "content that was popular in the past will remain popular in the future." It establishes a content popularity ranking by statistically analyzing historical access frequency, temporal locality, and other indicators, such as the LRU (Least Recently Used) method. The core idea of ​​learning-based predictive caching strategies is to use machine learning models to mine complex patterns in request data and predict content popularity by learning from historical data. The core idea of ​​edge caching strategies based on deep reinforcement learning (DRL) is to model caching decisions as Markov processes, learning the optimal strategy through interaction with the environment and dynamically adjusting cached content to adapt to the constantly changing device service requests and network conditions.

[0004] However, in the context of modern urban communication networks, traditional methods struggle to adapt to the dynamic changes in device demands and the uncertainty of network conditions. Specifically, while popularity-based predictive caching strategies can utilize content popularity information for caching decisions to some extent, they struggle to adjust cached content promptly when faced with rapid changes in popularity, leading to a decrease in cache hit rate. Learning-based predictive caching strategies predict content popularity by learning from historical data; however, due to the diversity of device preferences and real-time requirements, such predictions are often inaccurate, affecting the effectiveness of caching decisions and limiting their ability to guarantee Quality of Service (QoS). This strategy also ignores the diversity and real-time requirements of data services, resulting in less than ideal QoS in practical applications. Because modern urban communication network systems possess high-dimensional state spaces, complex decision-making processes, and non-Markovian characteristics, DRL algorithms are prone to problems such as unstable latency, insufficient data diversity, and unstable exploration strategies during training. DRL-based edge caching strategies lack adaptability to rapidly changing network conditions and device behavior, causing caching strategies to fail to respond promptly to changes in user device demands, leading to low content distribution efficiency and drastic latency fluctuations. When dealing with certain CF-MIMO environments with non-Markovian characteristics, traditional DRL-based algorithms often fail to accurately capture state transition features that rely on historical information, leading to biased estimation of the dominance function. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a cache optimization method for cellless multi-input / output environments based on the VPPO algorithm, in order to solve the problem that traditional methods are difficult to adapt to the dynamic changes in device requirements and the uncertainty of network conditions.

[0006] The technical solution of this invention is as follows:

[0007] On the one hand, this invention provides a cache optimization method for cellless multiple-input-output environments based on the VPPO algorithm, including the following specific steps:

[0008] Multiple independent environment instances are generated. Each environment instance simulates the interaction scenarios of different edge servers and smart devices. The environment instances are completely isolated from each other and run in parallel. Each environment instance includes a smart agent, which corresponds to the edge server in the interaction scenario.

[0009] Construct an Actor policy network and a Critic value network. Train the Actor policy network based on multiple independent environment instances and the Critic value network to obtain the trained Actor policy network.

[0010] Input the current cache ratio of the edge server into the trained Actor policy network to obtain the adjustment of the cache ratio of the edge server.

[0011] Furthermore, the step of training the Actor policy network based on multiple independent environment instances and the Critic value network to obtain the trained Actor policy network specifically includes:

[0012] S1: In each environment instance In this process, the agent independently generates a state. ,in Indicates the number of the environment instance, and , For the number of environment instances, For a specific moment;

[0013] The state Indicates in The current cache ratio of the edge server corresponding to the intelligent agent at any given time, wherein the cache ratio is the ratio of cache requests faced by the edge server to the currently stored content;

[0014] S2: Utilize an Actor policy network to simultaneously process the states generated by agents in all environment instances. Output an action for each agent. ;

[0015] Specifically, the Actor policy network first considers the states generated by the agents in all environment instances. Generate a Gaussian policy distribution as the probability distribution of the agent's next action, and then generate the agent's action based on the probability distribution of the next action. The actions of the intelligent agent This indicates an adjustment to the cache ratio of the edge server corresponding to the intelligent agent;

[0016] S3: Agents in each environment instance receive actions distributed by the Actor policy network. Then, it executes the calculation of rewards, and takes the average of the rewards obtained by each environment instance to obtain the final reward. Each environment instance generates the next state. All trajectory data generated by all agents in their respective environment instances are collected into the transition information pool; the trajectory data includes states. ,action ,award and the next state ;

[0017] S4: The Critic value network calculates the evaluation state value function based on the trajectory data in the transition information pool. The parameters of the Critic value network are updated using the TD-error method.

[0018] S5: Update the Actor policy network parameters based on the trajectory data in the transition information pool and the Critic value network;

[0019] S6: Determine whether the current iteration count has reached the set number or whether the Actor policy network has converged. If so, obtain the trained Actor policy network; otherwise, return to S2.

[0020] Furthermore, the Gaussian policy distribution is as follows:

[0021] (1);

[0022] in, In the state Take action below The probability density, For the Actor policy network, based on the state The mean of the output Gaussian distribution. For the Actor policy network, based on the state The standard deviation of the output Gaussian distribution.

[0023] Furthermore, the reward The calculation method is as follows:

[0024] (2);

[0025] (3);

[0026] (4);

[0027] (5);

[0028] in, For service quality, For the service number, For the total number of services that need to be satisfied, The number of services that can be served at present. express Time of the first Total demand for each service express Time of the first Service priority For the first The delay of the service Due to pre-release delay, For subsequent transmission delay, This refers to the latency parameters during the transmission process from the edge server to the smart device. Store the first for edge servers The amount of content for each service This refers to the latency parameters for edge servers downloading from cloud servers. This is a balancing parameter used to dynamically adjust latency and service fulfillment rate.

[0029] Furthermore, the step of updating the Actor policy network parameters based on the trajectory data in the transition information pool and the Critic value network specifically includes:

[0030] S5.1: Calculate the cutoff weight using the Retrace algorithm Then calculate the advantage function. ;

[0031] The truncation weight for:

[0032] (6);

[0033] in, In the state Take action below The probability density, i.e., the current action strategy. For the target probability strategy;

[0034] Advantage function The calculation formula is:

[0035] (7);

[0036] in, For hyperparameters, for Discount factor of time, for Dominance function estimation at time 1;

[0037] S5.2: Based on the dominant function The probability ratio of the new and old strategies is calculated and corrected using the Clip pruning mechanism to obtain the corrected probability ratio of the new and old strategies.

[0038] The revised probability ratio of the old and new strategies The calculation formula is:

[0039] (8);

[0040] in, This represents the ratio of the probability of the new strategy to the probability of the old strategy after modification. Expressing expectations, The importance sampling ratio represents the probability of using the new and old strategies. The clipping function represents the process of clipping the data. Limited to and between, For control strategy hyperparameters;

[0041] S5.3: Update the parameters of the Actor policy network based on the corrected probability ratio of the old and new policies;

[0042] (9);

[0043] in, For the parameters of the updated Actor policy network, The parameters of the Actor policy network before the update. The learning rate of the Actor policy network. Indicates to Perform gradient calculations.

[0044] Secondly, this application proposes an electronic device, comprising: one or more processors, and a memory for storing instructions, which, when executed by the one or more processors, cause the one or more processors to perform the cache optimization method for cellless multiple input / output environments based on the VPPO algorithm.

[0045] Thirdly, this application proposes a computer-readable storage medium storing executable instructions that, when executed, cause a processor to perform the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm.

[0046] Fourthly, this application proposes a computer program product, including a computer program or instructions that, when executed by a processor, implement the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] 1. Employing multi-environment parallel interaction, SubprocVecEnv is used to build independent isolated environment instances to simulate complex dynamic scenarios and collect multi-scenario data in parallel. This solves the problems of single-environment data being singular and inefficient, allowing the model to be exposed to richer scenarios, efficiently utilizing multi-scenario data, improving adaptability and decision-making ability to dynamic changes (such as sudden changes in content popularity), and accurately optimizing decision-making strategies.

[0049] 2. The policy network and value network work together, responsible for action decision-making and state value assessment, respectively. During each training round, data is first collected through parallel interaction across multiple environments. After calculating the advantage function using Retrace, the policy and value network parameters are updated synchronously. The updated policy is then distributed to each environment instance, iterating until metrics such as QoS and hit rate converge. This ensures the model stably outputs optimal decisions in complex dynamic scenarios (such as sudden changes in content popularity or smart device mobility) (e.g., edge server caching policy adjustments). Retrace advantage estimation is adapted to non-Markovian environments. By truncating weights and recursive calculation, it handles situations where state transitions depend on history, addressing the bias issues of traditional advantage estimation in such scenarios. This makes the advantage function calculation more accurate, laying the foundation for policy optimization and improving model performance in scenarios with complex dependencies. Attached Figure Description

[0050] Figure 1 This is a flowchart of VPPO model training in an embodiment of the present invention;

[0051] Figure 2 This is a flowchart of the model in an embodiment of the present invention;

[0052] Figure 3 This is a diagram showing the model runtime in an embodiment of the present invention;

[0053] Figure 4 This is a comparison of model delays in an embodiment of the present invention;

[0054] Figure 5 This is a diagram showing the hit rate effect in an embodiment of the present invention;

[0055] Where (a) is the hit rate with 20 services; and (b) is the hit rate with 40 services.

[0056] Figure 6 This is a QoS effect diagram in an embodiment of the present invention;

[0057] (a) represents the service quality for a service quantity of 20; (b) represents the service quality for a service quantity of 40. Detailed Implementation

[0058] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0059] Example 1:

[0060] A cache optimization method for cellless multi-input / output environments based on the VPPO algorithm, such as Figure 1 and Figure 2 As shown, the specific steps include the following:

[0061] Step 1: Generate multiple independent environment instances. Each environment instance simulates the interaction scenarios of different edge servers and smart devices. The environment instances are completely isolated from each other and run in parallel. Each environment instance includes an agent, which corresponds to the edge server in the interaction scenario.

[0062] In this embodiment, the environment instance is established through SubprocVecEnv, and the smart devices include mobile phones, drones, laptops, etc.

[0063] Step 2: Construct an Actor policy network and a Critic value network. Train the Actor policy network based on multiple independent environment instances and the Critic value network to obtain the trained Actor policy network.

[0064] Step 2.1: In each environment instance In this process, the agent independently generates a state. ,in Indicates the number of the environment instance, and , For the number of environment instances, For a specific moment;

[0065] The state Indicates in The current cache ratio of the edge server corresponding to the intelligent agent at any given time, wherein the cache ratio is the ratio of cache requests faced by the edge server to the currently stored content;

[0066] Step 2.2: Use an Actor policy network to process the states generated by agents in all environment instances simultaneously. Output an action for each agent. ;

[0067] Specifically, the Actor policy network first considers the states generated by the agents in all environment instances. Generate a Gaussian policy distribution as the probability distribution of the agent's next action (i.e., when the policy parameters are...). In the case of, in state Select action The logarithmic probability of the next action is used to generate the agent's action based on the probability distribution of the next action. The actions of the intelligent agent This represents the adjustment of the cache ratio of the edge server corresponding to the agent (a specific value in the continuous action space).

[0068] The Gaussian strategy distribution is as follows:

[0069] (1);

[0070] in, In the state Take action below The probability density, For Actor policy network (parameters are) According to the status The mean of the Gaussian distribution of the output, i.e., the mean of the Actor policy network in state. The "optimal" action direction (optimal cache ratio strategy) is determined. For Actor policy network (parameters are) According to the status The standard deviation of the output Gaussian distribution is used to control the agent's exploration strategy;

[0071] Step 2.3: Agents in each environment instance receive actions distributed by the Actor policy network. Then, it executes the calculation of rewards, and takes the average of the rewards obtained by each environment instance to obtain the final reward. Each environment instance generates the next state. All trajectory data generated by all agents in their respective environment instances are collected into the transition information pool; the trajectory data includes states. ,action ,award and the next state ;

[0072] The reward is calculated according to the reward formula. :

[0073] (2);

[0074] (3);

[0075] (4);

[0076] (5);

[0077] in, For service quality, For the service number, For the total number of services that need to be satisfied, The number of services that can be served at present. express Time of the first Total demand for each service express Time of the first Service priority For the first The delay of the service Due to pre-release delay, For subsequent transmission delay, This refers to the latency parameters during the transmission process from the edge server to the smart device. Store the first for edge servers The amount of content for each service This refers to the latency parameters for edge servers downloading from cloud servers. These are balancing parameters used to dynamically adjust latency and service fulfillment rate.

[0078] Step 2.4: The Critic value network calculates the evaluation state value function based on the trajectory data in the transition information pool. The parameters of the Critic value network are updated using the TD-error method.

[0079] Step 2.5: Update the Actor policy network parameters based on the trajectory data in the transition information pool and the Critic value network;

[0080] Step 2.5.1: Calculate the truncation weights using the Retrace algorithm. Then calculate the advantage function. ;

[0081] The truncation weight for:

[0082] (6);

[0083] in, In the state Take action below The probability density, i.e., the current action strategy. For the target probability strategy;

[0084] To reduce the impact of non-Markovian characteristics in the training environment, an advantage function is adopted. The calculation formula is as follows:

[0085] (7);

[0086] in, These are hyperparameters used to balance single-step and multi-step rewards for the agent. for Discount factor of time, for Dominance function estimation at time 1;

[0087] Step 2.5.2: Based on the dominance function The probability ratio of the new and old strategies is calculated and corrected using the Clip pruning mechanism to obtain the corrected probability ratio of the new and old strategies.

[0088] The revised probability ratio of the old and new strategies The calculation formula is:

[0089] (8);

[0090] in, This represents the ratio of the probability of the new strategy to the probability of the old strategy after modification. Expressing expectations, The importance sampling ratio represents the probability of using the new and old strategies. The clipping function represents the process of clipping the data. Limited to and between, For control strategy hyperparameters;

[0091] Step 2.5.3: Update the parameters of the Actor policy network based on the corrected probability ratio of the old and new policies;

[0092] (9);

[0093] in, For the parameters of the updated Actor policy network, The parameters of the Actor policy network before the update. The learning rate of the Actor policy network. Indicates to Perform gradient calculation;

[0094] Step 2.6: Determine whether the current iteration count has reached the set number or whether the Actor policy network has converged. If so, obtain the trained Actor policy network; otherwise, return to step 2.2.

[0095] Step 3: Input the current cache ratio of the edge server into the trained Actor policy network to obtain the adjustment of the cache ratio of the edge server;

[0096] By accelerating the reinforcement learning process of agents and improving the efficiency of data diversity collection through a distributed parallel training framework, a model that can flexibly adapt to dynamic changes in the network environment is constructed. At the same time, the Retrace algorithm is introduced to improve the calculation of the advantage function, effectively solving the decision problem under the non-Markov characteristics of CF-MIMO system, and improving the generalization ability and convergence speed of the model in dynamic network environment.

[0097] like Figure 3 As shown, with the number of services With the increase of [unclear], the model runtime increases more gradually, maintaining lower time complexity compared to other DRL algorithms such as PPO, A2C, and SAC. It still maintains efficient computing even when the value is 40.

[0098] like Figure 4 As shown, in the CF-MIMO high-dynamic scenario, the model adapts to changes in device requests more quickly through asynchronous data acquisition and Gaussian policy distribution output exploration strategy. The average latency of the model is better than other DRL algorithms, and the model latency distribution is more concentrated with a lower maximum latency value.

[0099] like Figure 5 As shown, in terms of service quantity =20 and When the value is 40, the model combines a priority weighting mechanism to more accurately match content request patterns, reduce cache misses, and achieve a higher hit rate.

[0100] like Figure 6 As shown, the number of services is also... =20 and When the value is 40, the model enhances dynamic adaptability through parallel learning of multiple environment instances and optimizes policy updates by combining the CLIP pruning mechanism, which can maximize QoS in dynamic environments.

[0101] Example 2:

[0102] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the cache optimization method for cellless multiple input / output environments based on the VPPO algorithm.

[0103] The electronic device may be a mobile phone, computer, or tablet computer, etc., and includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm as described in the embodiments. It is understood that the electronic device may also include input / output (I / O) interfaces and communication components.

[0104] The processor is used to execute all or part of the steps in the cache optimization method for a cellless multiple-input-output environment based on the VPPO algorithm as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions for any application or method in the electronic device, as well as application-related data.

[0105] The processor can be implemented as an Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic components, and is used to execute the cache optimization method for cellless multi-input / output environment based on VPPO algorithm described in the above embodiments.

[0106] Example 3:

[0107] This embodiment proposes a computer-readable storage medium that stores executable instructions. When these instructions are executed, if they are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0108] The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the cache optimization method for cellless multiple input / output environment based on the VPPO algorithm described in the various embodiments of this application.

[0109] The aforementioned storage media include: flash memory, hard disks, multimedia cards, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR) memory), random access memory (RAM), static random-access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, disks, optical discs, servers, APP (Application) app stores, and other media capable of storing program verification codes. These media store computer programs, which, when executed by a processor, can implement the various steps of the cache optimization method for cellless multi-input / output environments based on the VPPO algorithm described above.

[0110] Example 4:

[0111] This embodiment proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm.

[0112] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.

[0113] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0114] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of this disclosure and its equivalents, then the intent of this disclosure also includes these modifications and variations.

Claims

1. A cache optimization method for cellless multiple-input-output environments based on the VPPO algorithm, characterized in that, The specific steps include the following: Multiple independent environment instances are generated. Each environment instance simulates the interaction scenarios of different edge servers and smart devices. The environment instances are completely isolated from each other and run in parallel. Each environment instance includes a smart agent, which corresponds to the edge server in the interaction scenario. Construct an Actor policy network and a Critic value network. Train the Actor policy network based on multiple independent environment instances and the Critic value network to obtain the trained Actor policy network. Input the current cache ratio of the edge server into the trained Actor policy network to obtain the adjustment of the cache ratio of the edge server.

2. The cache optimization method for cellless multi-input / output environments based on the VPPO algorithm according to claim 1, characterized in that, The process of training the Actor policy network based on multiple independent environment instances and the Critic value network to obtain the trained Actor policy network specifically includes: S1: In each environment instance In this process, the agent independently generates a state. ,in Indicates the number of the environment instance, and , For the number of environment instances, For a specific moment; The state Indicates in The current cache ratio of the edge server corresponding to the intelligent agent at any given time, wherein the cache ratio is the ratio of cache requests faced by the edge server to the currently stored content; S2: Utilize an Actor policy network to simultaneously process the states generated by agents in all environment instances. Output an action for each agent. ; Specifically, the Actor policy network first considers the states generated by the agents in all environment instances. Generate a Gaussian policy distribution as the probability distribution of the agent's next action, and then generate the agent's action based on the probability distribution of the next action. The actions of the intelligent agent This indicates an adjustment to the cache ratio of the edge server corresponding to the intelligent agent; S3: Agents in each environment instance receive actions distributed by the Actor policy network. Then, it executes the calculation of rewards, and takes the average of the rewards obtained by each environment instance to obtain the final reward. Each environment instance generates the next state. All trajectory data generated by all agents in their respective environment instances are collected into the transition information pool; the trajectory data includes states. ,action ,award and the next state ; S4: The Critic value network calculates the evaluation state value function based on the trajectory data in the transition information pool. The parameters of the Critic value network are updated using the TD-error method. S5: Update the Actor policy network parameters based on the trajectory data in the transition information pool and the Critic value network; S6: Determine whether the current iteration count has reached the set number or whether the Actor policy network has converged. If so, obtain the trained Actor policy network; otherwise, return to S2.

3. The cache optimization method for cellless multi-input / output environments based on the VPPO algorithm according to claim 2, characterized in that, The Gaussian strategy distribution is as follows: (1); in, In the state Take action below The probability density, For the Actor policy network, based on the state The mean of the output Gaussian distribution. For the Actor policy network, based on the state The standard deviation of the output Gaussian distribution.

4. The cache optimization method for cellless multi-input / output environments based on the VPPO algorithm according to claim 2, characterized in that, The reward The calculation method is as follows: (2); (3); (4); (5); in, For service quality, For the service number, For the total number of services that need to be satisfied, The number of services that can be served at present. express Time of the first Total demand for each service express Time of the first Service priority For the first The delay of the service Due to pre-release delay, For subsequent transmission delay, This refers to the latency parameters during the transmission process from the edge server to the smart device. Store the first for edge servers The amount of content for each service This refers to the latency parameters for edge servers downloading from cloud servers. This is a balancing parameter used to dynamically adjust latency and service fulfillment rate.

5. The cache optimization method for cellless multi-input / output environments based on the VPPO algorithm according to claim 2, characterized in that, The step of updating the Actor policy network parameters based on the trajectory data in the transition information pool and the Critic value network specifically includes: S5.1: Calculate the cutoff weight using the Retrace algorithm Then calculate the advantage function. ; The truncation weight for: (6); in, In the state Take action below The probability density, i.e., the current action strategy. For the target probability strategy; Advantage function The calculation formula is: (7); in, For hyperparameters, for Discount factor of time, for Dominance function estimation at time 1; S5.2: Based on the dominant function The probability ratio of the new and old strategies is calculated and corrected using the Clip pruning mechanism to obtain the corrected probability ratio of the new and old strategies. The revised probability ratio of the old and new strategies The calculation formula is: (8); in, This represents the ratio of the probability of the new strategy to the probability of the old strategy after modification. Expressing expectations, The importance sampling ratio represents the probability of using the new and old strategies. The clipping function represents the process of clipping the data. Limited to and between, For control strategy hyperparameters; S5.3: Update the parameters of the Actor policy network based on the corrected probability ratio of the old and new policies; (9); in, For the parameters of the updated Actor policy network, The parameters of the Actor policy network before the update. The learning rate of the Actor policy network. Indicates to Perform gradient calculations.

6. An electronic device, characterized in that, include: One or more processors, and a memory for storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm as described in any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed, cause the processor to perform the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm as described in any one of claims 1-5.

8. A computer program product, characterized in that, Includes a computer program or instructions that, when executed by a processor, implement the cache optimization method for a cellless multiple input / output environment based on the VPPO algorithm as described in any one of claims 1-5.

Citation Information

Cited By

  • Drainage decision-making method, system and equipment for rice field irrigation and medium

    CN122089113A