Edge computing and cache enabling meta-universe intelligent optimization method and system

Through the deep reinforcement learning method of active reasoning, the complexity and real-time problems of resource management in the metaverse system are solved, the optimization of content caching, task offloading and resource allocation is achieved, and the adaptability and efficiency of the system are improved.

CN120706207APending Publication Date: 2025-09-26XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510568944.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The high dynamism and uncertainty of user needs and network environment in the metaverse system lead to complex resource management, which is difficult to be effectively solved by traditional single-objective optimization methods. Resource allocation needs to consider multiple performance indicators such as latency, energy consumption and resource utilization at the same time, and has high real-time requirements.

Method used

The active inference-based deep reinforcement learning (ADRL) method is adopted to initialize the global network parameters of the intelligent agent, generate candidate strategies, calculate the free energy value, update the strategy distribution, and optimize the global network parameters in combination with the backpropagation algorithm to achieve optimal control of joint content caching, task offloading and resource allocation.

Benefits of technology

It significantly improves the system's dynamic environmental adaptability and resource utilization, reduces task offloading delays and system energy consumption, achieves multi-objective optimization balance, and meets the efficient operation requirements of the Metaverse Network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706207A_ABST
    Figure CN120706207A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of wireless communication, and discloses an edge computing and cache enabling meta-universe system and an intelligent optimization method, and the technical scheme of the invention comprises the following core optimization objectives: firstly, through an intelligent content cache strategy, a user cache hit rate is maximized, and unnecessary data transmission is reduced; secondly, task unloading decisions are optimized, computing resources are reasonably distributed, and energy consumption of user terminals is reduced; and thirdly, dynamic intelligent allocation of computing resources is realized, and the overall resource utilization efficiency of the system is improved. The ADRL algorithm provided by the invention has the following unique advantages: 1, by introducing an active reasoning mechanism, the decision ability of the algorithm in an uncertain environment is enhanced; secondly, in combination with preference information of the intelligent agent, an optimization strategy is more targeted; and thirdly, comprehensive balance of a multi-dimensional optimization target of the system is realized, and powerful technical support is provided for efficient operation of the element universe network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication technology, and in particular relates to a metaverse intelligent optimization method and system for edge computing and cache empowerment. Background Art

[0002] The metaverse, a comprehensive virtual space that integrates virtual reality (VR), augmented reality (AR), blockchain, and artificial intelligence, has rapidly emerged in recent years. Its core goal is to provide users with an immersive, interactive experience, building a digital world that is parallel to and interconnected with the real world. The four key characteristics of the metaverse—immersion, interactivity, persistence, and affordability—are achieved through VR / AR technologies, natural language processing, gesture recognition, and blockchain. However, the implementation of the metaverse faces significant technical challenges, including the need for efficient computing and storage resources, low-latency network transmission, and intelligent resource allocation mechanisms. These challenges directly impact system performance and user experience.

[0003] Edge computing, a distributed computing paradigm, effectively reduces network latency and bandwidth consumption by deploying computing and storage resources close to users or data sources, making it a key approach to addressing the technological challenges of the metaverse. The core concept of edge computing is to migrate computing tasks from centralized cloud computing data centers to edge nodes on the network, enabling localized data processing and rapid response. In the metaverse, edge computing can provide users with a low-latency service experience. For example, when interacting in real time or watching high-definition videos, edge computing nodes can process data locally, reducing transmission delays. Furthermore, edge computing can offload the computational burden from cloud computing centers, improving system scalability and reliability. Specifically, typical applications of edge computing in the metaverse include task offloading, content caching, and resource allocation. Task offloading reduces the computing pressure on terminal devices by migrating computationally intensive tasks (such as 3D rendering and physics simulation) to edge nodes. Content caching stores popular content (such as virtual scenes and user data) on edge nodes to reduce data transmission latency and bandwidth consumption. Resource allocation dynamically allocates computing, storage, and network resources to edge nodes to optimize system performance and user experience. However, the application of edge computing in the metaverse also faces many challenges, such as the limited resources of edge nodes, the complexity of coordinated scheduling and resource allocation, and the high dynamics and uncertainty of service demands and user behavior. These all require innovative optimization methods to solve.

[0004] Deep reinforcement learning (DRL), a machine learning method that combines deep learning and reinforcement learning, has demonstrated tremendous potential in the field of resource optimization in recent years. Its core concept is that an intelligent agent learns optimal decision-making strategies based on trial and error through interaction with the environment to maximize cumulative rewards. Deep learning, on the other hand, enhances the learning capabilities of intelligent agents by extracting features and recognizing patterns from high-dimensional input data through neural network models. Deep reinforcement learning offers significant advantages in resource optimization, including adaptability, high dimensionality, and end-to-end learning capabilities. These advantages enable it to dynamically adjust decision-making strategies based on environmental changes, handle complex optimization problems, and learn and optimize strategies directly from raw data without the need for manually designed features or models. In the context of the metaverse, deep reinforcement learning can be applied to optimization problems such as task offloading, content caching, and resource allocation. For example, in task offloading, intelligent agents can dynamically decide whether to offload tasks to edge nodes based on task complexity, edge node computing power, and network conditions. In content caching, intelligent agents can develop optimal caching strategies based on user demand patterns, content popularity, and storage resource constraints. In resource allocation, intelligent agents can learn from dynamic environmental changes and task requirements to develop adaptive resource allocation strategies. However, the application of deep reinforcement learning in resource optimization also faces challenges, including high training data requirements, poor interpretability of policy models, and potential slow convergence and poor generalization in complex and changing environments.

[0005] In edge computing and the metaverse, task offloading, content caching, and resource allocation are three core optimization problems. Task offloading involves migrating compute-intensive tasks from end devices to edge nodes or cloud computing centers for processing, while content caching focuses on storing popular content at edge nodes to reduce data transmission latency and bandwidth consumption. Resource allocation requires dynamically allocating computing, storage, and network resources within edge and cloud computing environments to optimize system performance and user experience. In recent years, with the advancement of deep reinforcement learning, task offloading, content caching, and resource allocation problems have gradually shifted from traditional optimization models to intelligent decision-making models. However, research on this emerging scenario, the metaverse, is still in its infancy. Its complexity and dynamism place higher demands on resource management, and innovative optimization methods are urgently needed.

[0006] In the context of the metaverse, task offloading, content caching, and resource allocation face numerous technical challenges. First, the highly dynamic and uncertain nature of user needs and network environments requires adaptive resource management mechanisms. Second, computing, storage, and network resources are distributed across multiple layers, making resource allocation highly complex. Furthermore, resource optimization requires considering multiple performance metrics simultaneously (such as latency, energy consumption, and resource utilization), making it difficult to effectively address traditional single-objective optimization methods. Finally, the application scenarios of the metaverse place extremely high demands on real-time resource management, requiring decisions to be made within a limited timeframe.

[0007] Through the above analysis, the problems and defects of the existing technology are as follows:

[0008] (1) The high dynamism and uncertainty of user needs and network environments require an adaptive resource management mechanism. Secondly, computing, storage, and network resources are distributed at multiple levels, and resource allocation problems are highly complex. In addition, resource optimization problems require considering multiple performance indicators (such as latency, energy consumption, resource utilization, etc.) at the same time, which is difficult to be effectively solved by traditional single-objective optimization methods.

[0009] (2) The application scenarios of the metaverse have extremely high requirements for the real-time nature of resource management, and decisions need to be made within a limited time. Summary of the Invention

[0010] In response to the problems existing in the existing technology, the present invention provides a metaverse intelligent optimization method and system enabled by edge computing and cache.

[0011] The present invention is implemented as follows: a metaverse intelligent optimization method for edge computing and cache empowerment includes:

[0012] S101, initializing the agent global network parameters;

[0013] Initialize the agent's global network parameters, iteration number N ep , maximum number of iterations N, number of candidate strategies J, planning horizon H, number of rounds T max , the number of users N, the integrated network parameter θ1, the reward network parameter θ2, the number of optimal candidate strategies k, etc.; at the same time, the initialization strategy distribution η(π) and the initialization transition probability distribution δ(s t |s t-1 ,θ,π), etc., to provide initial conditions for subsequent optimization algorithms;

[0014] S102, candidate strategy generation;

[0015] At the beginning of each round, set the initial state s t , and randomly select J candidate strategies from the initial strategy distribution as the strategy set to be evaluated in the current round;

[0016] S103, action sampling and reward calculation;

[0017] Perform action sampling on each candidate strategy extracted in S102 to obtain J actions. At the same time, calculate the conditional transition probability distribution corresponding to each action, and calculate the corresponding immediate reward based on the result of the interaction between the action and the environment.

[0018] S104, free energy calculation;

[0019] According to active inference theory, the free energy value of each candidate strategy is calculated using the rewards and conditional transition probability distributions obtained from interacting with the environment to evaluate the quality of the strategy.

[0020] S105, strategy distribution update;

[0021] Select k candidate strategies with the smallest free energy values, average their free energy values, and update the strategy distribution based on the average value, so that the strategy distribution gradually converges to a better area;

[0022] S106, action execution and status update;

[0023] Based on the updated policy distribution in S105, a policy is sampled, and an action is further sampled from the policy as the current action of the agent; the agent executes the action and interacts with the environment to obtain the next state;

[0024] S107, experience storage;

[0025] The current experience tuple, including state, action, reward, and next state, is stored in the replay buffer. If the replay buffer is full, the first-in-first-out principle is used to overwrite the oldest stored experience with the new experience.

[0026] S108, network parameter update;

[0027] Using the error function between the predicted free energy value output by the global network and the actual free energy value, the back propagation algorithm and gradient descent method are used to update the global network parameters to improve the prediction accuracy and strategy optimization effect;

[0028] S109, iterative convergence;

[0029] Repeat the above steps until the algorithm converges; eventually, an optimized policy distribution is obtained. The agent can randomly extract policies based on this distribution to achieve optimal control of joint content caching, task offloading, and resource allocation, thereby maximizing system performance.

[0030] Furthermore, in S102, at the beginning of each round, the system initializes and sets the initial value of the current state; then, J candidate strategies are randomly selected from the strategy distribution, and corresponding actions are further randomly selected from each candidate strategy; wherein, the state of the agent is represented as a set of the following multidimensional vectors:

[0031]

[0032] in: in is the number of CPU cycles required for user i to render the front-end interaction in time slot t; in is the data size of the front-end interactive information of user i in time slot t; b(t) = {b i (t)}, where b i (t) is the index of the background model requested by user i in time slot t; in is the set of background objects requested by user i in time slot t; q(t) = {q i (t)}, where q i (t) is the resolution level required by user i to request content from the metaverse in time slot t; R b (t) is the wireless transmission rate between the base station and the remote Metaverse server in time slot t; is the set of models cached in the MEC server in time slot t-1; is the set of objects of model s cached in the MEC server in time slot t-1; rm (t-1) is the remaining available cache resources of the MEC server after the content of time slot t-1 is placed; Life md (t)=Life s (t), is the life cycle of model s in time slot t; Life obj (t)=Life s,l (t), is the life cycle of object l of model s in time slot t; V md (t)={v s (t)}, is the removal indicator of model s in time slot t; V obj (t)={v s,l (t)}, is the removal indicator of object l of model s in time slot t; self-policy π t The sampled action is represented as:

[0033]

[0034] Among them, the model's cache decision κ(t) = {κ s (t)}, The cache decision of the object ζ(t) = {ζ s,l (t)}, Offloading decision of UE-side interaction Computing resource allocation in MEC servers mec (t) = {f i mec (t)}, Computational resource allocation in UE loc (t) = {f i loc (t)}, Wireless transmission rate allocation R(t) = {R i (t)},

[0035] Furthermore, S103 specifically includes: calculating the conditional transition probability distribution for each candidate strategy; calculating the instant reward value at the current moment according to the actions corresponding to each candidate strategy; the calculation formula of the instant reward is as follows:

[0036] r(t)=v(t)r imm (t),

[0037] The binary variable v(t) serves as an indicator. If all constraints are satisfied, then v(t) = 1 and the immediate reward is set to the system utility at the current time slot t. Otherwise, if the action violates any constraint, the system will be penalized and v(t) = 0 will be set. Therefore, r(t) = 0 in these cases. imm (t) is the immediate reward the user receives when the constraint is satisfied, namely:

[0038]

[0039] This formula reflects the weighted sum of the UE's QoE, the number of cache hits, and the UE's energy consumption in each time slot.

[0040] Furthermore, the step S104 specifically includes: calculating the free energy of each candidate strategy based on active inference theory and the free energy principle, combined with the cumulative reward value and conditional transition probability distribution obtained in step S103; wherein the calculation formula for the inverse value of free energy is as follows:

[0041]

[0042] Furthermore, S105 is specifically as follows: k minimum values ​​are screened out from the free energies of all candidate strategies and the arithmetic average is calculated for these values, and the probability distribution of the current optimal strategy is generated based on the obtained free energy average value; this process is equivalent to taking the k maximum values ​​of opposite free energy values, sorting these maximum values ​​in descending order, and then calculating the average value to determine the strategy distribution, that is:

[0043]

[0044] And get the policy distribution:

[0045]

[0046] Here, σ(·) represents the diagonal Gaussian distribution, a typical multivariate probability distribution characterized by the independence of its dimensions. This distribution is determined by two key parameters: the mean vector and the variance vector on the diagonal. In practical applications, the diagonal Gaussian distribution is often used as a prior distribution or a posteriori distribution. Its advantage is that it can effectively prevent model overfitting while simplifying complex inference processes. In the field of reinforcement learning, the diagonal Gaussian distribution can be used to model state uncertainty. This property enables intelligent agents to make more reasonable decisions in uncertain environments. By quantifying state uncertainty, intelligent agents can achieve a better balance between exploration and exploitation, thereby improving learning effects and decision quality.

[0047] Furthermore, in S106, we first perform strategy sampling based on the obtained strategy distribution δ(π), and then further extract specific actions a from the sampled strategies. t , as the agent's decision-making behavior in the current state; this behavior is then executed by the agent, interacting with the environment in real time, thereby leading to the next state of system evolution and obtaining corresponding environmental feedback (such as reward value); the entire extraction and execution process is as follows:

[0048] π t ~δ(π),a t ~π t ;

[0049] Implementation process of S108: After the global network is responsible for outputting the predicted value of free energy, the error function between the predicted value and the target value is first calculated. Then, based on the error information, the back propagation algorithm combined with the gradient descent method is used to iteratively optimize the global network parameters. The loss function in this optimization process is given by the following formula:

[0050]

[0051] Among them, Q(s t ,a t ; θ) is the predicted value of the global network output free energy, is the target free energy; the gradient of the loss function is given by:

[0052]

[0053] Then update the parameters of the global network by gradient descent as follows:

[0054]

[0055] Among them, θ is the internal parameter of the global model network, and lr is the learning rate.

[0056] Another object of the present invention is to provide an edge computing and cache-enabled metaverse intelligent optimization system comprising:

[0057] System Initialization Module: Responsible for initializing the parameters of the deep deterministic policy gradient algorithm. This module consists of two parts: a configuration module for setting key hyperparameters such as the network learning rate, discount factor, and replay buffer size; and a network construction module for defining basic parameters such as the environment state dimension, user location information, and network topology.

[0058] Agent module: In each iteration, the agent module generates optimized actions based on the current network state. This module uses an active inference mechanism and the principle of free energy minimization to calculate and select the minimum value among the maximum preceding free energies, and uses the mean of these minimum values ​​as a benchmark to fit the optimal strategy distribution.

[0059] Action execution module: responsible for executing content caching strategies and computing task offloading decisions, mapping optimization strategies to specific system behaviors;

[0060] Reward Acquisition Module: Calculates the immediate reward value after executing an action. This module comprehensively considers the user experience quality and system overhead of all terminal devices in the Metaverse system, and performs quantitative evaluation based on a preset reward function.

[0061] State transfer module: After the agent performs an action, it transfers the system state from the current state to the next state and records the relevant state transfer information;

[0062] Experience replay module: used to store key information such as system status, executed actions, immediate rewards, cumulative rewards, and next state during each interaction process, forming experience tuples to provide data support for subsequent model training;

[0063] Sampling module: includes the strategy distribution sampling unit and the strategy action sampling unit, which are responsible for sampling candidate strategies from the strategy distribution and extracting specific actions from the selected strategy respectively;

[0064] Global network update module: uses the back-propagation algorithm to optimize the global network. Its core is the parameter optimization unit, which uses the gradient descent method to adjust the network parameters to minimize the loss function, thereby improving the overall performance of the system.

[0065] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the edge computing and cache-enabled metaverse system and intelligent optimization method.

[0066] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the edge computing and cache-enabled metaverse system and intelligent optimization method.

[0067] Another object of the present invention is to provide an information data processing terminal, which is used to implement the edge computing and cache-enabled metaverse system and intelligent optimization system.

[0068] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0069] To address the technical challenges of high dynamism, information acquisition uncertainty, and data incompleteness in metaverse network scenarios, this paper proposes a multi-access edge computing (MEC) and cache-enabled metaverse system and intelligent optimization method based on active inference enabled deep reinforcement learning (ADRL). This method achieves multi-dimensional system optimization by modeling the network optimization problem as a Markov decision process, comprehensively considering content caching strategy, computing task offloading decision-making, and resource allocation scheme. This technical solution significantly improves the system's adaptability to dynamic environments. By introducing agent preference parameters, it effectively balances user experience quality (QoE) and system resource utilization. At the same time, it achieves a multi-objective optimization balance, significantly improves cache hit rate, and reduces task offloading delay and system energy consumption. In addition, the algorithm adopts the free energy minimization principle and diagonal Gaussian distribution policy update mechanism, showing good convergence and stability. The method also has good real-time and scalability. In large-scale user concurrent testing, it can respond to user requests in real time, and the system performance remains stable as the number of users increases. In summary, the present invention not only effectively solves the problems of the existing technology, but also provides strong technical support for the efficient operation and resource management of the metaverse system, and has significant creativity and practical value.

[0070] The technical solution of the present invention includes the following core optimization goals: first, maximizing the user cache hit rate and reducing unnecessary data transmission through intelligent content caching strategies; second, optimizing task offloading decisions, rationally allocating computing resources, and reducing user terminal energy consumption; and third, achieving dynamic intelligent allocation of computing resources and improving the overall resource utilization efficiency of the system. Through the coordinated optimization of the above technical solutions, the present invention effectively solves the resource constraints and service quality assurance issues faced in the Metaverse network scenario, significantly reducing system operating costs while ensuring the quality of user experience.

[0071] The ADRL algorithm proposed in this invention has the following unique advantages: First, by introducing an active reasoning mechanism, the algorithm's decision-making ability in an uncertain environment is enhanced; second, by combining the agent's preference information, the optimization strategy is made more targeted; third, a comprehensive balance of the system's multi-dimensional optimization goals is achieved, providing strong technical support for the efficient operation of the metaverse network.

[0072] This invention aims to propose an optimization method based on deep reinforcement learning, combining the distributed advantages of edge computing to achieve intelligent optimization of task offloading, content caching, and resource allocation. Specific objectives include: dynamically formulating task offloading strategies by learning the characteristics of tasks and network conditions in the metaverse; developing adaptive content caching mechanisms by learning user access patterns and environmental changes; achieving efficient resource allocation through multi-objective optimization and collaborative scheduling; and designing a lightweight deep reinforcement learning model to ensure real-time decision-making and support the scalability of large-scale metaverse systems. Through this research, this invention will provide a theoretical basis and solutions for resource management in metaverse systems, promoting the development and practical application of metaverse technology.

[0073] This invention proposes a method for optimizing task offloading, content caching, and resource allocation for the metaverse system. By innovatively integrating edge computing and distributed storage technologies, it significantly improves the system's overall performance and flexibility. The system employs an intelligent task offloading mechanism to dynamically allocate compute-intensive tasks to edge servers or the cloud, optimizing the order of task execution and reducing the computational burden on terminal devices, ensuring efficient processing of complex tasks in high-concurrency scenarios. Furthermore, based on user behavior analysis and content access patterns, the system implements a dynamic cache allocation strategy to preload popular content to edge nodes or local devices, significantly reducing content retrieval latency and alleviating backhaul bandwidth pressure. Furthermore, the invention uses a resource allocation optimization algorithm to monitor system resource status in real time and dynamically adjust the allocation of computing, storage, and network resources, avoiding resource waste and system bottlenecks. Furthermore, by combining artificial intelligence and machine learning techniques to predict load changes and proactively optimize resource allocation strategies, the system further enhances its adaptability and stability. This technical solution provides an efficient, flexible, and reliable solution for the metaverse system, meeting the needs of low-latency, high-concurrency scenarios. It promotes the innovation and development of metaverse technology and opens up new possibilities for future intelligent application scenarios.

[0074] This paper proposes an optimization method based on active inference in the MEC and cache-enabled metaverse system, which achieves significant progress in the following key aspects:

[0075] 1) Optimization of content caching strategy:

[0076] Through intelligent content pre-caching, we optimize the type and quantity of content stored on edge servers, significantly reducing content retrieval latency and improving system responsiveness. This optimization not only reduces the load on remote servers but also improves the overall performance of the Metaverse system, demonstrating greater efficiency and stability in dynamic environments.

[0077] 2) Efficient strategies for task offloading and resource allocation:

[0078] By precisely optimizing the allocation of base station resources to user devices and offloading local computing tasks to users, the Metaverse network significantly improves resource utilization. This strategy not only reduces system overhead and latency, but also ensures efficient and reliable task processing, providing more flexible support for complex application scenarios.

[0079] 3) Fusion of Deep Reinforcement Learning and Active Inference:

[0080] By combining the DRL algorithm with active inference methods, the system not only relies on traditional reward mechanisms during decision-making but also fully utilizes additional information from the environment. This combination gives the system the ability to autonomously learn and adapt based on real-time data, enabling it to make optimal decisions without explicit instructions, significantly improving its processing capabilities in complex and dynamic environments.

[0081] 4) Balance optimization of user QoE, cache hit count, and system energy consumption:

[0082] This paper designs a reward mechanism that focuses on improving QoE and cache hits while reducing system energy consumption. This optimization not only improves system energy efficiency but also meets environmental requirements. This has important practical significance in the context of rising energy costs and the increasing importance of sustainable development.

[0083] 5) Enhanced system reliability and adaptability:

[0084] By screening and iteratively optimizing multiple candidate strategies, the predicted free energy values ​​gradually approach the true values, significantly improving the system's reliability. Furthermore, by expanding the agent's exploration space, the system is no longer limited to a deterministic reward function, but instead comprehensively considers multiple factors, further enhancing its adaptability and robustness in dynamic environments.

[0085] 6) The network's autonomous learning and optimization capabilities:

[0086] Through continuous iterative training and experience-based network adjustments, the system can continuously optimize its decision-making mechanisms and improve overall performance. This self-learning capability has brought significant performance improvements to the MEC and cache-enabled metaverse system, including greater stability, lower overhead, and stronger adaptability.

[0087] The core of the active inference-based optimization method in the MEC and cache-enabled metaverse system provided by this invention is to use mathematical models to guide the behavior and learning process of intelligent agents. The technical effects brought by these mathematical models can be explored based on their characteristics:

[0088] 1) Free energy

[0089] The free energy calculation not only considers user QoE, cache hits, and system energy consumption, but also includes a factor of environmental information, called information gain, which additionally considers agent preferences and increases agent initiative. By directly linking rewards to user QoE, cache hits, and system energy consumption, this approach encourages agents to explore environments where the weighted sum tends to be minimized, thereby improving user service quality.

[0090] 2) Global network loss function and update

[0091] Through a free energy evaluation mechanism based on approximate calculations, the global network innovatively integrates backpropagation and gradient descent algorithms to achieve dynamic parameter updates.

[0092] Strategy optimization: The system reconstructs the strategy distribution by autonomously optimizing the global network parameter space, constructing a set of decision-making strategies that are both diverse and efficient, thereby generating specific execution strategies that are adaptable to the environment. Learning stability control: The system adopts a prediction error constraint mechanism under the framework of active reasoning theory. By establishing an algebraic model of the objective function based on free energy minimization, it strictly regulates the parameter update path and effectively balances the dynamic relationship between theoretical exploration and model convergence. This mechanism transforms the traditional empirical weight update into an optimization process with information-theoretic constraints through mathematical methods, significantly enhancing the robustness of the learning process. Performance optimization: The system designs a loss function calculation architecture that includes regularization constraints, and cooperates with the dynamic batch update mechanism of network parameters to achieve a coordinated improvement in strategy valuation accuracy and decision-making effectiveness. This rigorous multi-objective optimization system based on active reasoning enables the network to achieve Pareto optimal improvements in multiple key indicators such as learning rate, strategy effectiveness, and execution reliability.

[0093] 3) Parameter update formula

[0094] The parameter update method of the global network is described.

[0095] Policy Gradual Approach: By gradually updating the global model network parameters, the system can smoothly transition to the new policy and prevent performance fluctuations due to drastic changes.

[0096] Continuous learning and adaptation: This continuous parameter update mechanism ensures that the system can adapt to long-term environmental changes.

[0097] The application of the mathematical model provided by this invention not only improves the operational efficiency and decision-making quality of MEC and cache-enabled metaverse systems, but also enhances their adaptability to environmental changes and long-term stability. These technical effects are crucial for modern edge computing environments that process large amounts of data and high-frequency interactions.

[0098] The present invention provides an optimization method based on active reasoning in the MEC and cache-enabled metaverse system, which optimizes the performance of the network through the interaction between the intelligent agent and the environment.

[0099] Initialize the agent's state, which includes multiple variables, such as the amount of data and computational complexity in the base station's cache, the user set's content requests, the base station's cached content set, and the lifecycle of the base station's cached content. These variables collectively define the agent's environmental state at a specific moment, which in turn influences the agent's decision-making.

[0100] The agent initially simulates executing actions and earns immediate rewards, then simulates generating actions and earning rewards. After a series of tedious screening processes, the agent determines the optimal action, at which point it executes the actual action and transitions to another state. The reward calculation formula considers the QoE of all users, cache hits, and system energy consumption. This is a comprehensive system design goal: to improve user experience quality and cache hits while reducing system energy consumption.

[0101] Calculating the loss function and updating the network, using gradient descent to update the current global network, can help improve the global network's prediction accuracy of free energy, thereby optimizing system performance. This part of the operation can be compared to updating the value function in deep reinforcement learning. The key is to improve prediction accuracy to guide policy improvement.

[0102] These steps and the application of mathematical models have led to significant technological advances:

[0103] Policy optimization: Through deep reinforcement learning, the system can self-learn and optimize strategies to adapt to the ever-changing network environment.

[0104] Efficient resource utilization: Optimization of content caching task offload distribution ensures that all device computing power is efficiently utilized.

[0105] Cache accuracy: By optimizing the number of user cache hits, we ensure that edge nodes store high-value data, effectively reducing the redundant cache rate and the frequency of request back to the source;

[0106] Minimize energy consumption: By optimizing system energy consumption, it helps achieve green communications and reduce the impact on the environment.

[0107] System stability and adaptability: It adopts a unique active reasoning mechanism that pays attention to additional information besides rewards, increases the agent's attention to its own preferences, and thus enhances stability and universality.

[0108] These advances demonstrate the potential of active inference-based deep reinforcement learning in complex network systems, especially in realizing intelligent and efficient metaverse communication networks.

[0109] The application of the present invention has achieved significant technological progress. On the one hand, through the learning and decision-making of the intelligent agent, the method can adapt to dynamically changing environments and achieve efficient resource utilization and task execution. On the other hand, by introducing an experience replay mechanism and a gradient ascent and descent update strategy, the present invention improves the stability and convergence speed of learning, thereby further improving the overall performance of the system. In addition, the method exhibits low computational complexity when processing the Metaverse network, meeting real-time requirements. Finally, by optimizing the relationship between user QoE, cache hit counts, and system energy consumption, the present invention achieves an effective balance between Metaverse network service deployment and resource allocation.

[0110] In summary, the present invention solves the problems existing in the metaverse network service deployment and resource allocation methods by introducing MEC and ADRL technologies, and realizes intelligent, efficient and dynamic management of metaverse network services.

[0111] This approach deploys edge servers on base stations. It aims to maximize user QoE while minimizing user overhead, taking into account cache replacement decisions for background models and objects, task offloading decisions for user device foreground information, computational resource allocation to MEC servers, local computational resource allocation, and wireless transmission rate allocation strategies for MEC servers. Its application not only improves resource utilization and task execution success rates, but also reduces computational complexity and meets real-time requirements, bringing significant technological advancement and application value to the development of the metaverse industry.

[0112] The proposed MEC and cache-enabled metaverse system optimization method solves existing technical challenges in resource scheduling, network bandwidth, and content caching efficiency by innovatively combining active reasoning, free energy models, and joint content caching, task offloading, and resource allocation algorithms, bringing significant technological progress. This is specifically reflected in the following aspects:

[0113] 1. Efficient resource scheduling and optimization in dynamic environments

[0114] Existing technical problems: Traditional MEC systems use mostly static or fixed strategies for resource scheduling, which are difficult to adapt to the needs of multi-user and highly dynamic environments, resulting in low resource allocation efficiency and inability to achieve real-time response.

[0115] Technological Advancement: This invention combines active inference with free energy theory to construct a dynamic adjustment mechanism. The agent updates its policy distribution based on free energy calculations within each time slot. Using transition probability distributions and cumulative rewards, it performs multiple iterations of training, enabling resource scheduling to efficiently adapt to the dynamic needs of different users. Through continuous training, the agent is able to dynamically and in real time perform task offloading and resource optimization in complex multi-user environments, improving system responsiveness and resource allocation efficiency.

[0116] 2. Accurate content caching and task offloading strategies

[0117] Existing technical problems: Existing content caching and task offloading technologies usually cannot optimize cache content and offloading strategies at the same time, especially in the metaverse system, which faces the dual challenges of limited edge server resources and dynamic cache requirements.

[0118] Technological Advancement: This invention combines optimized content caching and task offloading strategies, enabling intelligent agents to dynamically adjust base station cache content and offloading strategies based on user needs. Leveraging random sampling of policy distributions and a backpropagation algorithm, the system accurately caches appropriate content, maximizing user QoE and cache hits while avoiding resource waste caused by excessive caching or frequent offloading. Compared to traditional methods, this optimization strategy significantly improves caching and offloading accuracy and reduces system resource consumption.

[0119] 3. Efficient integration of free energy model and active inference

[0120] Existing technical issues: Traditional optimization methods typically use fixed or single reward functions, making it difficult to fully represent the multi-dimensional and heterogeneous user needs and environmental characteristics. This limitation results in rigid behavior of intelligent agents in complex and changing scenarios, lacking the ability to adapt to different task modes and environmental dynamics. Traditional methods often exhibit significant performance degradation, especially in extreme scenarios where demand distributions shift rapidly or resource constraints suddenly change.

[0121] Technological progress: This invention innovatively embeds free energy theory into the active reasoning framework to construct a multimodal policy optimization mechanism. Specifically, by defining the free energy function as the core indicator of policy evaluation, the system quantifies the diversity of user needs, the uncertainty of environmental states, and the complexity of resource constraints into an optimizable mathematical model. During the decision-making process, the intelligent agent dynamically evaluates the free energy of candidate strategies based on the joint calculation of the conditional transition probability distribution and the cumulative reward function, and selects the optimal strategy by minimizing the free energy. This method successfully achieves efficient adaptability to complex scenarios by modeling the environmental interaction process as a Bayesian reasoning framework.

[0122] 4. Rapid algorithm convergence and precise strategy generation

[0123] Existing technical issues: Existing reinforcement learning algorithms face significant convergence issues in the Metaverse system. These issues manifest themselves in challenges such as an overly large policy search space, unclear gradient update directions, and sparse reward signals, making it difficult for the algorithms to quickly generate efficient policies. This lack of convergence severely impacts the system's real-time performance, making it unable to effectively adapt to the rapidly changing task demands and resource constraints of the Metaverse's dynamic environment, and making it difficult to meet the high-responsiveness and low-latency requirements of applications.

[0124] Technological progress: In response to the above problems, the present invention proposes a fast convergence algorithm based on multi-dimensional parameter joint optimization. By initializing the core parameter space of the agent (including the number of candidate strategies, planning horizon, learning rate, discount factor, etc.), and combining the backpropagation algorithm with the adaptive momentum gradient descent method, an efficient policy distribution optimization framework is constructed. During the training process, the algorithm gradually guides the agent to approach the global optimal strategy by dynamically adjusting the entropy regularization coefficient and gradient update step size of the policy distribution. In addition, the present invention introduces an experience pool management mechanism based on priority experience replay, which significantly improves the utilization rate of rare reward samples and further accelerates the convergence process of the algorithm.

[0125] 5. Experience replay mechanism and data utilization efficiency improvement

[0126] Existing technical issues: Metaverse systems suffer from significant issues during reinforcement learning training, such as low utilization of empirical data and insufficient training efficiency. Due to the lack of effective management and screening mechanisms for massive amounts of interactive data, a large number of high-value empirical samples are ignored, while redundant or low-quality data is reused. This results in slow convergence of the policy optimization process and makes it difficult to adapt to the highly dynamic and non-stationary environment of the Metaverse.

[0127] Technological progress: To address the above-mentioned issues, the present invention innovatively designs an efficient experience replay mechanism based on priority allocation and dynamic updating. The system constructs a two-layer buffer architecture to categorize and store the state, action, and reward data generated by the interaction between the agent and the environment. It also combines sample importance sampling with an adaptive weight adjustment strategy to significantly improve the utilization frequency of high-quality experience data. At the same time, when the buffer reaches its capacity limit, the system dynamically evaluates the data value based on the temporal difference error of the sample, giving priority to retaining the latest experience with high information content, thereby ensuring that the buffer always stores the most representative environmental interaction information. This method optimizes the computational efficiency of the training process and the quality of strategy generation by reducing data redundancy and improving sample utilization.

[0128] 6. System energy consumption and cost optimization

[0129] Existing technical problems: The content caching and resource allocation methods of existing MEC systems mostly rely on fixed parameters, making it difficult to balance energy consumption and cost in a dynamic environment.

[0130] Technological Advancement: This invention integrates the system's energy consumption, latency, and service cost into a weighted sum, which serves as the denominator of the immediate reward. By maximizing the inverse of this objective function, the intelligent agent can effectively balance resource allocation and cost consumption, achieving optimal energy efficiency. Combined with dynamic threshold control for QoE, this system significantly reduces resource consumption while meeting user needs, improving the system's overall economic benefits.

[0131] In summary, this invention achieves a series of technological breakthroughs in a multi-user dynamic environment: efficient resource scheduling in virtual-reality fusion scenarios through a free-energy-driven adaptive allocation mechanism; cross-platform data collaborative management based on priority replay and dynamic cache updates; a fast-convergence algorithm architecture combining policy entropy regularization and adaptive gradient optimization; and a joint energy-latency optimization model for global gaming of heterogeneous computing units. These innovative achievements establish the core algorithmic framework of the Metaverse edge computing system, providing key technical support for typical scenarios such as virtual reality collaboration, digital twin simulation, and cross-platform virtual interaction.

[0132] Second, as auxiliary evidence for the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:

[0133] First, the technical solution of this invention is expected to bring significant commercial value and revenue upon transformation. By optimizing task offloading, content caching, and resource allocation within the Metaverse system, this invention can effectively improve user experience quality, reduce system energy consumption and operating costs, and thus provide a more efficient and economical solution for Metaverse-related industries. This solution has promising market applications and potential economic benefits.

[0134] Furthermore, this invention demonstrates significant innovation in the field of metaverse network optimization. Currently, effective resource management and optimization methods for the highly dynamic, complex, and real-time-critical network environment of the metaverse remain insufficient. This invention combines edge computing, deep reinforcement learning, and active inference theory to propose a novel optimization framework that can effectively adapt to the various changes in the metaverse environment, providing a viable technical foundation for the efficient operation of the metaverse system.

[0135] This invention also provides solutions to long-standing technical challenges in the metaverse. How to achieve efficient resource utilization and minimize energy consumption while ensuring a positive user experience has always been a major challenge facing the industry. This invention, through intelligent content caching strategies, task offloading decisions, and resource allocation optimization, provides an effective solution to this problem, providing important technical support for the further development and application of metaverse technology.

[0136] Finally, this invention overcomes the limitations of traditional technologies to a certain extent. Previous solutions often rely on static or fixed optimization strategies, making them difficult to adapt to the dynamic environment of the metaverse. This invention utilizes active inference and free energy minimization principles, combined with deep reinforcement learning, to enable the system to autonomously learn and adapt to environmental changes, thus providing a new approach and method for optimizing metaverse networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0137] Figure 1This is a flow chart of the edge computing and cache-enabled metaverse system and intelligent optimization method provided by an embodiment of the present invention.

[0138] Figure 2 This is a structural block diagram of the edge computing and cache-enabled metaverse system and intelligent optimization system provided by an embodiment of the present invention.

[0139] Figure 3 This is an applicable scene graph provided by an embodiment of the present invention.

[0140] Figure 4 This is a flow chart of a method for collaborative content caching decision, task offloading decision and resource allocation provided by an embodiment of the present invention.

[0141] Figure 5 This is a relationship diagram between system QoE and base station storage capacity after the simulation algorithm provided by an embodiment of the present invention converges.

[0142] Figure 6 This is a relationship diagram between user energy consumption and base station bandwidth after the simulation algorithm provided by an embodiment of the present invention converges.

[0143] Figure 7 This is a relationship diagram between the number of user cache hits after convergence of the simulation algorithm provided by an embodiment of the present invention and the number of objects included in each model. DETAILED DESCRIPTION

[0144] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0145] like Figure 1 As shown, an edge computing and cache-enabled metaverse system and intelligent optimization method provided by an embodiment of the present invention include the following steps:

[0146] S101, initializing the agent global network parameters;

[0147] Initialize the agent's global network parameters, iteration number N ep , maximum number of iterations N, number of candidate strategies J, planning horizon H, number of rounds T max , the number of users N, the integrated network parameter θ1, the reward network parameter θ2, the number of optimal candidate strategies k, etc.; at the same time, the initialization strategy distribution η(π) and the initialization transition probability distribution δ(s t |s t-1 ,θ,π), etc., to provide initial conditions for subsequent optimization algorithms;

[0148] S102, candidate strategy generation;

[0149] At the beginning of each round, set the initial state s t , and randomly select J candidate strategies from the initial strategy distribution as the strategy set to be evaluated in the current round;

[0150] S103, action sampling and reward calculation;

[0151] Perform action sampling on each candidate strategy extracted in S102 to obtain J actions. At the same time, calculate the conditional transition probability distribution corresponding to each action, and calculate the corresponding immediate reward based on the result of the interaction between the action and the environment.

[0152] S104, free energy calculation;

[0153] According to active inference theory, the free energy value of each candidate strategy is calculated using the rewards and conditional transition probability distributions obtained from interacting with the environment to evaluate the quality of the strategy.

[0154] S105, strategy distribution update;

[0155] Select k candidate strategies with the smallest free energy values, average their free energy values, and update the strategy distribution based on the average value, so that the strategy distribution gradually converges to a better area;

[0156] S106, action execution and status update;

[0157] Based on the updated policy distribution in S105, a policy is sampled, and an action is further sampled from the policy as the current action of the agent; the agent executes the action and interacts with the environment to obtain the next state;

[0158] S107, experience storage;

[0159] The current experience tuple, including state, action, reward, and next state, is stored in the replay buffer. If the replay buffer is full, the first-in-first-out principle is used to overwrite the oldest stored experience with the new experience.

[0160] S108, network parameter update;

[0161] Using the error function between the predicted free energy value output by the global network and the actual free energy value, the back propagation algorithm and gradient descent method are used to update the global network parameters to improve the prediction accuracy and strategy optimization effect;

[0162] S109, iterative convergence;

[0163] Repeat the above steps until the algorithm converges; eventually, an optimized policy distribution is obtained. The agent can randomly extract policies based on this distribution to achieve optimal control of joint content caching, task offloading, and resource allocation, thereby maximizing system performance.

[0164] In S102 provided in the embodiment of the present invention, at the beginning of each round, the system initializes and sets the initial value of the current state; then, J candidate strategies are randomly extracted from the strategy distribution, and corresponding actions are further randomly extracted from each candidate strategy; wherein, the state of the agent is represented as a set of the following multidimensional vectors:

[0165]

[0166] in: in is the number of CPU cycles required for user i to render the front-end interaction in time slot t; in is the data size of the front-end interactive information of user i in time slot t; b(t) = {b i (t)}, where b i (t) is the index of the background model requested by user i in time slot t; in is the set of background objects requested by user i in time slot t; q(t) = {q i (t)}, where q i (t) is the resolution level required by user i to request content from the metaverse in time slot t; R b (t) is the wireless transmission rate between the base station and the remote Metaverse server in time slot t; is the set of models cached in the MEC server in time slot t-1; is the set of objects of model s cached in the MEC server in time slot t-1; rm (t-1) is the remaining available cache resources of the MEC server after the content of time slot t-1 is placed; Life md (t)=Life s (t), is the life cycle of model s in time slot t; Life obj (t)=Life s,l (t), is the life cycle of object l of model s in time slot t; V md (t)={v s (t)}, is the removal indicator of model s in time slot t; V obj (t)={v s,l (t)}, is the removal indicator of object l of model s in time slot t; self-policy π t The sampled action is represented as:

[0167]

[0168] Among them, the model's cache decision κ(t) = {κ s (t)}, The cache decision of the object ζ(t) = {ζ s,l (t)}, Offloading decision of UE-side interaction Computing resource allocation in MEC servers mec (t) = {f i mec (t)}, Computational resource allocation in UE loc (t) = {f i loc (t)}, Wireless transmission rate allocation R(t) = {R i (t)},

[0169] S103 provided in the embodiment of the present invention specifically includes: calculating the conditional transition probability distribution for each candidate strategy; calculating the instant reward value at the current moment based on the actions corresponding to each candidate strategy; the calculation formula of the instant reward is as follows:

[0170] r(t)=v(t)r imm (t),

[0171] The binary variable v(t) serves as an indicator. If all constraints are satisfied, then v(t) = 1 and the immediate reward is set to the system utility at the current time slot t. Otherwise, if the action violates any constraint, the system will be penalized and v(t) = 0 will be set. Therefore, r(t) = 0 in these cases. imm (t) is the immediate reward the user receives when the constraint is satisfied, namely:

[0172]

[0173] This formula reflects the weighted sum of the UE's QoE, the number of cache hits, and the UE's energy consumption in each time slot.

[0174] S104 provided in the embodiment of the present invention specifically includes: based on active inference theory and the free energy principle, combined with the cumulative reward value and conditional transition probability distribution obtained in step S103, calculating the free energy of each candidate strategy; wherein the calculation formula for the inverse value of free energy is as follows:

[0175]

[0176] S105 provided in the embodiment of the present invention is specifically as follows: k minimum values ​​are screened out from the free energy of all candidate strategies and arithmetic averages are calculated for these values, and the probability distribution of the current optimal strategy is generated based on the obtained free energy average value. This process is equivalent to taking the k maximum values ​​of opposite free energy values, sorting these maximum values ​​in descending order, and then calculating the average value to determine the strategy distribution, that is:

[0177]

[0178] And get the policy distribution:

[0179]

[0180] Here, σ(·) represents the diagonal Gaussian distribution, a typical multivariate probability distribution characterized by the independence of its dimensions. This distribution is determined by two key parameters: the mean vector and the variance vector on the diagonal. In practical applications, the diagonal Gaussian distribution is often used as a prior distribution or a posteriori distribution. Its advantage is that it can effectively prevent model overfitting while simplifying complex inference processes. In the field of reinforcement learning, the diagonal Gaussian distribution can be used to model state uncertainty. This property enables intelligent agents to make more reasonable decisions in uncertain environments. By quantifying state uncertainty, intelligent agents can achieve a better balance between exploration and exploitation, thereby improving learning effects and decision quality.

[0181] In S106 provided by the embodiment of the present invention, we first perform strategy sampling based on the obtained strategy distribution δ(π), and then further extract specific actions a from the sampled strategies. t , as the agent's decision-making behavior in the current state; this behavior is then executed by the agent, interacting with the environment in real time, thereby leading to the next state of system evolution and obtaining corresponding environmental feedback (such as reward value); the entire extraction and execution process is as follows:

[0182] π t ~δ(π),a t ~π t ;

[0183] Implementation process of S108: After the global network is responsible for outputting the predicted value of free energy, the error function between the predicted value and the target value is first calculated. Then, based on the error information, the back propagation algorithm combined with the gradient descent method is used to iteratively optimize the global network parameters. The loss function in this optimization process is given by the following formula:

[0184]

[0185] Among them, Q(s t ,a t; θ) is the predicted value of the global network output free energy, is the target free energy; the gradient of the loss function is given by:

[0186]

[0187] Then update the parameters of the global network by gradient descent as follows:

[0188]

[0189] Among them, θ is the internal parameter of the global model network, and lr is the learning rate.

[0190] like Figure 2 As shown, an edge computing and cache-enabled metaverse system and intelligent optimization system provided by an embodiment of the present invention include:

[0191] System Initialization Module: Responsible for initializing the parameters of the deep deterministic policy gradient algorithm. This module consists of two parts: a configuration module for setting key hyperparameters such as the network learning rate, discount factor, and replay buffer size; and a network construction module for defining basic parameters such as the environment state dimension, user location information, and network topology.

[0192] Agent module: In each iteration, the agent module generates optimized actions based on the current network state. This module uses an active inference mechanism and the principle of free energy minimization to calculate and select the minimum value among the maximum preceding free energies, and uses the mean of these minimum values ​​as a benchmark to fit the optimal strategy distribution.

[0193] Action execution module: responsible for executing content caching strategies and computing task offloading decisions, mapping optimization strategies to specific system behaviors;

[0194] Reward Acquisition Module: Calculates the immediate reward value after executing an action. This module comprehensively considers the user experience quality and system overhead of all terminal devices in the Metaverse system, and performs quantitative evaluation based on a preset reward function.

[0195] State transfer module: After the agent performs an action, it transfers the system state from the current state to the next state and records the relevant state transfer information;

[0196] Experience replay module: used to store key information such as system status, executed actions, immediate rewards, cumulative rewards, and next state during each interaction process, forming experience tuples to provide data support for subsequent model training;

[0197] Sampling module: includes the strategy distribution sampling unit and the strategy action sampling unit, which are responsible for sampling candidate strategies from the strategy distribution and extracting specific actions from the selected strategy respectively;

[0198] Global network update module: uses the back-propagation algorithm to optimize the global network. Its core is the parameter optimization unit, which uses the gradient descent method to adjust the network parameters to minimize the loss function, thereby improving the overall performance of the system.

[0199] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the edge computing and cache-enabled metaverse system and intelligent optimization method.

[0200] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the edge computing and cache-enabled metaverse system and intelligent optimization method.

[0201] Another object of the present invention is to provide an information data processing terminal, which is used to implement the edge computing and cache-enabled metaverse system and intelligent optimization system.

[0202] The present invention is specifically implemented:

[0203] In order to effectively solve the content caching, task offloading and resource allocation problems in the metaverse system, the problem becomes extremely challenging considering the dynamic characteristics of the network and the complexity of the QoE of user access to content, the number of cache hits and system energy consumption. The present invention models the problem as a partially observable Markov decision process. Since the variables involved (such as the cache decision of the base station, the resource allocation ratio, etc.) are both continuous and discrete, a deep reinforcement learning algorithm based on active inference is used. This is an algorithm that combines active inference methods to solve complex decision-making problems and can support real-time online decision-making.

[0204] As an emerging technology, multi-access edge computing has been widely used in multiple industrial fields, especially in the metaverse field, providing effective technical support for solving related problems.

[0205] The following are three specific application examples of the present invention in the “Metaverse System Optimization Method for Joint Content Caching, Task Offloading, and Resource Allocation Based on Active Inference”:

[0206] Application Example 1: Cross-platform virtual reality collaborative conference system

[0207] Application Scenario: In a multi-user collaboration scenario within the Metaverse, participants globally access a virtual meeting space through VR devices, enabling simultaneous 3D model editing, streaming media sharing, and real-time voice interaction. The rendering delays and data transmission conflicts caused by massive user concurrency require millisecond-level response times from edge computing systems.

[0208] Implementation steps:

[0209] 1. System Initialization: Deploy a virtual conference server cluster in the cloud and build a collaborative architecture with regional edge nodes. Set core parameters for the user's VR device, including rendering quality thresholds, transmission bandwidth limits, and model compression rates.

[0210] 2. Policy generation and execution: Based on differences in user device performance and environmental network fluctuations, a set of candidate policies for task offloading is dynamically generated to allocate rendering computing resources, audio and video stream cache space, and bandwidth allocation plans to edge nodes.

[0211] 3. Free energy calculation and update: Collect user interaction data in real time (such as operation delay, screen frame rate, and voice synchronization error), quantify the strategy execution effect through a multi-objective free energy model, and give priority to the delay-energy consumption Pareto optimal strategy.

[0212] 4. Content caching and task offloading: Pre-rendering and caching of frequently called 3D model components to edge nodes, offloading real-time voice processing tasks to the nearest heterogeneous computing unit, and dynamically adjusting data compression rates to balance image quality and bandwidth consumption.

[0213] Application Example 2: Industrial Digital Twin Simulation Optimization Platform

[0214] Application scenarios: In the metaverse industrial scenario, the digital twin of the production line based on the physics engine needs to process sensor data streams and simulation operation requests in real time, and support multiple engineers to collaborate on process optimization and fault prediction.

[0215] Implementation steps:

[0216] 1. System initialization: Build a factory-level digital twin edge computing architecture, linking physical devices with virtual production line models. Set constraints such as physical simulation accuracy thresholds, data synchronization cycles, and fault detection response times.

[0217] 2. Strategy generation and execution: Based on the computing power requirements of different processes, a three-level strategy set is generated, including local computing, edge offloading, and hybrid execution, respectively configuring the parallel computing power of the GPU cluster and the real-time stream processing capabilities of the edge nodes.

[0218] 3. Free energy calculation and update: Monitor simulation error rates, data synchronization delays, and computing node load indicators, and dynamically optimize strategy distribution through adaptive free energy weights to ensure that sudden fault detection tasks receive priority computing resources.

[0219] 4. Content caching and task offloading: Distributed caching is implemented for frequently called production line model components, complex fluid dynamics simulation tasks are disassembled and offloaded to multiple edge nodes for parallel processing, and key quality inspection data is transmitted back to the cloud for persistent storage in real time.

[0220] Application Example 3: Large-Scale Immersive Metaverse Game Ecosystem

[0221] Application scenarios: In a holographic game scenario with more than 10,000 people online simultaneously, dynamic terrain reconstruction driven by player behavior, swarm intelligence of non-player characters driven by artificial intelligence, and physical special effects rendering place extremely high concurrent processing demands on the edge computing system.

[0222] Implementation steps:

[0223] 1. System Initialization: Deploy globally distributed game edge nodes, define player perspective rendering priorities, AI behavior tree computational load, and dynamic event triggering rules. Establish a player device performance gradient model and QoE evaluation matrix.

[0224] 2. Strategy Generation and Execution: Based on the player distribution heat map and real-time network topology, a candidate strategy set for regional computing load balancing is generated, and the heterogeneous computing power of each edge node (such as ray tracing rendering, collision detection, and non-player character decision tree calculations) is dynamically allocated.

[0225] 3. Free energy calculation and update: Taking into account player operation delays, screen refresh stability, and server frame synchronization errors, a reinforcement learning-based free energy optimizer is used to dynamically adjust strategy weights to ensure that high-value players (such as participants in competitive events) receive priority resource guarantees.

[0226] 4. Content caching and task offloading: Pre-load frequently accessed game scene resources on edge nodes, offload physical collision detection tasks in player-dense areas to the nearest FPGA acceleration unit, and perform path planning calculations for AI non-player characters in a hierarchical manner based on player interaction frequency.

[0227] These three examples demonstrate the specific applications of multi-access edge computing and active inference-based deep reinforcement learning in Metaverse multi-user collaboration scenarios, Metaverse industrial scenarios, and large-scale immersive Metaverse gaming ecosystems, as well as the technical advantages and value they bring. As these technologies continue to develop and mature, they will be widely applied and promoted in more Metaverse-related industries.

[0228] An embodiment of the present invention provides a MEC and cache-enabled metaverse system and intelligent optimization method, including the following steps:

[0229] Step 1: Construct a system initialization module and an agent module. The system initialization module includes a configuration module and a network construction module. The agent module uses active inference mechanism and free energy principle to fit the strategy distribution.

[0230] Step 2: Initialize the parameters in the system initialization module, including network parameters, number of rounds, number of generations, and network layout parameters such as the amount of background content;

[0231] Step 3: In the agent module, an action is generated based on the current network state at the beginning of each cycle.

[0232] Step 4: Execute task offloading and resource allocation strategies through the action execution module;

[0233] Step 5: Use the reward acquisition module to execute the action and calculate the immediate reward. The reward is calculated based on the QoE of all devices in the system and the energy consumption of cache hits. At the same time, the system state is transferred from the current state to the next state.

[0234] Step 6: The experience replay module stores the experience tuples of each system state, action, reward, cumulative reward, and next state;

[0235] Step 7, the sampling module samples from the distribution of strategies and samples from the strategies;

[0236] Step 8: Using the parameter optimization unit in the global network update module, the network parameters are adjusted by using the back propagation algorithm and the gradient descent method;

[0237] Step nine: determine whether the network has converged. If so, the final optimal solution is obtained. Otherwise, start again from step three.

[0238] This invention provides a MEC and cache-enabled metaverse system and intelligent optimization method, which optimizes content caching and resource allocation through active reasoning and the principle of free energy. Its working principle is as follows:

[0239] First, initialize the agent's global network parameters. These parameters include the number of rounds, maximum number of iterations, number of candidate strategies, planning horizon, number of users, and number of optimal candidate strategies. Furthermore, initialize the strategy distribution and transition probability distribution, as well as hyperparameters (such as the QoE threshold). These parameter initializations provide the foundation for the entire system's training process, ensuring the agent can perform strategy evaluation and optimization in subsequent operations.

[0240] Second, candidate policies are generated and the state is initialized. At the beginning of each round, the agent sets its initial state and randomly samples J candidate policies from the policy distribution. These candidate policies represent a set of possible actions the agent can take in the current environment. These candidate policies provide input for subsequent action sampling and the calculation of the conditional transition probability distribution, while also enabling the agent to maintain a certain level of diversity and exploration in each iteration.

[0241] Third, actions are sampled based on the candidate strategy and rewards are calculated. By randomly sampling the candidate strategy, a corresponding set of actions is obtained. These actions are used to calculate the conditional transition probability distribution, and rewards are generated based on the interaction between the current state and action and the environment. These rewards are important indicators for measuring the quality of the strategy and directly reflect its performance in content caching and resource allocation.

[0242] Fourth, optimize the strategy distribution using the free energy principle. After obtaining the cumulative rewards and conditional transition probability distributions for all candidate strategies, the free energy value of each candidate strategy is calculated using the free energy principle. The top k strategies with the smallest free energy values ​​are averaged to generate a new strategy distribution. This step optimizes the strategies by minimizing free energy, ensuring that the agent gradually converges to the globally optimal strategy.

[0243] Fifth, environment interaction and experience storage. Based on the optimized policy distribution, the agent further samples actions to interact with the environment, obtaining the next state and simultaneously storing the experience tuple in a replay buffer. If the buffer is full, the oldest stored experience is overwritten with the new one. This experience replay mechanism effectively avoids overfitting and improves training efficiency and model generalization.

[0244] Sixth, the global network parameters are updated and training is iteratively conducted. The global network outputs a predicted free energy value and calculates the error function between it and the actual free energy. The global network parameters are updated using backpropagation and gradient descent. The entire training process is repeated over multiple rounds until the algorithm converges, ultimately generating a policy distribution. Using this policy distribution, the agent can randomly select the optimal policy at each step, thereby optimizing joint content caching and resource offloading.

[0245] This method uses active reasoning and the free energy principle to achieve intelligent control of content caching and resource allocation in complex environments, greatly improving the service quality and resource utilization efficiency of the metaverse system.

[0246] Figure 3 This is a scenario diagram of the system involved in the present invention. The system includes a base station, multiple mobile user devices, and a metaverse server deployed at a remote end. The edge server is deployed on the base station to form a metaverse system to provide services for the user devices. The number and set of mobile user devices are represented as I and The present invention divides time into multiple time slots and regards each time slot as a discrete time unit for executing tasks and processing communication processes. At the beginning of each time slot, each user device creates a service request, which includes a content request and a local computing task. Due to the limited processing power of the user device, it may be necessary to offload the local computing task to the base station for processing. In order to optimize the QoE and overhead of the user device obtaining content, the present invention considers caching part of the content on the base station and dynamically adjusting the caching decision and resource allocation according to the user request and system status. In the relevant scenario, there are K different environments and objects, where the data set is represented by. Among them, C f (t+1) is the amount of analog signal data of the environment or object; D f (t+1) is the amount of digital signal data of the environment or object; (t+1) is the CPU cycles required for processing the corresponding data;

[0247] like Figure 4 As shown, the method for collaborative content caching decision, task offloading decision and resource allocation of the present invention includes the following steps:

[0248] S101, initializing the agent global network parameters;

[0249] Initialize the agent's global network parameters, iteration number N ep , maximum number of iterations N, number of candidate strategies J, planning horizon H, number of rounds T max , the number of users N, the integrated network parameter θ1, the reward network parameter θ2, the number of optimal candidate strategies k, etc. At the same time, the strategy distribution η(π) and the transition probability distribution δ(s t |s t-1 ,θ,π), etc., to provide initial conditions for subsequent optimization algorithms;

[0250] S102, candidate strategy generation;

[0251] At the beginning of each round, set the initial state s t , and randomly select J candidate strategies from the initial strategy distribution as the strategy set to be evaluated in the current round;

[0252] S103, action sampling and reward calculation;

[0253] Perform action sampling on each candidate strategy extracted in S102 to obtain J actions. At the same time, calculate the conditional transition probability distribution corresponding to each action, and calculate the corresponding immediate reward based on the result of the interaction between the action and the environment.

[0254] S104, free energy calculation;

[0255] According to active inference theory, the free energy value of each candidate strategy is calculated using the rewards and conditional transition probability distributions obtained from interacting with the environment to evaluate the quality of the strategy.

[0256] S105, strategy distribution update;

[0257] Select k candidate strategies with the smallest free energy values, average their free energy values, and update the strategy distribution based on the average value, so that the strategy distribution gradually converges to a better area;

[0258] S106, action execution and status update;

[0259] Based on the updated policy distribution in S105, a policy is sampled, and an action is further sampled from the policy as the agent's current action. The agent executes the action and interacts with the environment to obtain the next state;

[0260] S107, experience storage;

[0261] The current experience tuple (including state, action, reward, next state, etc.) is stored in the replay buffer. If the replay buffer is full, the first-in-first-out principle is adopted to overwrite the oldest stored experience with the new experience;

[0262] S108, network parameter update;

[0263] Using the error function between the predicted free energy value output by the global network and the actual free energy value, the back propagation algorithm and gradient descent method are used to update the global network parameters to improve the prediction accuracy and strategy optimization effect;

[0264] S109, iterative convergence;

[0265] Repeat the above steps until the algorithm converges. Finally, an optimized policy distribution is obtained. Based on this distribution, the agent can randomly extract policies to achieve optimal control of joint content caching, task offloading, and resource allocation, thereby maximizing system performance.

[0266] Furthermore, in S102, at the beginning of each round, the initial state is set; J candidate strategies are randomly sampled from the strategy distribution, and then J actions are randomly sampled from each of the J strategies; the state of the agent is represented as:

[0267]

[0268] in: in is the number of CPU cycles required for user i to render the front-end interaction in time slot t; in is the data size of the front-end interactive information of user i in time slot t; b(t) = {bi (t)}, where b i (t) is the index of the background model requested by user i in time slot t; in is the set of background objects requested by user i in time slot t; q(t) = {q i (t)}, where q i (t) is the resolution level required by user i to request content from the metaverse in time slot t; R b (t) is the wireless transmission rate between the base station and the remote Metaverse server in time slot t; is the set of models cached in the MEC server in time slot t-1; is the set of objects of model s cached in the MEC server in time slot t-1; rm (t-1) is the remaining available cache resources of the MEC server after the content of time slot t-1 is placed; Life md (t)=Life s (t), is the life cycle of model s in time slot t; Life obj (t)=Life s,l (t), is the life cycle of object l of model s in time slot t; V md (t)={v s (t)}, is the removal indicator of model s in time slot t; V obj (t)={v s,l (t)}, is the removal indicator of object l of model s in time slot t. Self-policy π t The sampled action is represented as:

[0269]

[0270] Among them, the model's cache decision κ(t) = {κ s (t)}, The cache decision of the object ζ(t) = {ζ s,l (t)}, Offloading decision of UE-side interaction Computing resource allocation in MEC servers mec (t) = {f i mec (t)}, Computational resource allocation in UE loc (t) = {f i loc (t)}, Wireless transmission rate allocation R(t) = {R i (t)},

[0271] Furthermore, in step S103, for each candidate strategy, the conditional transition probability distribution is calculated; and based on the actions corresponding to each candidate strategy, the instant reward value at the current moment is calculated respectively. The calculation formula of the instant reward is as follows:

[0272] r(t)=v(t)r imm (t),

[0273] Among them, the binary variable v(t) serves as an indicator. If all constraints are satisfied, then v(t) = 1 and the immediate reward is set to the system utility at the current time slot t; otherwise, if the action violates any constraint, the system will be penalized and v(t) = 0 will be set, so r(t) = 0 in these cases. imm (t) is the immediate reward the user receives when the constraint is satisfied, namely:

[0274]

[0275] This formula reflects the weighted sum of UE QoE, cache hit count, and UE energy consumption in each time slot;

[0276] Furthermore, in step S104, based on active inference theory and the free energy principle, the free energy of each candidate strategy is calculated in combination with the cumulative reward value and conditional transition probability distribution obtained in step S103; wherein the calculation formula for the inverse value of free energy is as follows:

[0277]

[0278] Furthermore, in S105, k minimum values ​​are screened out from the free energies of all candidate strategies and the arithmetic average is calculated for these values. Based on the obtained average free energy value, the probability distribution of the current optimal strategy is generated. This process is equivalent to taking the k maximum values ​​of opposite free energy values, sorting these maximum values ​​in descending order, and then calculating the average value to determine the strategy distribution, that is:

[0279]

[0280] And get the policy distribution:

[0281]

[0282] Here, σ(·) represents the diagonal Gaussian distribution, which is a typical multivariate probability distribution characterized by the independence of each dimension. The distribution is determined by two key parameters: the mean vector and the variance vector on the diagonal. In practical applications, the diagonal Gaussian distribution is often used as a prior distribution or a posterior distribution. Its advantage is that it can effectively prevent model overfitting while simplifying complex inference processes. In the field of reinforcement learning, the diagonal Gaussian distribution can be used to model state uncertainty. This feature enables intelligent agents to make more reasonable decisions in uncertain environments. By quantifying the uncertainty of the state, the intelligent agent can achieve a better balance between exploration and exploitation, thereby improving learning effects and decision-making quality;

[0283] S106, firstly, performs strategy sampling based on the obtained strategy distribution δ(π), and then further extracts specific actions a from the sampled strategies. t , as the agent's decision-making behavior in the current state. This behavior is then executed by the agent, interacting with the environment in real time, thereby leading to the next state of system evolution and obtaining corresponding environmental feedback (such as reward value). The entire extraction and execution process is as follows:

[0284] π t ~δ(π),a t ~π t ;

[0285] Furthermore, in step S108, after the global network outputs the predicted free energy value, the error function between the predicted value and the target value is first calculated. Then, based on the error information, the global network parameters are iteratively optimized using a backpropagation algorithm combined with a gradient descent method. The loss function in this optimization process is given by the following formula:

[0286]

[0287] Among them, Q(s t ,a t ; θ) is the predicted value of the global network output free energy, is the target free energy. The gradient of the loss function is given by:

[0288]

[0289] Then update the parameters of the global network by gradient descent as follows:

[0290]

[0291] Among them, θ is the internal parameter of the global model network, lr is the learning rate;

[0292] In order to prove the creativity and technical value of the technical solution of the present invention, this section provides application examples of the claimed technical solution on specific products or related technologies.

[0293] A computer device supporting metaverse scenarios includes a high-performance computing unit and a massive storage system. The storage unit is composed of a non-volatile storage medium and is embedded with an executable code set for implementing an active reasoning optimization method in a MEC and cache-enabled metaverse system. The processing unit includes a multi-core heterogeneous computing architecture and can fully implement the active reasoning-based resource allocation optimization method by parsing and executing the program code in the storage unit, specifically involving a closed-loop control process of virtual environment state perception, free energy calculation, strategy generation, and action execution.

[0294] A computer-readable storage medium for Metaverse services, characterized in that the medium utilizes a distributed storage cluster architecture and contains program code for the optimization method. When the code is loaded by a processor with multi-core computing capabilities, it triggers processing steps including but not limited to: multi-dimensional state space modeling, policy network gradient updates, virtual entity behavior decision-making mechanisms, and a dynamic reward feedback system based on QoE.

[0295] The Metaverse Information Processing Terminal is characterized by: It integrates a high-performance environmental perception unit and a task scheduling center, implementing the aforementioned optimization system through a cloud-edge collaborative architecture; it includes an immersive interactive interface module that can collect user actions and virtual environment data in real time; it is equipped with a distributed computing interface module for coordinating physical layer computing resources with virtual layer service requirements; and it also deploys an autonomous decision-making engine module with a built-in policy network inference component and a free energy optimization core to support low-latency, highly reliable intelligent operation and maintenance of Metaverse scenarios.

[0296] Example 1: Cross-platform virtual reality collaborative conference system

[0297] In a multi-user collaboration scenario within the Metaverse, participants globally access a virtual meeting space through VR devices, enabling simultaneous 3D model editing, streaming media sharing, and real-time voice interaction. The rendering delays and data transmission conflicts caused by this massive user concurrency require edge computing systems to achieve millisecond-level response times.

[0298] 1) System Initialization: Deploy a virtual conference server cluster in the cloud and build a collaborative architecture with regional edge nodes. Set core parameters such as the rendering quality threshold, transmission bandwidth limit, and model compression rate for the user's VR device.

[0299] 2) Policy generation and execution: Based on the performance differences of user devices and environmental network fluctuations, a set of candidate policies for task offloading is dynamically generated to allocate rendering computing resources, audio and video stream cache space, and bandwidth allocation solutions to edge nodes.

[0300] 3) Free energy calculation and update: Collect user interaction data (such as operation delay, screen frame rate, and voice synchronization error) in real time, quantify the strategy execution effect through a multi-objective free energy model, and give priority to the delay-energy consumption Pareto optimal strategy.

[0301] 4) Content caching and task offloading: Pre-rendering and caching of frequently called 3D model components to edge nodes, offloading real-time voice processing tasks to the nearest heterogeneous computing unit, and dynamically adjusting data compression rates to balance image quality and bandwidth consumption.

[0302] Example 2: Industrial Digital Twin Simulation Optimization Platform

[0303] In the metaverse industrial scenario, the digital twin of the production line based on the physics engine needs to process sensor data streams and simulation calculation requests in real time, and support multiple engineers to collaborate on process optimization and fault prediction.

[0304] 1) System Initialization: Build a factory-level digital twin edge computing architecture, linking physical devices with virtual production line models. Set constraints such as physical simulation accuracy thresholds, data synchronization cycles, and fault detection response times.

[0305] 2) Strategy generation and execution: Based on the computing power requirements of different processes, a three-level strategy set is generated, including local computing, edge offloading, and hybrid execution, which respectively configures the parallel computing power of the GPU cluster and the real-time stream processing capabilities of the edge nodes.

[0306] 3) Free energy calculation and update: Monitor simulation error rates, data synchronization delays, and computing node load indicators, and dynamically optimize strategy distribution through adaptive free energy weights to ensure that sudden fault detection tasks receive priority computing resources.

[0307] 4) Content caching and task offloading: Distributed caching is implemented for frequently called production line model components, complex fluid dynamics simulation tasks are disassembled and offloaded to multiple edge nodes for parallel processing, and key quality inspection data is transmitted back to the cloud for persistent storage in real time.

[0308] Example 3: Large-Scale Immersive Metaverse Gaming Ecosystem

[0309] In a holographic game scenario with more than 10,000 people online at the same time, dynamic terrain reconstruction driven by player behavior, group intelligence of artificial intelligence non-player characters, and physical special effects rendering place extremely high concurrent processing demands on the edge computing system.

[0310] 1) System Initialization: Deploy globally distributed game edge nodes, define player perspective rendering priorities, AI behavior tree computational load, and dynamic event triggering rules. Establish a player device performance gradient model and QoE evaluation matrix.

[0311] 2) Strategy Generation and Execution: Based on the player distribution heat map and real-time network topology, a candidate strategy set for regional computing load balancing is generated, and the heterogeneous computing power of each edge node (such as ray tracing rendering, collision detection, and non-player character decision tree calculations) is dynamically allocated.

[0312] 3) Free Energy Calculation and Update: Taking into account player operation latency, screen refresh stability, and server frame synchronization errors, a reinforcement learning-based free energy optimizer is used to dynamically adjust strategy weights to ensure that high-value players (such as participants in competitive events) receive priority resource protection.

[0313] 4) Content caching and task offloading: Pre-loading of frequently accessed game scene resources is implemented on edge nodes, and physical collision detection tasks in player-dense areas are offloaded to the nearest FPGA acceleration unit. Path planning calculations for AI non-player characters are processed in a hierarchical manner based on the frequency of player interactions.

[0314] Figure 5 This is a comparison diagram of user QoE obtained when different base station storage capacities are used according to an embodiment of the present invention and existing collaborative content caching and replacement decision-making, computation offloading decision-making, user resource allocation method, and transmission power control.

[0315] Figure 6 This is a user-side energy consumption comparison diagram of different user processing powers provided by an embodiment of the present invention and existing collaborative content caching and replacement decision-making, computation offloading decision-making, user resource allocation method and transmission power control.

[0316] Figure 7 This is a comparison chart of the number of user cache hits when a single model contains different numbers of objects, provided by an embodiment of the present invention and existing collaborative content caching and replacement decision-making, computational offloading decision-making, user resource allocation method, and transmission power control.

[0317] Figure 5 The storage capacity of the base station is set to 0.001×10 9 bit, 2×10 9 bit, 4×10 9 bit, 6×10 9 bit. When the storage capacity is small, the base station can cache very little content, and users often need to obtain content from the remote Metaverse server, resulting in a very low QoE. After the algorithm converges, due to the increase in storage capacity, the base station can cache more and more environments and objects. As a result, the need to download background information from the remote Metaverse server decreases, which will lead to a reduction in the total end-to-end latency. Clearly, the QoE value increases with the increase in the π value. In addition, the performance of the proposed algorithm is significantly better than other baseline algorithms.

[0318] Figure 6 The relationship between user segment energy consumption and user local processing power after algorithm convergence is described. As the processing power of user devices gradually increases, their energy consumption levels also rise accordingly. The figure clearly shows that compared with the joint optimization algorithm proposed in this article, the model-free caching strategy, the object-free caching mechanism, and the DDPG algorithm all exhibit higher system energy consumption. The energy consumption performance of the average mobile edge computing resource allocation scheme and the average transmission rate allocation strategy is similar to that of the algorithm in this study. This is because when processing power is dynamically adjusted, the computing resource allocation decision and maximum transmission rate setting of the edge server mainly affect the base station's computing processing delay and data transmission delay, and have a relatively limited direct impact on system energy consumption-related indicators. It is worth noting that the intelligent algorithm designed in this study has demonstrated significant energy consumption optimization advantages when facing user device systems with different processing powers. Its system energy consumption minimization effect surpasses all benchmark comparison schemes, including the traditional deep reinforcement learning algorithm DDPG.

[0319] Figure 7 The relationship between the user cache hit rate and the number of different objects contained in a single model after the algorithm converges is described. The figure sets the number of objects contained in a single model to 1, 5, 10, and 15, respectively. When each model contains only a single object, the probability of the object being cached in the base station is high, and thus when the user device requests the object, the probability of a cache hit is significantly increased. However, as the number of objects in the model increases, the probability of a user device request for a specific object in the model triggering a cache hit gradually decreases, which is consistent with the trend shown in the figure. In the case where the base station does not cache any models or objects, all user device requests for models and objects must be downloaded remotely from the Metaverse server, so the cache hit count of the two algorithms in the figure is always zero. Overall, compared with other baseline algorithms, the algorithm proposed in this study shows a significant advantage in the key indicator of cache hit count.

[0320] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0321] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. An edge computing and cache-enabled metaverse system and intelligent optimization method, characterized in that: The following steps are involved: S101, initializing the agent global network parameters; Initialize the agent's global network parameters, iteration number N ep , maximum number of iterations N, number of candidate strategies J, planning horizon H, number of rounds T max , the number of users N, the integrated network parameter θ1, the reward network parameter θ2, the number of optimal candidate strategies k, etc.; at the same time, the initialization strategy distribution η(π) and the initialization transition probability distribution δ(s t |s t-1 ,θ,π), etc., to provide initial conditions for subsequent optimization algorithms; S102, candidate strategy generation; At the beginning of each round, set the initial state s t , and randomly select J candidate strategies from the initial strategy distribution as the strategy set to be evaluated in the current round; S103, action sampling and reward calculation; Perform action sampling on each candidate strategy extracted in S102 to obtain J actions. At the same time, calculate the conditional transition probability distribution corresponding to each action, and calculate the corresponding immediate reward based on the result of the interaction between the action and the environment. S104, free energy calculation; According to active inference theory, the free energy value of each candidate strategy is calculated using the rewards and conditional transition probability distributions obtained from interacting with the environment to evaluate the quality of the strategy. S105, strategy distribution update; Select k candidate strategies with the smallest free energy values, average their free energy values, and update the strategy distribution based on the average value, so that the strategy distribution gradually converges to a better area; S106, action execution and status update; Based on the updated policy distribution in S105, a policy is sampled, and an action is further sampled from the policy as the current action of the agent; the agent executes the action and interacts with the environment to obtain the next state; S107, experience storage; Store the current experience tuple including state, action, reward, and next state into the replay buffer; If the replay buffer is full, the first-in-first-out principle is adopted to overwrite the oldest stored experience with the new experience; S108, network parameter update; Using the error function between the predicted free energy value output by the global network and the actual free energy value, the back propagation algorithm and gradient descent method are used to update the global network parameters to improve the prediction accuracy and strategy optimization effect; S109, iterative convergence; Repeat the above steps until the algorithm converges; eventually, an optimized policy distribution is obtained. The agent can randomly extract policies based on this distribution to achieve optimal control of joint content caching, task offloading, and resource allocation, thereby maximizing system performance.

2. The edge computing and cache-enabled metaverse system and intelligent optimization method according to claim 1, characterized in that: In S102, at the beginning of each round, the system initializes and sets the initial value of the current state; then, J candidate strategies are randomly selected from the strategy distribution, and corresponding actions are further randomly selected from each candidate strategy; wherein, the state of the agent is represented as a set of the following multi-dimensional vectors: in: in is the number of CPU cycles required for user i to render the front-end interaction in time slot t; in is the data size of the front-end interaction information of user i in time slot t; where b i (t) is the index of the background model requested by user i in time slot t; in is the set of background objects requested by user i in time slot t; where q i (t) is the resolution level required by user i to request content from the metaverse in time slot t; R b (t) is the wireless transmission rate between the base station and the remote Metaverse server in time slot t; is the set of models cached in the MEC server in time slot t-1; is the set of objects of model s cached in the MEC server in time slot t-1; rm (t-1) is the remaining available cache resources of the MEC server after the content of time slot t-1 is placed; Life md (t)=Life s (t), is the life cycle of model s in time slot t; Life obj (t)=Life s,l (t), is the life cycle of object l of model s in time slot t; V md (t)={v s (t)}, is the removal indicator of model s in time slot t; V obj (t)={v s,l (t)}, is the removal indicator of object l of model s in time slot t; self-policy π t The sampled action is represented as: Among them, the model's cache decision κ(t) = {κ s (t)}, The cache decision of the object ζ(t) = {ζ s,l (t)}, Offloading decision of UE-side interaction Computing resource allocation in MEC servers Computing resource allocation in UE Wireless transmission rate allocation 3. The edge computing and cache-enabled metaverse system and intelligent optimization method according to claim 1, characterized in that: S103 specifically includes: calculating the conditional transition probability distribution for each candidate strategy; calculating the instant reward value at the current moment according to the actions corresponding to each candidate strategy; the calculation formula of the instant reward is as follows: r(t)=v(t)r imm (t), The binary variable v(t) serves as an indicator. If all constraints are satisfied, then v(t) = 1 and the immediate reward is set to the system utility at the current time slot t. Otherwise, if the action violates any constraint, the system will be penalized and v(t) = 0 will be set. Therefore, r(t) = 0 in these cases. imm (t) is the immediate reward the user receives when the constraint is satisfied, namely: This formula reflects the weighted sum of the UE's QoE, the number of cache hits, and the UE's energy consumption in each time slot.

4. The edge computing and cache-enabled metaverse system and intelligent optimization method according to claim 1, characterized in that: S104 specifically includes: calculating the free energy of each candidate strategy based on active inference theory and the free energy principle, combined with the cumulative reward value and conditional transition probability distribution obtained in step S103; wherein the calculation formula for the inverse value of free energy is as follows:

5. The edge computing and cache-enabled metaverse system and intelligent optimization method according to claim 1, characterized in that: S105 specifically includes: selecting k minimum values ​​from the free energy of all candidate strategies and calculating the arithmetic average thereof, and generating the probability distribution of the current optimal strategy based on the obtained free energy average value; this process is equivalent to taking the k maximum values ​​of opposite free energy values, sorting these maximum values ​​in descending order, and then calculating the average value to determine the strategy distribution, that is: And get the policy distribution: Here, σ(·) represents the diagonal Gaussian distribution, a typical multivariate probability distribution characterized by the independence of its dimensions. This distribution is determined by two key parameters: the mean vector and the variance vector on the diagonal. In practical applications, the diagonal Gaussian distribution is often used as a prior distribution or a posteriori distribution. Its advantage is that it can effectively prevent model overfitting while simplifying complex inference processes. In the field of reinforcement learning, the diagonal Gaussian distribution can be used to model state uncertainty. This property enables intelligent agents to make more reasonable decisions in uncertain environments. By quantifying state uncertainty, intelligent agents can achieve a better balance between exploration and exploitation, thereby improving learning effects and decision quality.

6. The edge computing and cache-enabled metaverse system and intelligent optimization method according to claim 1, characterized in that: In S106, we first perform strategy sampling based on the obtained strategy distribution δ(π), and then further extract specific actions a from the sampled strategies. t , as the agent's decision-making behavior in the current state; this behavior is then executed by the agent, interacting with the environment in real time, thereby leading to the next state of system evolution and obtaining corresponding environmental feedback (such as reward value); the entire extraction and execution process is as follows: p t ~δ(π),a t ~π t ; Implementation process of S108: After the global network is responsible for outputting the predicted value of free energy, the error function between the predicted value and the target value is first calculated. Then, based on the error information, the back propagation algorithm combined with the gradient descent method is used to iteratively optimize the global network parameters. The loss function in this optimization process is given by the following formula: Among them, Q(s t ,a t ; θ) is the predicted value of the global network output free energy, is the target free energy; the gradient of the loss function is given by: Then update the parameters of the global network by gradient descent as follows: θ←θ-lr·▽ θ L(θ), Among them, θ is the internal parameter of the global model network, and lr is the learning rate.

7. A metaverse system and intelligent optimization system that implements the edge computing and cache-enabled optimization method according to any one of claims 1 to 6, characterized in that: The system comprises: System Initialization Module: Responsible for initializing the parameters of the deep deterministic policy gradient algorithm. This module consists of two parts: a configuration module for setting key hyperparameters such as the network learning rate, discount factor, and replay buffer size; and a network construction module for defining basic parameters such as the environment state dimension, user location information, and network topology. Agent module: In each iteration, the agent module generates optimized actions based on the current network state. This module uses an active inference mechanism and the principle of free energy minimization to calculate and select the minimum value among the maximum preceding free energies, and uses the mean of these minimum values ​​as a benchmark to fit the optimal strategy distribution. Action execution module: responsible for executing content caching strategies and computing task offloading decisions, mapping optimization strategies to specific system behaviors; Reward Acquisition Module: Calculates the immediate reward value after executing an action. This module comprehensively considers the user experience quality and system overhead of all terminal devices in the Metaverse system, and performs quantitative evaluation based on a preset reward function. State transfer module: After the agent performs an action, it transfers the system state from the current state to the next state and records the relevant state transfer information; Experience replay module: used to store key information such as system status, executed actions, immediate rewards, cumulative rewards, and next state during each interaction process, forming experience tuples to provide data support for subsequent model training; Sampling module: includes the strategy distribution sampling unit and the strategy action sampling unit, which are responsible for sampling candidate strategies from the strategy distribution and extracting specific actions from the selected strategy respectively; Global network update module: uses the back-propagation algorithm to optimize the global network. Its core is the parameter optimization unit, which uses the gradient descent method to adjust the network parameters to minimize the loss function, thereby improving the overall performance of the system.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the edge computing and cache-enabled metaverse system and intelligent optimization method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the edge computing and cache-enabled metaverse system and intelligent optimization method as described in any one of claims 1-6.

10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the edge computing and cache-enabled metaverse system and intelligent optimization system as described in claim 7.

Citation Information

Cited By

  • Metacosmic dynamic and static scene dual-target cache optimization method

    CN121000785A

  • A meta-universe static and dynamic scene dual-target cache optimization method

    CN121000785B

  • Metacosm player behavior off-line hosting method and system

    CN122020409A