Resource collaborative allocation method and device, electronic equipment and storage medium

By acquiring user visual attention data and dividing it into sub-tasks, and combining deep reinforcement learning and a two-layer Stackelberg game model, the problem of low resource allocation efficiency in existing technologies is solved. This enables refined resource allocation and a win-win business incentive mechanism, thereby improving user experience and resource utilization efficiency.

CN121967388APending Publication Date: 2026-05-01STATE GRID LIAONING ELECTRIC POWER CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID LIAONING ELECTRIC POWER CO LTD
Filing Date
2025-12-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technical solutions mostly use objective KPIs as optimization targets, without incorporating user attention and subjective experience into the model. This results in low resource utilization efficiency and poor user experience. Furthermore, they fail to build a profit-sharing and constraint mechanism that links users, platforms, and infrastructure, leading to an incomplete price signal and revenue transmission chain and making it difficult to form effective business incentives.

Method used

By acquiring users' visual attention data, the metaverse interaction task is divided into multiple subtasks with different priorities. A deep reinforcement learning algorithm is used to assign the subtasks to the optimal edge nodes under the condition of satisfying the time delay constraint. A two-layer Stackelberg game model is constructed to achieve dynamic resource collaborative allocation.

Benefits of technology

It enables refined resource allocation based on perceived value, improves resource utilization efficiency, enhances user experience, and forms a dynamic economic incentive mechanism that benefits all parties, ensuring that all parties obtain reasonable returns in the resource allocation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967388A_ABST
    Figure CN121967388A_ABST
Patent Text Reader

Abstract

The invention relates to a resource collaborative allocation method and device, electronic equipment and a storage medium, and relates to the technical field of resource allocation, and the method comprises the steps: employing a deep reinforcement learning algorithm based on a user demand vector, an experience quality value and heterogeneous resources under a unified measurement system, and obtaining a resource collaborative allocation result; assigning each subtask to an optimal edge node to obtain a task unloading decision scheme; and a deep reinforcement learning algorithm is adopted to jointly solve the constructed double-layer Stackelberg game model and the task unloading decision scheme, and dynamic resource collaborative allocation is realized. According to the scheme, the subjective experience quality value of the user is included in the consideration category of resource allocation, refined resource release based on perception value is realized, the resource utilization efficiency and the user experience are improved, benefit games and balance among the user, the platform and infrastructures are considered, and the user experience is improved. And a complete price signal and income transmission chain and effective commercial incentive are formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of resource allocation technology, specifically to a resource collaborative allocation method, apparatus, electronic device, and storage medium. Background Technology

[0002] As an evolutionary form of the next-generation Internet, the core characteristic of the metaverse is to provide a persistent, highly immersive, and low-latency shared virtual space for a large number of users. To achieve this vision, the system must process massive amounts of 3D rendering, physical simulation, and multi-user interactive data in real time.

[0003] Unlike traditional audio and video services, metaverse interaction exhibits significant attention selectivity. Users' visual focus is typically concentrated in the concave region of their field of vision, while unfocused areas show significantly lower sensitivity to quality. If resource allocation is statically or uniformly configured based solely on objective network performance metrics such as throughput and latency, it is highly likely that excessive resources will be invested in low-interest areas, while quality assurance in high-interest areas will be inadequate. This not only results in severe resource waste but also leads to an imbalance in the user's subjective experience. Therefore, in edge-cloud collaborative networks, resource allocation and pricing mechanisms must consider the diversity and spatiotemporal heterogeneity of services, and explicitly consider the differentiated requirements of user attention and subjective experience at the strategy level to avoid the dilemma that "objective optimality does not equal subjective satisfaction."

[0004] Furthermore, in the commercialization of the metaverse, in addition to the technical aspects of resource orchestration and scheduling, the economic aspects of incentives and pricing must also be addressed. Without matching, refined billing and revenue-sharing rules, cost-saving measures on the technical side will be difficult to translate into sustainable commercial revenue, and users will lack the motivation to pay for quality assurance in high-value areas.

[0005] To better depict the interactive relationship between "price-demand-supply," existing research and patents often incorporate game theory and mechanism design theory. For example, auction mechanisms are used to allocate scarce resources among multiple users and nodes; contract theory is employed to design menu-style terms for different user types to guide their genuine declarations and self-selection; or a leader-follower structure of "pricing-unloading-supply" is established based on Stackelberg games to clarify the order of price prioritization, decision response, and resource settlement.

[0006] However, after searching and analyzing existing technologies, it was found that current technical solutions still take objective KPIs such as latency and speed as the core optimization targets, ignoring the marginal value differences of different attention areas in the metaverse business. This results in resource allocation strategies failing to accurately map users' subjective perceptions, easily leading to resource misallocation in low-attention areas and insufficient protection for high-attention areas, ultimately causing low resource utilization efficiency and poor user experience.

[0007] Furthermore, most existing technological solutions focus on the bilateral game between "user-edge" or "platform-edge," failing to include users, platforms, and infrastructure providers as independent entities with interconnected interests within a unified profit-sharing and constraint framework. This results in an incomplete transmission chain of price signals and returns, making it difficult for technological resource optimization to be transformed into a verifiable and sustainable settlement and incentive loop, thus failing to support a healthy and win-win metaverse economic ecosystem. Summary of the Invention

[0008] In view of this, this application provides a resource collaborative allocation method, device, electronic device, and storage medium. The main purpose is to solve the technical problems of current solutions, which mostly take objective KPIs as optimization targets, fail to incorporate user attention and subjective experience into the model, and cannot achieve refined resource allocation based on perceived value, resulting in low resource utilization efficiency and poor user experience. Furthermore, existing solutions mostly focus on the two-party relationship of "user-edge" or "platform-edge", failing to build a profit-sharing and constraint mechanism covering the linkage of users, platforms, and infrastructure, resulting in an incomplete price signal and revenue transmission chain, making it difficult to form effective business incentives.

[0009] According to a first aspect of this application, a resource collaborative allocation method is provided, the method comprising: Acquire the user's visual attention data, divide the user's metaverse interaction task into multiple subtasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each subtask. The requirement vector of the subtask, the heterogeneous resources of the edge nodes, and the user's experience quality value are quantified into a unified measurement system to generate a unified requirement vector, experience quality value, and heterogeneous resources for resource allocation and settlement. Based on the unified demand vector, experience quality value, and heterogeneous resources, a deep reinforcement learning algorithm is used to assign each subtask to the optimal edge node under the condition of meeting the latency constraint, thereby obtaining the task offloading decision scheme for each subtask. The constructed two-layer Stackelberg game model and the task offloading decision scheme are jointly solved using the deep reinforcement learning algorithm to achieve dynamic resource collaborative allocation.

[0010] According to a second aspect of this application, a resource collaborative allocation apparatus is provided, the apparatus comprising: The acquisition module is used to acquire the user's visual attention data, divide the user's metaverse interaction task into multiple sub-tasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each sub-task. The generation module is used to quantify the demand vector of the subtask, the heterogeneous resources of the edge nodes, and the user's experience quality value into a unified measurement system, and generate a unified demand vector, experience quality value, and heterogeneous resources for resource allocation and settlement. The assignment module is used to assign each subtask to the optimal edge node based on the unified demand vector, the experience quality value, and the heterogeneous resources, using a deep reinforcement learning algorithm, under the condition of meeting the latency constraint, so as to obtain the task unloading decision scheme for each subtask. The solution module is used to jointly solve the constructed two-layer Stackelberg game model and the task unloading decision scheme using the deep reinforcement learning algorithm, so as to realize dynamic resource collaborative allocation.

[0011] According to a third aspect of this application, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method of the first aspect described above.

[0012] According to a fourth aspect of this application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect described above.

[0013] By employing the above technical solutions, this application provides a resource collaborative allocation method, apparatus, electronic device, and storage medium. Compared with existing technologies, this application can acquire user visual attention data, divide the user's metaverse interaction task into multiple sub-tasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each sub-task. The demand vectors of the sub-tasks, the heterogeneous resources of edge nodes, and the user's experience quality value are quantified into a unified measurement system to generate a unified demand vector, experience quality value, and heterogeneous resources for resource allocation and settlement. Based on the unified demand vector, experience quality value, and heterogeneous resources, a deep reinforcement learning algorithm is used to assign each sub-task to the optimal edge node under the condition of satisfying latency constraints, obtaining a task offloading decision scheme for each sub-task. A deep reinforcement learning algorithm is used to jointly solve the constructed two-layer Stackelberg game model and the task offloading decision scheme to achieve dynamic resource collaborative allocation.

[0014] The solution described in this application acquires user visual attention data and then meticulously divides the user's metaverse interaction tasks into multiple sub-tasks with different priorities based on this data. This approach introduces user attention factors into the task processing flow, matching the importance of different sub-tasks with the user's level of attention, thus laying the foundation for subsequent resource allocation based on user perceived value.

[0015] By assigning a demand vector containing resource requirements to each subtask, and quantifying the subtask demand vectors, heterogeneous resources of edge nodes, and user experience quality values ​​into a unified measurement system, a unified demand vector, experience quality value, and heterogeneous resources are generated for resource allocation and settlement. This quantified and unified approach incorporates the user's subjective experience quality value into the resource allocation considerations, enabling resource allocation to fully take into account the user's subjective feelings. This achieves refined resource allocation based on perceived value, thereby improving resource utilization efficiency and enhancing user experience.

[0016] Based on a unified demand vector, experience quality score, and heterogeneous resources, a deep reinforcement learning algorithm is employed to assign each subtask to the optimal edge node while satisfying latency constraints, thus obtaining a task offloading decision scheme for each subtask. The deep reinforcement learning algorithm can dynamically adjust the task offloading strategy according to the constantly changing environment and user needs, further optimizing resource allocation and ensuring that resource utilization efficiency is improved while meeting user needs, thereby enhancing the user experience.

[0017] This application employs a deep reinforcement learning algorithm to jointly solve a constructed two-layer Stackelberg game model and a task offloading decision scheme. This joint solution approach ensures that the resource allocation process not only considers the optimization of task offloading but also takes into account the game and balance of interests among users, the platform, and infrastructure. This method enables the formation of a complete price signal and benefit transmission chain, allowing all parties to obtain reasonable benefits during resource allocation, thereby creating effective business incentives and promoting the healthy development of the entire system.

[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A flowchart illustrating a resource collaborative allocation method provided in an embodiment of this application is shown; Figure 2 This paper illustrates a system architecture diagram of a resource collaborative allocation method provided in an embodiment of this application. Figure 3 A flowchart illustrating the algorithm of a resource collaborative allocation method provided in an embodiment of this application is shown. Figure 4 A schematic diagram of the structure of a resource collaborative allocation device provided in an embodiment of this application is shown. Detailed Implementation

[0022] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0023] The resource collaborative allocation method, apparatus, electronic device, and storage medium of this application are described below with reference to the accompanying drawings.

[0024] To address the technical problems that current solutions often focus on objective KPIs for optimization, failing to incorporate user attention and subjective experience into the model, thus hindering refined resource allocation based on perceived value and resulting in low resource utilization efficiency and poor user experience, and that existing solutions often focus on two-way relationships such as "user-edge" or "platform-edge," failing to construct a profit-sharing and constraint mechanism covering the interaction of users, platforms, and infrastructure, leading to an incomplete price signal and revenue transmission chain and difficulty in forming effective business incentives, this application provides a resource collaborative allocation method, such as... Figure 1 As shown, the method includes: Step 101: Obtain the user's visual attention data, divide the user's metaverse interaction task into multiple subtasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each subtask.

[0025] To address the profound deficiencies in the aforementioned background technology, this application provides a metaverse resource collaborative allocation method based on attention perception and hierarchical game theory. Its core objective is: First, to achieve perception-driven refined resource allocation, fundamentally changing the extensive resource allocation model oriented towards objective KPIs. This application introduces user visual attention as the core input, precisely allocating resources to the content most relevant to the user, aiming to eliminate "perceptual resource waste" and maximize system resource utilization efficiency while ensuring or even exceeding the user's subjective experience.

[0026] Secondly, this application constructs a dynamic economic incentive mechanism that benefits all parties: Addressing the conflict of interest among the user, service provider (MSP), and edge node, this application designs an innovative two-layer Stackelberg game framework (i.e., a two-layer Stackelberg game model). This framework, through market-based pricing and procurement mechanisms, deeply couples and dynamically engages the interests of the three parties, aiming to find a Stackelberg equilibrium point that simultaneously optimizes the benefits for all parties, forming a healthy and sustainable business loop.

[0027] Thirdly, the system is endowed with real-time adaptive intelligent decision-making capabilities: To cope with the high dynamism of the metaverse environment, this application employs advanced deep reinforcement learning algorithms as solvers for the game model. This enables the system to learn from continuous interaction with the environment, making near-optimal pricing, unloading, and resource procurement decisions in real time without the need for pre-setting static models, exhibiting strong robustness and scalability.

[0028] Fourthly, achieving deep collaboration among communication, computation, and rendering resources: This application incorporates the three key heterogeneous resources required for the Metaverse service into a unified quantification and optimization framework. Through a service demand awareness model and a joint allocation strategy, it achieves on-demand and proportional collaborative supply of the three types of resources, aiming to break down resource silos and achieve end-to-end global performance optimization.

[0029] In this embodiment, as Figure 2 As shown, the resource collaborative allocation system of this application consists of three layers of entities, which may include a user layer, a service provider layer, and an infrastructure layer.

[0030] The user layer may contain: Each user of the Metaverse is equipped with a head-mounted display device with eye-tracking capabilities.

[0031] Service provider layer: The core of the system, which logically includes four modules: task perception and decomposition module, first-layer game engine (MSP - user game), task offloading decision module, and second-layer game engine (MSP - edge node game). Task awareness and decomposition module: It can receive user requests, process eye-tracking data, and decompose tasks into priority sub-tasks; The first-layer game engine (MSP - User Game): can be based on the TD3 algorithm and is responsible for formulating the quality of service (QoE) and pricing strategies; Task unloading decision module: Based on the Dueling DQN algorithm, it is responsible for selecting the best edge node for execution for each subtask; The second-layer game engine (MSP - edge node game): can be based on the TD3 algorithm and is responsible for formulating resource procurement quantity and pricing strategies.

[0032] Infrastructure layer: may contain Each of the heterogeneous edge nodes has a different number of communication, computing, and rendering resources and can dynamically quote prices.

[0033] For the embodiments of this application, such as Figure 3 As shown, based on visual attention data, the user's metaverse interaction task is divided into multiple subtasks with different priorities, and a corresponding demand vector containing resource requirements is assigned to each subtask. Specifically, this may include: Based on the human visual hierarchy of visual attention data, the user's visual field is divided into multiple regions with different attention levels. These regions include at least three attention levels: fovea, parafovea, and peripheral vision. Configure a corresponding resolution level for each attention level region, and decompose the metaverse interaction task into multiple subtasks with different priorities corresponding to the resolution level, based on the regionalized task scale. Configure corresponding attention weights for each subtask. Based on the attention weights and preset service quality levels, assign a corresponding demand vector containing resource requirements to each subtask. The service quality requirements include at least the end-to-end latency limit and the resolution level.

[0034] This application addresses the pain points of "subjective attention-driven," "multi-agent collaborative incentive," and "unified metering and settlement of heterogeneous resources across edge and cloud" in multi-user metaverse scenarios, and proposes an end-to-end closed-loop method of "representation-unloading-incentive": First, use a unified representation model to achieve co-domain quantification of resources, prices, demand, and QoE; Second point: The attention-aware task offloading method prioritizes the high-concern area under latency / consistency constraints; Thirdly, a two-layer game theory model is used to achieve a closed loop of price and profit sharing among users, MSPs, and infrastructure. The three parties are interconnected, which not only ensures subjective experience but also supports an auditable incentive mechanism.

[0035] In the metaverse scenario, users wear head-mounted display devices with eye-tracking capabilities, allowing the system to acquire their visual attention data in real time. Based on the hierarchical characteristics of human vision reflected in the visual attention data, the user's field of vision (FoV) is divided into multiple different attention levels. Specifically, it is divided into three different attention levels: fovea (<5%), parafovea (<30%), and peripheral vision (<60%).

[0036] After projecting panoramic video onto a 2D plane, the instantaneous field of view for each user is... For different partitions and corresponding attention levels Furthermore, for each attention level region, a corresponding resolution level can be configured. For example, the central concave area, which receives the most attention from users, is configured with the highest resolution level to ensure that users can clearly see details; the peripheral concave area receives the second most attention and is configured with a medium resolution level; the peripheral visual area receives the least attention and is configured with a lower resolution level.

[0037] The regionalized task size is defined as follows:

[0038] in, Let be the task size after user i is regionalized at time t. Attention level The area representing the level of attention. The resolution level configured for user i's attention level. This refers to the number of bits per pixel.

[0039] Based on the regionalized task scale, the metaverse interaction task is decomposed into multiple subtasks with different priorities corresponding to the resolution level. For example, in a metaverse exhibition scene, if a user is viewing a virtual exhibit (a blue and white porcelain vase) and their gaze is focused on the texture details of the vase for 90% of the time, then the task of rendering the vase is classified as a high-priority subtask, while the task of rendering distant visitors and buildings is classified as a low-priority subtask.

[0040] The above-mentioned tiering allows service providers to allocate resources to high-attention areas based on user attention, thereby significantly reducing the overall rendering load without sacrificing the subjective experience.

[0041] Configure appropriate attention weights for each subtask based on its attention level. (For example, central concave) =0.8, concave center =0.15, peripheral view =0.05). Based on attention weights and preset service quality levels, a corresponding demand vector containing resource requirements is assigned to each subtask. The service quality requirements at least cover the end-to-end latency ceiling and resolution level.

[0042] To ensure a better user experience for subtasks of different priorities, each attention level... Set rendering latency limit At edge nodes Up, level The rendering latency and downlink transmission latency are respectively:

[0043] in, For the task on edge node j Relevant and attention level is Rendering latency, For the task on edge node j Relevant and attention level is downlink transmission latency, The number of computation cycles required per bit For nodes Rendering speed For demand-weight mapping, For downlink speed, Let be the task size after user i is regionalized at time t.

[0044] The total latency of a session can be broken down into three parts: transmission, rendering, and computation. ,in, The total latency of the session. For downlink transmission delay, To reduce rendering latency, To calculate latency and link it with a Weber–Fechner type QoE function, a logarithmic term is used to reflect the marginal gain in image quality while meeting the latency threshold:

[0045] in, To experience quality indicators, For adjustment coefficients, This is the upper limit threshold for latency. For the total delay, Attention level For user i at attention level The following image quality related indicators is the baseline image quality index, and n is the exponential parameter.

[0046] Optimization must meet the following requirements ,in, For the task on edge node j Relevant and attention level is Rendering latency, Attention level The settings include a maximum rendering latency limit and a total latency threshold.

[0047] Step 102: Quantify the subtask requirement vector, the heterogeneous resources of the edge nodes, and the user experience quality value into a unified measurement system to generate a unified requirement vector, experience quality value, and heterogeneous resources for resource allocation and settlement.

[0048] In order to address the problems of difficulty in modeling heterogeneous resources of edge nodes in the same domain, difficulty in comparable pricing, and difficulty in aligning with user subjective experience, this application proposes a unified representation model that normalizes multidimensional resources, prices, and task requirements to the same computable and settlement coordinate system, thereby realizing cross-layer mapping from "perception-resources-price-QoE".

[0049] The process of constructing a unified measurement system may specifically include: The heterogeneous resources of edge nodes are abstracted into a multi-dimensional resource capacity vector and a corresponding unit price vector. The heterogeneous resources include communication resources, general computing resources, rendering resources and storage resources. Establish a quantitative model for the subjective user experience quality, and map the end-to-end latency of each subtask to an experience quality value; By using preset normalization coefficients, all resource capacity vectors, unit price vectors, demand vectors, and experience quality values ​​are normalized to construct a unified measurement system.

[0050] In this embodiment, the process of constructing a unified measurement system is as follows: First, such as Figure 3 As shown, the heterogeneous resources of edge nodes can be abstracted and represented by multi-dimensional resource capacity vectors and corresponding unit price vectors. Heterogeneous resources may include communication resources (such as available bandwidth), general-purpose computing resources (such as available general-purpose computing power), rendering resources (such as graphics rendering capabilities), and storage resources (such as available storage). For example, the capacity of an edge node is represented by a four-dimensional capacity vector. .

[0051] in, It is a four-dimensional capacity vector. For available bandwidth, For the availability of general-purpose computing power, For rendering graphics capabilities, Available storage.

[0052] The corresponding unit price vector is represented as

[0053] in, A unit price vector, Price per unit bandwidth Price per unit of general computing power Price per unit of graphics rendering capability Price per unit of storage.

[0054] Secondly, users can Session task Based on the field of view (FoV) attention mechanism, it is divided into a set of regions. (High / Medium / Low Attention), each area At any moment The attention weights are:

[0055] in, Let z be the attention weight of the user's session task k in time t. The smoothed attention heatmap Let z be the set of pixels in region z. Map the user's quality-latency requirements to a regionalized demand vector:

[0056] in, Let z be the regionalized demand vector of region z in user session task k at time t. Let z be the attention weight of region z in user session task k at time t. The base bandwidth required for region z to achieve the target quality. The computing power required for region z to achieve the target quality. The rendering resources required for region z to achieve the target quality. Storage resources required for region z to achieve the target quality. The weights are adaptive to the desired quality level (e.g., resolution / frame rate). The computing power weights are adaptive to the desired quality level. For rendering resource weights that adapt to the desired quality level, Storage resource weights that adapt to the desired quality level.

[0057] A quantitative model of user subjective experience quality can be established. The end-to-end latency of each subtask is mapped to an experience quality value. The regional QoE is defined as:

[0058] in, This represents the regional-level experience quality score. As a proxy indicator for image quality / smoothness, As weight, For end-to-end delay, This represents the upper limit of regional latency.

[0059] Session QoE is the sum of the QoE of each zone. This approach achieves a unified quantification of attention, demand, and QoE.

[0060] Finally, to facilitate comparisons across nodes and time periods, preset normalization coefficients can be used. This involves normalizing all resource capacity vectors, unit price vectors, demand vectors, and experience quality values. For example, a normalization coefficient is introduced to normalize the resource capacity vector (i.e., the four-dimensional capacity vector). ) and unit price vector ( Perform element-wise division and element-wise multiplication operations to obtain the normalized capacity. ) and standardized price ( ),in, To divide element by element, This involves element-wise multiplication. Through this series of processes, a unified measurement system is built, enabling the quantification and comparison of resource, price, demand, and experience quality values ​​across different dimensions within the same coordinate system. This provides a computable interface for subsequent resource allocation and settlement, greatly simplifying the complexity of managing multidimensional resources.

[0061] Step 103: Based on a unified demand vector, experience quality value, and heterogeneous resources, a deep reinforcement learning algorithm is used to assign each subtask to the optimal edge node under the condition of meeting the latency constraint, thereby obtaining the task offloading decision scheme for each subtask.

[0062] For the embodiments of this application, following steps 101 and 102 of the embodiments, this application constructs an attention-aware QoE model based on visual saliency and combines dual connectivity (end-to-end decomposition of communication latency and rendering / computation latency), and uses DuelingDQN to obtain the optimal discrete assignment in a multi-task-multi-edge node environment.

[0063] Specifically, the latency constraints may include: node capacity constraints, regional consistency constraints, and single-homing constraints. Node capacity constraints can be used to characterize that the total resource requirements of all subtasks assigned to the same edge node do not exceed its resource capacity vector. The node capacity constraint can be:

[0064] in, To perform a summation operation on all user session tasks k and regions z, Let z be the regionalized demand vector of region z in user session task k at time t. Let n be the resource capacity vector of edge node n.

[0065] Regional consistency constraints can be used to characterize that the end-to-end latency of subtasks assigned to edge nodes does not exceed a preset latency limit; The regional consistency constraint can be:

[0066] in, For adjustment coefficients, For the end-to-end latency of the high-interest region H in user session task k. Let z be the end-to-end latency of region z in user session task k. This is the preset latency upper limit threshold.

[0067] The single-homing constraint is used to characterize that each subtask is assigned to only one edge node; The single-ownership constraint can be:

[0068] in, Summation is performed on all edge nodes n.

[0069] like Figure 3 As shown, through the Unified Representation Model (URM), all entities are allocated and settled under a unified measurement system, providing a computable interface for subsequent unloading optimization and pricing game theory, greatly simplifying the management complexity of multi-dimensional resources. This model is a necessary prerequisite and key foundation for realizing subsequent integrated collaborative game theory and scheduling of communication, computing, and rendering, solving the fundamental problem in existing technologies where effective joint optimization is impossible due to inconsistent resource standards.

[0070] In multitasking With multiple edge nodes In scenarios where time delay constraints are met, deep reinforcement learning algorithms (such as Dueling DQN) can be used to solve the discrete assignment of tasks to nodes, which has better stability than direct Q estimation.

[0071] state space Data size that can be included for each task Service quality requirements and resource preference coefficient ; Right now

[0072] Action space For each task Select an edge node via binary assignment vector Indicates task attribution, where, Indicates allocation to In terms of value decomposition, the Q function is decomposed into state values. With advantage function This avoids conflating the overall value of a state with the relative advantage of actions, improving convergence and accuracy in unstable scenarios. During the inference phase, the Q-values ​​of each possible action for the current state are output, and the largest one is selected.

[0073] The objective is to minimize overall execution latency and indirectly maximize QoE while satisfying QoE / latency constraints. These constraints include: regional latency caps, single-task single-homing, node capacity, and link rate limits. Training employs empirical replay and the target network, updating according to a Dueling structure until a stable optimal assignment strategy is obtained in a multi-task, multi-node scenario, thus yielding the task offloading decision scheme for each subtask.

[0074] Step 104: Use deep reinforcement learning algorithms to jointly solve the constructed two-layer Stackelberg game model and task offloading decision scheme to achieve dynamic resource collaborative allocation.

[0075] In the embodiments of this application, a deep reinforcement learning algorithm is used to jointly solve the constructed two-layer Stackelberg game model and task offloading decision scheme to achieve dynamic resource collaborative allocation, which may specifically include: A two-layer Stackelberg game model with service providers as leaders can be constructed. The two-layer Stackelberg game model includes an upper-layer game and a lower-layer game. The upper-layer game is used to represent the game between the MSP and users regarding service quality and pricing, while the lower-layer game is used to represent the game between the MSP and edge nodes regarding resource procurement quantity and bidding. The two layers are coupled in a closed loop through a unified representation model of the demand vector and the assignment result of attention offloading. The whole system uses Twin Delayed DDPG (TD3) to solve the continuous decision variables (price / procurement quantity) and Dueling DQN to solve the discrete offloading. Finally, Stackelberg equilibrium is reached in both layers, and auditable profit sharing is achieved. A deep reinforcement learning algorithm is used to jointly solve a two-layer Stackelberg game model. When the game result causes the user's experience quality value to deviate or the service provider's cost to exceed a preset cost threshold, the task offloading decision scheme is fine-tuned in reverse, forming a closed-loop coupling between the game and offloading, so as to determine the optimal service quality and pricing strategy as well as the optimal resource procurement quantity and bidding strategy. Based on the optimal service quality and pricing strategy, as well as the optimal resource procurement and bidding strategy, subtasks are offloaded to the corresponding edge nodes for execution, thereby achieving the coordinated allocation of communication, computing, and rendering resources.

[0076] In this embodiment, in the upper-level game, the MSP, as the leader, can first formulate a pricing vector strategy. Users, as followers, can choose service quality strategies based on price and their own needs. The game objective can be to maximize platform profits and achieve the target user utility. This stage is modeled as a Markov Decision Process (MDP) and the TD3 algorithm is used to learn the optimal strategy in a continuous price space.

[0077] The state space consists of the user quality selection at time t, forming a state vector. ; Action space The MSP's actions involve feasible continuous pricing adjustments; rewards are based on platform system utility to guide the maximization of long-term revenue and user retention. This setup ensures dynamic response and optimal balance in the "price-demand-quality" dynamic.

[0078] Since price / quality are continuous variables, TD3's dual-commentator + delayed update strategy and soft update mechanism are used to suppress Q-value overestimation and training oscillations; MSP (Actor) observes Then directly output the action. The optimization of price and quality has been completed, resulting in a more robust convergence.

[0079] In the lower-level game, within a multi-task, multi-node environment, the MSP, as the leader, first determines the purchase quantity for each edge node; each edge node, as a follower, responds with pricing. The two sides engage in a dynamic game revolving around "minimizing MSP costs and maximizing node profits," seeking a Stackelberg equilibrium for resource quantity and price. The state for each node... definition ,in, Indicates the available resources of a node. Indicates remaining quantity / load. The current unit price and overall status are as follows. Actions of the MSP For the node The quantity of resources procured constitutes an action set. This action is directly linked to the quality and requirements given by the upper layer, and is used to meet the actual needs after the task is unloaded; in terms of the revenue function, the node profit is used to uniformly represent the output demand. Node pricing express:

[0080] in, To determine the output demand in a unified manner for node profits, Quote for the node, This is the penalty coefficient.

[0081] Reflecting marginal revenue and price smoothing penalties, nodes continuously adjust prices to maximize their profits.

[0082] MSP Costs / Benefits:

[0083] in, Costs / benefits of MSP For the platform's baseline revenue, This is a cost item calculated as purchase quantity × unit price. To punish excessive procurement, For items that deviate from QoE, it prompts a "cost-QoE" trade-off and alignment with upper-level user utility. , These are the penalty coefficients.

[0084] This layer also uses TD3 to learn in the continuous action space (MSP: purchase quantity; edge: bid). By taking the minimum value of two Critics, the Q value overestimation is reduced. Experience replay, delayed policy updates and soft updates ensure stable convergence in non-stationary game environment. After multiple iterations, the purchase quantity and node price converge around the Stackelberg equilibrium, forming a win-win situation.

[0085] From Representation Model to Two-Level Game: After obtaining the initial marginal bid, the MSP generates a demand vector using the URM. Price vector It serves as a unified measurement and settlement basis for lower-level games.

[0086] From task unloading to a two-layer game: prioritizing high-concern areas based on saliency, after QoE-aware unloading is completed, the lower layer uses the updated... By conducting resource procurement and bidding negotiations, a closed loop of consistency is achieved from task assignment to actual needs, and then to procurement and pricing.

[0087] Two-layer game to representation model / task offloading: After the price / supply boundary is formed at the lower layer, if QoE deviates or costs are too high (see...) and The penalty item will drive fine-tuning of upper-level price / quality and task offloading assignments until both QoE and cost constraints are met.

[0088] Initialization: The MSP and edge nodes each configure their Actor-Critic structures, initialize experience playback and the target network. The user submits a heterogeneous request containing minimum QoE / latency and service requirements, and the MSP generates an initial procurement plan.

[0089] Iterative game: In each round, the MSP gives an action based on the current state using Actors (price / quality at the upper level, purchase quantity at the lower level); followers (users / edges) respond (quality selection / bid), the environment generates an immediate reward and writes it into the replay pool; two Critics evaluate independently, and the smaller Q value is used to update the Actor; delayed update strategy and soft target update are adopted to improve stability.

[0090] Convergence and Equilibrium: After multiple rounds, the upper-level "price-quality" and the lower-level "purchase-bid" reach dynamic equilibrium with the support of TD3, corresponding to the Stackelberg solution in economic terms, which reduces the MSP's procurement cost and increases marginal profit, thereby improving resource allocation efficiency.

[0091] Two-layer game theory only handles continuous economic decisions (price / purchase), while discrete task-to-node assignment is accomplished by Dueling DQN: the Q-value is decomposed into state value V and advantage A, avoiding confusion between the overall value of the state and the relative advantage of actions, thus improving stability and accuracy in multi-task assignment; the state contains task data. Service quality requirements With resource preferences .

[0092] In a specific application scenario, in a high-fidelity metaverse exhibition, user A is carefully viewing a virtual exhibit (an intricate blue and white porcelain vase), while other tourists and background buildings are in the distance.

[0093] Eye-tracking showed that user A's gaze lingered on the texture and details of the blue and white porcelain vase for 90% of the time.

[0094] High-priority subtasks ( ): Rendering a porcelain vase, QoE requirements: 8K resolution, maximum latency. The calculated data size .

[0095] low priority subtasks ( ): Render distant tourists and buildings. QoE requirement: 2K resolution. Data size .

[0096] First layer of game theory: User A has a high willingness to pay for an "expert-level" viewing experience. MSP's TD3 agent tentatively proposes a high pricing strategy, and User A's client accepts the price and selects the highest QoE level. Both parties reach an agreement. =0.5 yuan / minute The equilibrium is 0.95.

[0097] The system has edge nodes (Powerful GPU rendering capabilities, 5ms network access latency) and (Computational power is balanced, and network access latency is 10ms).

[0098] After Dueling DQN agent evaluation Because it predicted Can better meet The stringent rendering and latency requirements. Therefore, .

[0099] for , The resources and latency are sufficient to meet its lenient requirements; to balance the load, the agent chooses to... Uninstall to , .

[0100] Second-level game: The MSP's TD3 agent sends a request to purchase a large amount of GPU resources (e.g., 10 TFLOPS) and high bandwidth (50 Mbps); to A request was issued to purchase a small number of GPUs (2 TFLOPS) and standard bandwidth (10 Mbps).

[0101] Detecting a surge in demand for its own GPU resources, the company raised its price from 0.01 yuan / TFLOPs to 0.012 yuan / TFLOPs. If resources are available, maintain the original price.

[0102] The MSP agent readjusts its purchase volume based on the new price, ultimately achieving a cost-optimal procurement equilibrium.

[0103] User A felt the details of the blue-and-white porcelain vase in front of them were incredibly clear, and there was no lag when rotating the view, resulting in an excellent immersive experience. Meanwhile, the system saved approximately 60% of rendering resources and 40% of bandwidth resources by degrading the background. MSP, while providing top-tier QoE, achieved high profits through meticulous cost control. and The effective use of resources has also increased revenue, resulting in a win-win situation for all parties.

[0104] In summary, based on the resource collaborative allocation method provided in this application, compared with existing technologies, this application can obtain user visual attention data, divide the user's metaverse interaction task into multiple sub-tasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each sub-task; quantify the sub-task demand vector, heterogeneous resources of edge nodes, and user experience quality value into a unified measurement system to generate a unified demand vector, experience quality value, and heterogeneous resources for resource allocation and settlement; based on the unified demand vector, experience quality value, and heterogeneous resources, a deep reinforcement learning algorithm is used to assign each sub-task to the optimal edge node under the condition of satisfying latency constraints, thereby obtaining a task offloading decision scheme for each sub-task; and a deep reinforcement learning algorithm is used to jointly solve the constructed two-layer Stackelberg game model and task offloading decision scheme to achieve dynamic resource collaborative allocation.

[0105] The solution described in this application acquires user visual attention data and then meticulously divides the user's metaverse interaction tasks into multiple sub-tasks with different priorities based on this data. This approach introduces user attention factors into the task processing flow, matching the importance of different sub-tasks with the user's level of attention, thus laying the foundation for subsequent resource allocation based on user perceived value.

[0106] By assigning a demand vector containing resource requirements to each subtask, and quantifying the subtask demand vectors, heterogeneous resources of edge nodes, and user experience quality values ​​into a unified measurement system, a unified demand vector, experience quality value, and heterogeneous resources are generated for resource allocation and settlement. This quantified and unified approach incorporates the user's subjective experience quality value into the resource allocation considerations, enabling resource allocation to fully take into account the user's subjective feelings. This achieves refined resource allocation based on perceived value, thereby improving resource utilization efficiency and enhancing user experience.

[0107] Based on a unified demand vector, experience quality score, and heterogeneous resources, a deep reinforcement learning algorithm is employed to assign each subtask to the optimal edge node while satisfying latency constraints, thus obtaining a task offloading decision scheme for each subtask. The deep reinforcement learning algorithm can dynamically adjust the task offloading strategy according to the constantly changing environment and user needs, further optimizing resource allocation and ensuring that resource utilization efficiency is improved while meeting user needs, thereby enhancing the user experience.

[0108] This application employs a deep reinforcement learning algorithm to jointly solve a constructed two-layer Stackelberg game model and a task offloading decision scheme. This joint solution approach ensures that the resource allocation process not only considers the optimization of task offloading but also takes into account the game and balance of interests among users, the platform, and infrastructure. This method enables the formation of a complete price signal and benefit transmission chain, allowing all parties to obtain reasonable benefits during resource allocation, thereby creating effective business incentives and promoting the healthy development of the entire system.

[0109] Based on the above Figure 1 To illustrate the specific implementation of the method shown, this embodiment provides a resource collaborative allocation device, such as... Figure 4 As shown, the device includes: an acquisition module 31, a generation module 32, an assignment module 33, and a solution module 34; The acquisition module 31 is used to acquire the user's visual attention data, divide the user's metaverse interaction task into multiple sub-tasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each sub-task. The generation module 32 is used to quantify the demand vector of the subtask, the heterogeneous resources of the edge nodes, and the user's experience quality value into a unified measurement system, and generate a unified demand vector, experience quality value, and heterogeneous resources for resource allocation and settlement. The assignment module 33 is used to assign each subtask to the optimal edge node based on the unified demand vector, the experience quality value and the heterogeneous resources, using a deep reinforcement learning algorithm, under the condition of meeting the latency constraint, so as to obtain the task unloading decision scheme for each subtask. The solution module 34 is used to jointly solve the constructed two-layer Stackelberg game model and the task unloading decision scheme using the deep reinforcement learning algorithm, so as to realize dynamic resource collaborative allocation.

[0110] In specific application scenarios, the acquisition module 31 can be used to divide the user's visual field into multiple regions with different attention levels based on the human eye visual hierarchy of the visual attention data; the multiple regions with different attention levels include: fovea, parafovea and peripheral visual field regions with at least three attention levels; Configure a corresponding resolution level for each attention level region, and decompose the metaverse interaction task into multiple sub-tasks with different priorities corresponding to the resolution level, based on the regionalized task scale. Configure a corresponding attention weight for each subtask, and assign a corresponding demand vector containing resource requirements to each subtask based on the attention weight and a preset service quality level, wherein the service quality requirements include at least an end-to-end latency limit and a resolution level.

[0111] In specific application scenarios, the generation module 32 can be used to abstract the heterogeneous resources of the edge nodes into a multi-dimensional resource capacity vector and a corresponding unit price vector, wherein the heterogeneous resources include communication resources, general computing resources, rendering resources and storage resources; Establish a quantitative model for the user's subjective experience quality, and map the end-to-end latency of each subtask to an experience quality value; By using a preset normalization coefficient, all the resource capacity vectors, unit price vectors, demand vectors, and experience quality values ​​are normalized to construct the unified measurement system.

[0112] In specific application scenarios, the assignment module 33 can be used for the aforementioned latency constraints, including: node capacity constraints, regional consistency constraints, and single-homing constraints. The node capacity constraint is used to characterize that the total resource requirement of all subtasks allocated to the same edge node does not exceed its resource capacity vector. The regional consistency constraint is used to characterize that the end-to-end latency of the subtasks assigned to the edge nodes does not exceed a preset latency limit. The single-homing constraint is used to characterize that each subtask is assigned to only one edge node.

[0113] In specific application scenarios, the solution module 34 can be used to construct a two-layer Stackelberg game model with the service provider as the leader. The two-layer Stackelberg game model includes an upper-layer game and a lower-layer game. The upper-layer game is used to represent the game between the service provider and the user regarding service quality and pricing, and the lower-layer game is used to represent the game between the service provider and the edge node regarding resource procurement quantity and bidding. A deep reinforcement learning algorithm is used to jointly solve the two-layer Stackelberg game model. When the game result causes the user's experience quality value to deviate or the service provider's cost to exceed a preset cost threshold, the task offloading decision scheme is fine-tuned in reverse, forming a closed-loop coupling between the game and offloading, so as to determine the optimal service quality and pricing strategy as well as the optimal resource procurement quantity and bidding strategy. Based on the optimal service quality and pricing strategy, as well as the optimal resource procurement quantity and bidding strategy, the subtasks are offloaded to the corresponding edge nodes for execution, thereby achieving the coordinated allocation of communication, computing, and rendering resources.

[0114] It should be noted that other corresponding descriptions of the functional units involved in the resource collaborative allocation device provided in this embodiment can be found in [reference]. Figure 1 The corresponding descriptions in [the document] will not be repeated here.

[0115] Based on the above, Figure 1 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 The method shown.

[0116] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0117] Based on the above, Figure 1 The method shown, and Figure 4 To achieve the above objectives, the present application also provides an electronic device, comprising a storage medium and a processor; the storage medium for storing a computer program; and the processor for executing the computer program to implement the above-described virtual device embodiments. Figure 1 The method shown.

[0118] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0119] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0120] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. Compared with the prior art, the technical solution of this application can obtain the user's visual attention data, divide the user's metaverse interaction task into multiple sub-tasks with different priorities according to the visual attention data, and assign a corresponding demand vector containing resource requirements to each sub-task; quantify the demand vector of the sub-task, the heterogeneous resources of the edge nodes, and the user's experience quality value into a unified measurement system to generate a unified demand vector, experience quality value, and heterogeneous resources for resource allocation and settlement; based on the unified demand vector, experience quality value, and heterogeneous resources, a deep reinforcement learning algorithm is used to assign each sub-task to the optimal edge node under the condition of satisfying the latency constraint, so as to obtain the task offloading decision scheme for each sub-task; the deep reinforcement learning algorithm is used to jointly solve the constructed two-layer Stackelberg game model and the task offloading decision scheme to realize dynamic resource collaborative allocation.

[0122] The solution described in this application acquires user visual attention data and then meticulously divides the user's metaverse interaction tasks into multiple sub-tasks with different priorities based on this data. This approach introduces user attention factors into the task processing flow, matching the importance of different sub-tasks with the user's level of attention, thus laying the foundation for subsequent resource allocation based on user perceived value.

[0123] By assigning a demand vector containing resource requirements to each subtask, and quantifying the subtask demand vectors, heterogeneous resources of edge nodes, and user experience quality values ​​into a unified measurement system, a unified demand vector, experience quality value, and heterogeneous resources are generated for resource allocation and settlement. This quantified and unified approach incorporates the user's subjective experience quality value into the resource allocation considerations, enabling resource allocation to fully take into account the user's subjective feelings. This achieves refined resource allocation based on perceived value, thereby improving resource utilization efficiency and enhancing user experience.

[0124] Based on a unified demand vector, experience quality score, and heterogeneous resources, a deep reinforcement learning algorithm is employed to assign each subtask to the optimal edge node while satisfying latency constraints, thus obtaining a task offloading decision scheme for each subtask. The deep reinforcement learning algorithm can dynamically adjust the task offloading strategy according to the constantly changing environment and user needs, further optimizing resource allocation and ensuring that resource utilization efficiency is improved while meeting user needs, thereby enhancing the user experience.

[0125] This application employs a deep reinforcement learning algorithm to jointly solve a constructed two-layer Stackelberg game model and a task offloading decision scheme. This joint solution approach ensures that the resource allocation process not only considers the optimization of task offloading but also takes into account the game and balance of interests among users, the platform, and infrastructure. This method enables the formation of a complete price signal and benefit transmission chain, allowing all parties to obtain reasonable benefits during resource allocation, thereby creating effective business incentives and promoting the healthy development of the entire system.

[0126] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0127] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A resource collaborative allocation method, characterized in that, The method includes: Acquire the user's visual attention data, divide the user's metaverse interaction task into multiple subtasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each subtask. The requirement vector of the subtask, the heterogeneous resources of the edge nodes, and the user's experience quality value are quantified into a unified measurement system to generate a unified requirement vector, experience quality value, and heterogeneous resources for resource allocation and settlement. Based on the unified demand vector, experience quality value, and heterogeneous resources, a deep reinforcement learning algorithm is used to assign each subtask to the optimal edge node under the condition of meeting the latency constraint, thereby obtaining the task offloading decision scheme for each subtask. The constructed two-layer Stackelberg game model and the task offloading decision scheme are jointly solved using the deep reinforcement learning algorithm to achieve dynamic resource collaborative allocation.

2. The resource collaborative allocation method according to claim 1, characterized in that, The step of dividing the user's metaverse interaction task into multiple sub-tasks with different priorities based on the visual attention data, and assigning a corresponding demand vector containing resource requirements to each sub-task, includes: Based on the human visual hierarchy of the aforementioned visual attention data, the user's visual field is divided into multiple regions with different attention levels; Configure a corresponding resolution level for each attention level region, and decompose the metaverse interaction task into multiple sub-tasks with different priorities corresponding to the resolution level, based on the regionalized task scale. Configure a corresponding attention weight for each subtask, and assign a corresponding demand vector containing resource requirements to each subtask based on the attention weight and a preset service quality level, wherein the service quality requirements include at least an end-to-end latency limit and a resolution level.

3. The resource collaborative allocation method according to claim 2, characterized in that, The multiple regions with different attention levels include: regions with at least three attention levels: fovea, parafovea, and peripheral vision.

4. The resource collaborative allocation method according to claim 2, characterized in that, The process of constructing the unified measurement system includes: The heterogeneous resources of the edge nodes are abstracted into a multi-dimensional resource capacity vector and a corresponding unit price vector, wherein the heterogeneous resources include communication resources, general computing resources, rendering resources and storage resources; Establish a quantitative model for the user's subjective experience quality, and map the end-to-end latency of each subtask to an experience quality value; By using a preset normalization coefficient, all the resource capacity vectors, unit price vectors, demand vectors, and experience quality values ​​are normalized to construct the unified measurement system.

5. The resource collaborative allocation method according to claim 1, characterized in that, The latency constraints include: node capacity constraints, regional consistency constraints, and single-homing constraints; The node capacity constraint is used to characterize that the total resource requirement of all subtasks allocated to the same edge node does not exceed its resource capacity vector. The regional consistency constraint is used to characterize that the end-to-end latency of the subtasks assigned to the edge nodes does not exceed a preset latency limit. The single-homing constraint is used to characterize that each subtask is assigned to only one edge node.

6. The resource collaborative allocation method according to claim 1, characterized in that, The method employs the deep reinforcement learning algorithm to jointly solve the constructed two-layer Stackelberg game model and the task offloading decision scheme, achieving dynamic resource collaborative allocation, including: Construct a two-level Stackelberg game model with the service provider as the leader; A deep reinforcement learning algorithm is used to jointly solve the two-layer Stackelberg game model. When the game result causes the user's experience quality value to deviate or the service provider's cost to exceed a preset cost threshold, the task offloading decision scheme is fine-tuned in reverse, forming a closed-loop coupling between the game and offloading, so as to determine the optimal service quality and pricing strategy as well as the optimal resource procurement quantity and bidding strategy. Based on the optimal service quality and pricing strategy, as well as the optimal resource procurement quantity and bidding strategy, the subtasks are offloaded to the corresponding edge nodes for execution, thereby achieving the coordinated allocation of communication, computing, and rendering resources.

7. The resource collaborative allocation method according to claim 1, characterized in that, The two-layer Stackelberg game model includes an upper-layer game and a lower-layer game. The upper-layer game is used to represent the game between the service provider and the user regarding service quality and pricing, while the lower-layer game is used to represent the game between the service provider and the edge node regarding resource procurement quantity and bidding.

8. A resource collaborative allocation device, characterized in that, include: The acquisition module is used to acquire the user's visual attention data, divide the user's metaverse interaction task into multiple sub-tasks with different priorities based on the visual attention data, and assign a corresponding demand vector containing resource requirements to each sub-task. The generation module is used to quantify the demand vector of the subtask, the heterogeneous resources of the edge nodes, and the user's experience quality value into a unified measurement system, and generate a unified demand vector, experience quality value, and heterogeneous resources for resource allocation and settlement. The assignment module is used to assign each subtask to the optimal edge node based on the unified demand vector, the experience quality value, and the heterogeneous resources, using a deep reinforcement learning algorithm, under the condition of meeting the latency constraint, so as to obtain the task unloading decision scheme for each subtask. The solution module is used to jointly solve the constructed two-layer Stackelberg game model and the task unloading decision scheme using the deep reinforcement learning algorithm, so as to realize dynamic resource collaborative allocation.

9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the resource collaborative allocation method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the resource collaborative allocation method according to any one of claims 1 to 7.