Energy-saving wireless computing power network resource management method and device

By building a task model and a deep reinforcement learning model for wireless computing power network, the resource waste problem caused by single goal optimization in traditional wireless computing power network resource management is solved, efficient resource utilization and energy consumption are achieved, task execution efficiency is improved and delay is reduced.

CN120343631APending Publication Date: 2025-07-18GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510568474.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The traditional wireless computing power network energy saving optimization method mainly focuses on local optimization of a single goal, and it is difficult to balance the effective utilization of computing resources and energy consumption.

Method used

Build a task model, computing model, communication model and queue model based on wireless computing power network, with the goal of minimizing the total energy consumption of the network, and convert it into a Markov decision-making process and build a deep reinforcement learning model. It can maximize iterative training to output task offload decisions.

Benefits of technology

It realizes the flexibility of using resources, reducing energy consumption, balancing task offloading and resource scheduling decisions, improving task execution efficiency and reducing latency while ensuring network stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343631A_ABST
    Figure CN120343631A_ABST
Patent Text Reader

Abstract

The invention discloses an energy-saving wireless computing power network resource management method and device, and relates to the technical field of communication, and the method comprises the steps: taking mutual transmission among computing nodes of a wireless computing power network as a basis, and based on a task model, a computing model, a communication model and a queue model of the wireless computing power network; constructing a task resource joint optimization problem by taking the minimization of the total energy consumption of the network as a target; converting the task resource joint optimization problem into a Markov decision process, correspondingly constructing a to-be-trained deep reinforcement learning model, performing iterative training on the to-be-trained deep reinforcement learning model by taking reward value maximization as a target, and determining a trained deep reinforcement learning model; and when the wireless computing power network receives a to-be-processed computing task, outputting a target task unloading decision of the to-be-processed computing task through the trained deep reinforcement learning model. Based on the above scheme, it can be ensured that the wireless computing power network can find a balance between stability and energy efficiency during task unloading and resource scheduling decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to an energy-saving wireless computing power network resource management method and apparatus. Background Art

[0002] With the rapid development of technologies such as mobile Internet, big data, Internet of Things (IoT), artificial intelligence (AI), virtual reality (VR), and augmented reality (AR), the user demand for computing power has increased sharply. Traditional computing platforms such as cloud computing and edge computing are difficult to meet the needs of a large number of users and various application scenarios in terms of real-time performance, flexibility, energy efficiency, and cost. As an emerging network architecture, the Wireless Computing Power Network (WPCN) connects computing resources (such as cloud computing, edge computing, and terminal devices) distributed at different locations through wireless technologies. Computing tasks can be dynamically offloaded from edge devices to remote computing nodes, and the computing power of cloud or edge computing platforms is utilized to achieve on-demand scheduling and allocation of computing resources to meet real-time computing requirements. It has advantages such as flexibility, intelligence, and low latency, and can reduce deployment costs.

[0003] However, the energy consumption problem in the wireless computing power network has become increasingly severe with the increase in computing resources, the increase in the number of users, and the change in communication load. Factors such as the offloading of computing tasks, the dynamic scheduling of resources, and the communication delay will all affect the energy efficiency of the network. Traditional energy-saving optimization methods for wireless computing power networks mainly focus on reducing the power consumption of computing nodes or optimizing the signal transmission power. Although they can reduce energy consumption to a certain extent, this kind of local optimization focusing on a single target (such as energy efficiency or computing resource scheduling) often ignores the impacts of resource scheduling, network load balancing, and task offloading, etc., and is prone to resource waste or performance degradation in some aspects of the wireless computing power network, and it is difficult to balance the effective utilization of computing resources and energy consumption. Summary of the Invention

[0004] The present invention provides an energy-saving wireless computing power network resource management method and apparatus, which solves the problem that traditional energy-saving optimization methods for wireless computing power networks mainly focus on local optimization of a single target and are difficult to balance the effective utilization of computing resources and energy consumption.

[0005] An energy-saving wireless computing power network resource management method provided by the first aspect of the present invention includes:

[0006] Based on the mutual transmission between computing nodes of the wireless computing power network, a joint task-resource optimization problem is constructed with the goal of minimizing the total network energy consumption based on the task model, computing model, communication model, and queue model of the wireless computing power network;

[0007] Convert the task resource joint optimization problem into a Markov decision process and correspondingly construct a deep reinforcement learning model to be trained. Iteratively train the deep reinforcement learning model to be trained with the goal of maximizing the reward value to determine the trained deep reinforcement learning model;

[0008] When the wireless computing power network receives a computing task to be processed, output the target task offloading decision of the computing task to be processed through the trained deep reinforcement learning model.

[0009] Further, the task model includes:

[0010] , ;

[0011] When the computing task will be processed at the local computing node When the computing task will be transmitted to other computing nodes for processing;

[0012] In the formula, is the time slot , is the computing node of the wireless computing power network , is the computing node of the wireless computing power network , is the task generation rate, is the Poisson distribution, is the computing task of the computing node , is the task computing requirement of the computing node , is the task transmission requirement of the computing node , is the task push requirement of the computing node , is the time slot is the transmission decision variable;

[0013] The computing model includes a local computing sub-model and other node computing sub-models;

[0014] The local computing sub-model includes:

[0015] ;

[0016] ;

[0017] In the formula, is the task processing time of the computing node in the time slot , To calculate the task computing requirements of compute nodes in a time slot let be the computing power of compute nodes let be the set of compute nodes in the wireless computing power network let be the set of time slots let be the task processing energy consumption of compute nodes in a time slot let be the energy coefficient of compute nodes

[0018] relative to the chip architecture

[0019] ;

[0020] ;

[0021] where ;

[0022] In the formula let be the task processing time of compute nodes in a time slot let be the computing resource allocation coefficient of compute nodes in a time slot let be the computing power of compute nodes let be the task processing energy consumption of compute nodes in a time slot let be the energy consumption of compute nodes representing the execution of CPU / GPU cycles let

[0023] The communication model includes

[0024] ;

[0025] ;

[0026] ;

[0027] ​ ;

[0028] ;

[0029] ;

[0030] ;

[0031] ;

[0032] ;

[0033] wherein, is the bandwidth allocated from computing node to computing node in time slot , is the allocation coefficient of bandwidth resources, is the total bandwidth resources of computing node , is the orthogonal sub-channel , is the set of available sub-channels, is the signal-to-noise ratio when computing node transmits tasks to computing node , is the transmission power of computing node , is the channel gain between computing node and computing node in time slot , is the inter-cell interference when computing node allocates the orthogonal sub-channel to computing node to transmit tasks in time slot , is the noise spectral density, is the rate at which computing node transmits tasks to computing node in time slot , is the delay when computing node transmits tasks to computing node in time slot , is the transmission energy consumption when computing node transmits tasks to computing node in time slot , is the computing node transmits to the computing node The signal-to-noise ratio when pushing tasks is the transmission power of the computing node ; is the rate at which the computing node pushes tasks to the computing node during the time slot ; is the task reception and processing delay when the computing node pushes tasks to the computing node during the time slot ; is the task push requirement of the computing node during the time slot ; is the push energy consumption when the computing node pushes tasks to the computing node during the time slot ;

[0034] The queue model includes a local queue model and other node queue models:

[0035] The local queue model includes:

[0036] ;

[0037] The other node queue models include:

[0038] ;

[0039] In the formula, is the length of the task queue to be calculated on the computing node at time slot ; is the length of the task queue to be calculated on the computing node at time slot ;

[0040] Furthermore, the task resource joint optimization problem includes:

[0041] ;

[0042] Among them, , , , , , ;

[0043] In the formula, is the time slot , is the computing node of the wireless computing power network , The computing node of the wireless computing power network , is the transmission decision variable of the time slot , is the set of computing nodes of the wireless computing power network is the set of time slots is the transmission decision is the computing resource is the computing resource allocation coefficient of the computing node in the time slot , is the bandwidth resource is the allocation coefficient of the bandwidth resource is the time length of the time slot is the total network energy consumption in the time slot , is the energy consumption of the computing node to complete the task in the time slot , is the total task delay in the time slot , is the task processing time of the computing node in the time slot , is the task processing time of the computing node in the time slot , is the delay of the task transmission from the computing node to the computing node in the time slot , is the task reception and processing delay when the computing node pushes the task to the computing node in the time slot , is the task processing energy consumption of the computing node in the time slot , is the transmission energy consumption of the task transmission from the computing node to the computing node in the time slot , is the push energy consumption when the computing node pushes the task to the computing node in the time slot , is the maximum tolerable delay is the task queue length in the time slot , is the expectation

[0044] Furthermore, the Markov decision process includes:

[0045] Define the computing power network state space with the task generation rate, task queue length, channel gain, inter-cell interference, and the amount of computing resources required for the computing nodes to complete tasks;

[0046] Define the computing power network action space with transmission decision variables, computing resource allocation coefficients of computing nodes, and transmission powers of computing nodes;

[0047] Define the reward function by integrating the total network energy consumption, total task delay, and task queue length;

[0048] Quantify the reward value of performing the computing power network actions in the computing power network action space under the computing power network state in the computing power network state space based on the reward function.

[0049] Furthermore, the reward function includes:

[0050] ;

[0051] In the formula, is the reward value, is the total network energy consumption in time slot , is the total task delay in time slot , is the task queue length in time slot , is the energy consumption weight factor, is the delay weight factor, is the discount factor.

[0052] Furthermore, when the wireless computing power network receives a computing task to be processed, the target task offloading decision of the computing task to be processed output by the trained deep reinforcement learning model includes:

[0053] When the wireless computing power network receives a computing task to be processed, determine the current computing power network state of the wireless computing power network;

[0054] Input the current computing power network state into the trained deep reinforcement learning model, and output the target task offloading decision of the computing task to be processed.

[0055] An energy-saving wireless computing power network resource management device provided in the second aspect of the present invention includes:

[0056] An optimization problem determination module, which is used to construct a joint task resource optimization problem with the goal of minimizing the total network energy consumption based on the mutual transmission between computing nodes of the wireless computing power network and based on the task model, computing model, communication model, and queue model of the wireless computing power network;

[0057] A reinforcement learning module, configured to transform the task resource joint optimization problem into a Markov decision process and correspondingly construct a deep reinforcement learning model to be trained, and iteratively train the deep reinforcement learning model to be trained with the maximization of the reward value as the goal, so as to determine the trained deep reinforcement learning model;

[0058] A task resource decision module, configured to, when the wireless computing power network receives a computing task to be processed, output a target task offloading decision of the computing task to be processed through the trained deep reinforcement learning model.

[0059] A computer device provided in the third aspect of the present invention includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the energy-saving wireless computing power network resource management method as described in any one of the above.

[0060] A computer-readable storage medium provided in the fourth aspect of the present invention stores a computer program thereon. When the computer program is executed, the energy-saving wireless computing power network resource management method as described in any one of the above is implemented.

[0061] A computer program product provided in the fifth aspect of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, the energy-saving wireless computing power network resource management method as described in any one of the above is implemented.

[0062] It can be seen from the above technical solutions that the present invention has the following advantages:

[0063] The above solution of the present invention provides an energy-saving wireless computing power network resource management method, including: based on the mutual transmission between computing nodes of the wireless computing power network, constructing a task resource joint optimization problem with the goal of minimizing the total network energy consumption based on the task model, computing model, communication model and queue model of the wireless computing power network; transforming the task resource joint optimization problem into a Markov decision process and correspondingly constructing a deep reinforcement learning model to be trained, and iteratively training the deep reinforcement learning model to be trained with the maximization of the reward value as the goal, so as to determine the trained deep reinforcement learning model; when the wireless computing power network receives a computing task to be processed, output a target task offloading decision of the computing task to be processed through the trained deep reinforcement learning model. Based on the above solution, by introducing the mutual transmission between computing nodes, resources can be utilized more flexibly and energy consumption can be reduced. Through the designed Markov decision process, global optimization is performed through deep reinforcement learning, and network resources are fully allocated, which helps to ensure that the task offloading and resource scheduling decisions can find a balance between stability and energy efficiency. Description of the Drawings

[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0065] Figure 1 The flowchart of the steps of an energy-saving wireless computing power network resource management method provided by an embodiment of the present invention;

[0066] Figure 2 The schematic diagram of the system architecture of the wireless computing power network provided by an embodiment of the present invention;

[0067] Figure 3 The block diagram of the structure of an energy-saving wireless computing power network resource management device provided by an embodiment of the present invention. Detailed implementation manners

[0068] The embodiments of the present invention provide an energy-saving wireless computing power network resource management method and device, which are used to solve the technical problem that the traditional energy-saving optimization method of the wireless computing power network mainly focuses on the local optimization of a single target and is difficult to balance the effective utilization of computing resources and energy consumption.

[0069] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0070] Please refer to Figure 1 , Figure 1 The flowchart of the steps of an energy-saving wireless computing power network resource management method provided by an embodiment of the present invention.

[0071] An energy-saving wireless computing power network resource management method provided by the present invention includes:

[0072] Step 101: Based on the mutual transmission between computing nodes of the wireless computing power network, and based on the task model, computing model, communication model, and queue model of the wireless computing power network, construct a task resource joint optimization problem with the goal of minimizing the total network energy consumption.

[0073] It should be noted that in this embodiment, it is considered that in the wireless computing power network of the Internet of Everything, the collaborative work of computing and network resources between devices is used to support diverse application programs, such as Figure 2As shown in the figure, it is mainly divided into two parts: the access layer (Access Plane) and the application layer (Application Plane). The access layer is responsible for connecting different types of computing entities such as user devices (Users), edge servers (Edge device), and cloud data centers (Cloud Data Center). These computing entities are called computing nodes. These computing nodes have computing power (Computing power) and are connected to each other through the Internet (Internet connection), and are remotely controlled and managed based on the control signals (Control Signal) of the remote controller (Remote controller). The application layer centrally processes the requirements of different application fields such as smart industry (Smartindustry), smart agriculture (Smart agriculture), and smart home (Smart home). Specifically, the application program will submit computing requests according to its business needs, and these requests will then be automatically directed to the computing nodes in the access layer, using the set of computing nodes represents the computing node. According to the interaction between the access layer and the application layer, this embodiment focuses on how the computing node dynamically adapts to user needs to ensure the efficiency, fairness, and minimum energy consumption of resource allocation;

[0074] Different from the existing solutions that mainly focus on task offloading between edge servers and users, this embodiment introduces the inter-node transfer of computing tasks, that is, computing tasks can be dynamically scheduled among multiple computing nodes, such as between users and users, between users and edge servers, between edge servers and edge servers, etc. Based on this, a task model, a computing model, a communication model, and a queue model of the wireless computing power network are constructed. In a specific implementation, the model construction process is as follows:

[0075] 1. Task model:

[0076] 1) Task generation: Each computing node generates computing tasks according to a certain arrival rate. The generation of computing tasks follows a Poisson process, and the generated computing tasks are independent in time; assuming the task generation rate is (unit: task / second), then the size of each generated task includes the task computing requirement , the task transmission requirement and the task push requirement , that is , and the generation process of the task is represented by It is shown that, where the task computing requirement refers to the amount of computing resources required for a computing task to be executed on a computing node, the task transmission requirement refers to the amount of communication resources required for task data to be offloaded from a local computing node to other computing nodes, and the task push requirement refers to the communication resources for the task result to be returned from other computing nodes to the local computing node. In the formula, is a time slot , is a computing node of the wireless computing power network , is a computing node of the wireless computing power network , is the task generation rate, follows a Poisson distribution, is a computing node 's computing task, is a computing node 's task computing requirement, is a computing node 's task transmission requirement, is a computing node 's task push requirement;

[0077] 2) Task decision-making: After each computing task is generated, the computing node needs to decide whether to transmit the task to other computing nodes for processing or to process it on the local device; we define the task transmission decision for time slot as . When , the computing task will be processed on the local computing node . When , the computing task will be transmitted to other computing nodes for processing. It can be understood that, for the sake of easy distinction, in this embodiment, computing node is defined as the computing node for local processing, and computing node is defined as the computing node for transmitting from the local to other computing nodes for processing;

[0078] 2. Computing model. The computing task can be processed locally on the computing node or offloaded to other computing nodes for processing, and is correspondingly divided into a local computing sub-model and an other-node computing sub-model:

[0079] 1) Local computing sub-model: When the computing task is processed on the local computing node (i.e., ), then the processing time of the task in time slot can be expressed as:

[0080] (1)

[0081] In the formula, To calculate the task processing time of the computing node in a time slot ; To calculate the task computing requirement of the computing node in a time slot ; Let be the computing power of the computing node and be the set of computing nodes in the wireless computing power network, and

[0082] be the set of time slots;

[0082] Correspondingly, the energy consumption of the computing node in the time slot is expressed by the following formula:

[0083] (2)

[0084] In the formula, is the task processing energy consumption of the computing node in the time slot ; is the energy coefficient of the computing node relative to the chip architecture, and is the amount of computing resources required for the computing node to complete the task;

[0085] 2) Other node computing sub-model: When the computing task is transmitted to another computing node (i.e., ), the computing node will allocate computing resources ( ) to process the task, where represents the computing resource allocation coefficient. Therefore, the task processing time of the computing node in the time slot is:

[0086] (3)

[0087] In the formula, is the task processing time of the computing node in the time slot ; is the computing resource allocation coefficient of the computing node in the time slot ; is the computing power of the computing node ;

[0088] Correspondingly, the computing energy consumption of the computing node in the time slot is expressed by the following formula:

[0089] (4)

[0090] Wherein, is the task processing energy consumption of the computing node in time slot , is the energy consumption of the computing node representing the execution of CPU / GPU cycles, is the total number of computing cycles required to complete the task, is the computing node the amount of computing resources required to complete the task;

[0091] 3) Communication model:

[0092] In the communication model, in this embodiment, the OFDMA scheme is adopted. On this basis, the bandwidth resource is evenly divided into sub-channels of size Hz, and the set of available sub-channels for the user equipment is denoted as , and each computing node occupies at most one sub-channel, that is, inter-cell interference is not considered; in time slot , from the computing node allocated to the computing node the bandwidth is as follows:

[0093] (5)

[0094] Wherein, is the bandwidth allocated from the computing node to the computing node in time slot , is the allocation coefficient of the bandwidth resource, is the computing node total bandwidth resource;

[0095] Without ignoring the inter-cell interference, that is, multiple computing nodes use one sub-channel at the same time, use to represent the computing node in time slot allocates the sub-channel to the computing node for the inter-cell interference during task transmission, then the signal-to-noise ratio is:

[0096] (6)

[0097] Wherein, is the orthogonal sub-channel , is the set of available sub-channels, For the computing node When transmitting a task to the computing node the signal-to-noise ratio is For the computing node the transmit power is At time slot the channel gain between the computing node and the computing node is For the computing node At time slot when the orthogonal sub-channel is assigned to the computing node the inter-cell interference during task transmission is and the noise spectral density is

[0098] Based on this, the rate at which the computing node transmits a task to the computing node at time slot can be obtained as :

[0099] (7)

[0100] Correspondingly, at time slot the delay for the computing node to transmit a task to the computing node is :

[0101] (8)

[0102] Then, at time slot the transmission energy consumption for the computing node to transmit a task to the computing node is as follows

[0103] (9)

[0104] Similarly, when the computing node pushes a task to the computing node the signal-to-noise ratio at this time is

[0105] (10)

[0106] In the formula for the computing node when pushing a task to the computing node the signal-to-noise ratio is for the computing node the transmit power is

[0107] Meanwhile, the rate of pushing tasks to the computing node in time slot Computing node to the computing node can be obtained as follows: :

[0108] (11)

[0109] Then, in time slot Computing node to the computing node The task reception and processing delay when pushing tasks is as follows:

[0110] (12)

[0111] In the formula, is the task pushing requirement of the computing node in time slot Computing node ;

[0112] Correspondingly, in time slot Computing node to the computing node The pushing energy consumption when pushing tasks is as follows:

[0113] (13)

[0114] It can be understood that since the energy consumption and transmission delay caused by data transmission such as computing result feedback and task message instructions are too small compared to the size of the input task data, this embodiment assumes that these data transmissions and energy consumptions are negligible;

[0115] 4. Queue model:

[0116] When a computing task is waiting to be processed locally by a computing node or transmitted to a corresponding service on other nodes for computing, it first enters the task queue. Correspondingly, it is divided into a local queue model and an other-node queue model. The length of the queue can reflect the queuing time of the task. Limiting the task queuing time is equivalent to maintaining the long-term stability of task queuing on the computing node. Otherwise, the length of the task queue will continuously increase with the increase of time slots, and the task processing performance will also decrease accordingly;

[0117] 1) Local queue model: If the computing task is placed on the computing node it can be defined as which is equivalent to the length of the task queue to be computed on the computing node in time slot (user equipment). The value is the queue length at the beginning. When , the queue length is 0, that is , and then is updated to:

[0118] (14)

[0119] Among them, indicates that the queue length cannot be negative;

[0120] 2) Other node queue models: Similarly, denote as the queue length of the tasks to be calculated on the computing node at time slot . There is also a situation where when , the queue length is 0, that is , and then is updated to:

[0121] (15)

[0122] Based on the above task model, computing model, communication model and queue model, define as the total task delay at time slot , which includes computing delay and transmission delay:

[0123] (16)

[0124] Similar to the task delay, define as the energy consumption of the computing node to complete the task at time slot . Then:

[0125] (17)

[0126] Considering the computing tasks of each computing node at time slot , the total network energy consumption of the entire wireless computing power network at time slot can be expressed as:

[0127] (18)

[0128] The objective to be optimized in this embodiment is to minimize the energy consumption of the wireless computing power network under the specified task delay constraint (maximum tolerable delay ) and while ensuring the stability of the queues on all computing nodes; based on the transmission decision , computing resources and bandwidth resources , the target to be optimized is formulated as a MINLP problem, which is defined as the task-resource joint optimization problem and is expressed as follows:

[0129] (19)

[0130] where, is the time slot length, is the task queue length of time slot , is the expectation; in the above constraints, C1 means that each computing node only selects to process computing tasks locally or transmit them to other nodes, C2 ensures that the execution time of the task should not exceed the maximum tolerable delay, C3 and C4 respectively represent the constraints on the total amount of available computing resources and bandwidth resources, that is, the resources allocated to all computing nodes cannot exceed the total amount of resources, and C5 ensures the long-term stability of the task queues of all computing nodes.

[0131] Step 102: Convert the task-resource joint optimization problem into a Markov decision process and correspondingly construct a deep reinforcement learning model to be trained. Iteratively train the deep reinforcement learning model to be trained with the goal of maximizing the reward value to determine the trained deep reinforcement learning model.

[0132] Deep Reinforcement Learning (DRL) is an intelligent optimization method that combines reinforcement learning and deep learning. In recent years, it has received extensive attention in the fields of wireless network resource management, control, and optimization; DRL continuously adjusts the behavior strategy of the agent through interaction with the environment to maximize the long-term cumulative reward. Different from traditional reinforcement learning, DRL uses deep neural networks to model complex state spaces and action spaces and can handle high-dimensional, non-linear, and variable resource management problems. In a wireless computing power network, DRL can dynamically perceive and learn the network environment (such as link quality, computing demand, network load, etc.), and can optimize the resource allocation strategy in real time. In terms of task offloading, computing resource scheduling, and network bandwidth allocation, DRL can adjust decisions in real time according to the state of the wireless computing power network, thereby improving resource utilization, reducing energy consumption, and improving the overall performance of the system. Compared with traditional resource management methods, DRL has strong adaptive and global optimization capabilities and can make optimal decisions when facing a complex network environment.

[0133] It should be noted that the task-resource joint optimization problem P1 is transformed into a Markov decision process by constructing the corresponding state space, action space, and reward function, so that it can be optimized and solved based on the deep reinforcement learning algorithm to dynamically adjust the task offloading strategy. After constructing the corresponding deep reinforcement learning model to be trained based on the Markov decision process, the model is iteratively trained with the goal of maximizing the reward value to determine the trained deep reinforcement learning model.

[0134] In a specific implementation manner of this embodiment, the Markov decision process includes:

[0135] Define the computing power network state space with the task generation rate, task queue length, channel gain, inter-cell interference, and the amount of computing resources required for the computing node to complete the task;

[0136] Define the computing power network action space with the transmission decision variable, the computing resource allocation coefficient of the computing node, and the transmission power of the computing node;

[0137] Define the reward function by integrating the total network energy consumption, total task delay, and task queue length;

[0138] Quantify the reward value of executing the computing power network action in the computing power network action space under the computing power network state in the computing power network state space based on the reward function.

[0139] It should be noted that this embodiment models the optimization problem as a Markov decision process (MDP), including the following five core elements:

[0140] 1. The computing power network state space (State, S) that defines the state of the wireless computing power network, including:

[0141] 1) Task queue state: task arrival rate , the number of waiting tasks and ;

[0142] 2) Wireless channel state: channel gain , interference situation ;

[0143] 3) Computing requirements of the task: amount of computing resources and , deadline ;

[0144] 2. The computing power network state action space (Action, A), which comprehensively considers the joint optimization of wireless transmission power control, task offloading, and computing resource allocation to achieve system-level energy efficiency improvement. It includes:

[0145] 1) Task offloading decision (i.e., whether to perform local computing or transmit to other computing nodes);

[0146] 2) Computing resource allocation (i.e., how much CPU / GPU resources to allocate);

[0147] 3) Wireless transmission power control and (adjust the transmission power to balance energy consumption and latency);

[0148] 3. State Transition (State Transition, P) follows the following conditions:

[0149] 1) After a task is offloaded to a certain computing node, the computing resource occupancy rate of that node increases, and the task queue may change;

[0150] 2) After successful wireless transmission, the task enters the computing stage, and after the computing is completed, the task leaves the network;

[0151] 3) Since the environment is dynamic (task arrival rate, channel conditions, resource availability, etc.), the state transition is random and can be trained through experience replay and policy optimization;

[0152] 4. Reward Function (Reward, R), in this embodiment, combined with the Lyapunov optimization objective, the reward function is defined as follows:

[0153] ;

[0154] In the formula, is the reward value, is the time slot of the total network energy consumption, is the total task delay in the time slot ; is the time slot of the task queue length, is the energy consumption weight factor, is the latency weight factor, is the discount factor;

[0155] 5. Discount Factor (Discount Factor, γ\gammaγ):

[0156] Set the discount factor , to balance short-term and long-term benefits. Generally, the value is close to 1 (such as 0.99) so that reinforcement learning pays more attention to long-term system optimization.

[0157] On the basis of dynamically optimizing task offloading and resource allocation using deep reinforcement learning, this embodiment also uses Lyapunov optimization to ensure the stability of the task queue, minimize energy consumption, and improve system adaptability; under the Lyapunov optimization framework, a Lyapunov function is defined. Reflect the queue stability: , and the goal is to optimize energy consumption when the Lyapunov drift is minimized: , while DRL indirectly minimizes the Lyapunov drift by maximizing the reward value, thus ensuring queue stability and optimizing energy consumption. Existing Lyapunov optimization methods often only focus on queue stability, while this embodiment further introduces an energy consumption awareness mechanism to explicitly optimize energy consumption in the reward function, guiding the network to adjust in a more energy-efficient direction. It can be understood that the Lyapunov optimization method, as a dynamic optimization method based on stability theory, has been widely used in network resource management; Lyapunov optimization describes the stability of the system by introducing a Lyapunov function and makes dynamic adjustments according to the current state of the system to minimize delays, energy consumption, and other load metrics in the resource management process. This method is particularly suitable for situations where there are uncertainties and time-varying characteristics in the network. By transforming the optimization goal of the system into a Lyapunov stability problem that can be calculated in real time, the stability of the system is ensured during the optimization process; in WPCN, Lyapunov optimization can effectively perform resource allocation and scheduling to ensure queue stability when the system performs task offloading, and thus achieve the goal of low energy consumption. The Lyapunov method can dynamically adjust the resource allocation strategy to avoid overloading or resource waste, balance the contradiction between computing tasks and energy consumption, and thus improve the energy efficiency of the wireless computing power network to a certain extent. Many existing studies attempt to use Lyapunov optimization to achieve network stability and efficient resource allocation, but they mostly rely on explicit network models and have poor adaptability to network environments that change rapidly. In practical applications, the state and task requirements of the network are often dynamic. Existing methods often cannot make real-time adjustments in highly dynamic and uncertain environments, resulting in the inability to fully utilize network resources and even possible overconsumption of resources.

[0158] It can be understood that since the action space is mixed (discrete decision-making + continuous resource allocation), a multi-agent reinforcement learning (Multi-Agent RL) algorithm can be selected. If collaborative decision-making among multiple computing nodes is considered, independent DQN or MADDPG can be adopted. At the same time, through a distributed learning framework, the training process can be dispersed to each edge node in the network, enabling multiple computing nodes to autonomously optimize resource management, improving the scalability and real-time decision-making ability of the system, adapting to complex and dynamic wireless computing power network environments, reducing the burden of centralized computing, accelerating the real-time optimization ability of the system. The specific implementation process of the algorithm can refer to the existing technology and will not be elaborated here. In addition, this embodiment can also introduce a pre-trained model. For example, the parameters of the pre-trained model are used as the initial model parameters of the deep reinforcement learning model to be trained, combined with the Lyapunov optimization framework, and multiple rounds of optimization are carried out during the training stage, enabling fast convergence with less data and computing resources.

[0159] Step 103: When the wireless computing power network receives a computing task to be processed, output the target task offloading decision of the computing task to be processed through the trained deep reinforcement learning model.

[0160] Step 103 includes the following sub-steps:

[0161] When the wireless computing power network receives a computing task to be processed, determine the current computing power network state of the wireless computing power network;

[0162] Input the current computing power network state into the trained deep reinforcement learning model and output the target task offloading decision of the computing task to be processed.

[0163] The current computing power network state refers to the parameters characterizing the current state of the wireless computing power network, which can be understood as the task generation rate, task queue length, channel gain, inter-cell interference, and the amount of computing resources required for the computing node to complete the task included in the computing power network state space.

[0164] It should be noted that when receiving a computing task to be processed, the trained deep reinforcement learning model dynamically senses the current computing power network state of the wireless computing power network, and thus outputs the corresponding action as the target task offloading decision according to the mapping strategy from state to action learned during training. The characteristics of deep reinforcement learning can enable the resource management method to adapt in real time in a highly dynamic network environment. Therefore, it can effectively cope with the rapid changes in network load, fluctuations in task requirements, and other external changes, and achieve highly adaptive resource scheduling and task offloading optimization.

[0165] In the embodiments of the present invention, by introducing the mutual transmission between computing nodes to more flexibly utilize resources, and based on the designed Markov decision process, global multi-dimensional optimization is performed through deep reinforcement learning. It can simultaneously consider the comprehensive impacts of multiple factors such as task offloading, communication latency, computing power, load balancing, and energy efficiency. It can fully allocate various resources on the premise of ensuring network stability, avoid resource waste or system performance imbalance caused by single-objective optimization, ensure that the task offloading and resource scheduling decisions can find the best balance between stability and energy efficiency, improve task execution efficiency and reduce latency.

[0166] Please refer to Figure 3 , Figure 3 which is a structural block diagram of an energy-saving wireless computing power network resource management device provided by the embodiments of the present invention.

[0167] An energy-saving wireless computing power network resource management device provided by the present invention includes:

[0168] An optimization problem determination module 301, configured to construct a joint task-resource optimization problem with the goal of minimizing the total network energy consumption based on the mutual transmission between the computing nodes of the wireless computing power network and based on the task model, computing model, communication model, and queue model of the wireless computing power network;

[0169] A reinforcement learning module 302, configured to convert the joint task-resource optimization problem into a Markov decision process and correspondingly construct a deep reinforcement learning model to be trained, and iteratively train the deep reinforcement learning model to be trained with the goal of maximizing the reward value to determine the trained deep reinforcement learning model;

[0170] A task-resource decision module 303, configured to output a target task offloading decision for the to-be-processed computing task through the trained deep reinforcement learning model when the wireless computing power network receives the to-be-processed computing task.

[0171] Furthermore, the task model includes:

[0172] , ;

[0173] When the computing task will be processed by the local computing node When the computing task will be transmitted to other computing nodes for processing;

[0174] Wherein, is the time slot , is the computing node of the wireless computing power network , is the computing node of the wireless computing power network , is the task generation rate, follows a Poisson distribution, is the computing node 's computing task, is the task computing requirement of the computing node , is the task transmission requirement of the computing node , is the task push requirement of the computing node , is the time slot 's transmission decision variable;

[0175] The computing model includes a local computing sub-model and other node computing sub-models;

[0176] The local computing sub-model includes:

[0177] ;

[0178] ;

[0179] In the formula, is the task processing time of the computing node in the time slot , is the task computing requirement of the computing node in the time slot , is the computing power of the computing node , is the set of computing nodes in the wireless computing power network, is the set of time slots, is the task processing energy consumption of the computing node in the time slot , is the energy coefficient of the computing node relative to the chip architecture, is the amount of computing resources required for the computing node to complete the task;

[0180] The other node computing sub-model includes:

[0181] ;

[0182] ;

[0183] Among them, , ;

[0184] In the formula, is in the time slot The task processing time of the computing node , For the computing node in time slot , The computing resource allocation coefficient For the computing node The computing power For the computing node in time slot , The task processing energy consumption To represent the execution The computing node energy consumption for CPU / GPU cycles The total number of computing cycles required to complete the task For the computing node The amount of computing resources required for the computing node to complete the task

[0185] The communication model includes:

[0186] ;

[0187] ;

[0188] ;

[0189] ;

[0190] ;

[0191] ;

[0192] ;

[0193] ;

[0194] ;

[0195] In the formula, For the bandwidth allocated from the computing node To the computing node In time slot , The bandwidth allocation coefficient For the computing node The total bandwidth resource For the orthogonal subchannel , The available subchannel set For the computing node When transmitting tasks to the computing node , For the computing node The transmission power is the channel gain between the computing node and the computing node in the time slot ; is the inter-cell interference when the computing node allocates the orthogonal sub-channel to the computing node for task transmission in the time slot ; is the noise spectral density is the rate at which the computing node transmits tasks to the computing node in the time slot ; is the delay at which the computing node transmits tasks to the computing node in the time slot ; is the transmission energy consumption when the computing node transmits tasks to the computing node in the time slot ; is the signal-to-noise ratio when the computing node pushes tasks to the computing node ; is the transmission power of the computing node ; is the rate at which the computing node pushes tasks to the computing node in the time slot ; is the task reception and processing delay when the computing node pushes tasks to the computing node in the time slot ; is the task push requirement of the computing node in the time slot ; is the push energy consumption when the computing node pushes tasks to the computing node in the time slot ;

[0196] Queue model, including local queue model and other node queue models:

[0197] The local queue model includes:

[0198] ;

[0199] The other node queue models include:

[0200] ;

[0201] wherein, is the length of the task queue to be calculated on the computing node at time slot ; is the length of the task queue to be calculated on the computing node at time slot ;

[0202] Furthermore, the task resource joint optimization problem includes:

[0203] ;

[0204] wherein, , , , , , ;

[0205] wherein, is the time slot , is the computing node of the wireless computing power network , is the computing node of the wireless computing power network , is the transmission decision variable of the time slot , is the set of computing nodes of the wireless computing power network is the set of time slots is the transmission decision is the computing resource is the computing resource allocation coefficient of the computing node at time slot , is the bandwidth resource is the allocation coefficient of the bandwidth resource is the time length of the time slot is the total network energy consumption of the time slot , is the energy consumption of the computing node to complete the task at time slot , is the total task delay at time slot , is the task processing time of the computing node at time slot , is the task processing time of the computing node at time slot , To calculate the delay of task transmission to the computing node Computing node To the computing node The delay of task transmission To calculate the task reception and processing delay when the computing node Computing node To the computing node Pushes tasks To calculate the task processing energy consumption of the computing node Computing node In a time slot To calculate the transmission energy consumption of task transmission from the computing node Computing node To the computing node In a time slot To calculate the push energy consumption when the computing node Computing node To the computing node Pushes tasks Is the maximum tolerable delay Is the time slot The length of the task queue Is the expectation

[0206] Furthermore, the Markov decision process includes:

[0207] Defining the computing power network state space with the task generation rate, task queue length, channel gain, inter-cell interference, and the amount of computing resources required for the computing node to complete tasks;

[0208] Defining the computing power network action space with the transmission decision variable, the computing resource allocation coefficient of the computing node, and the transmit power of the computing node;

[0209] Defining the reward function by integrating the total network energy consumption, total task delay, and task queue length;

[0210] Quantifying the reward value of performing the computing power network action in the computing power network action space under the computing power network state in the computing power network state space based on the reward function.

[0211] Furthermore, the reward function includes:

[0212] ;

[0213] In the formula, Is the reward value Is the time slot The total network energy consumption For the time slot The total task delay Is the time slot The length of the task queue is the energy consumption weight factor, is the delay weight factor, is the discount factor.

[0214] Furthermore, the task resource decision module 303 is specifically configured to:

[0215] When the wireless computing power network receives a computing task to be processed, determine the current computing power network state of the wireless computing power network;

[0216] Input the current computing power network state into the trained deep reinforcement learning model, and output the target task offloading decision for the computing task to be processed.

[0217] An embodiment of the present invention also provides a computer device, including a memory and a processor, and a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the energy-saving wireless computing power network resource management method in any of the above embodiments.

[0218] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by the processor, the steps of the energy-saving wireless computing power network resource management method in any of the above embodiments are implemented.

[0219] An embodiment of the present invention also provides a computer program product, including a computer program / instruction, and when the computer program / instruction is executed by the processor, the steps of the energy-saving wireless computing power network resource management method in any of the above embodiments are implemented.

[0220] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0221] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.

[0222] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0223] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0224] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0225] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. An energy-saving wireless computing power network resource management method, characterized in that, Including: Based on the mutual transmission between computing nodes of the wireless computing power network, a joint task-resource optimization problem is constructed with the goal of minimizing the total network energy consumption, based on the task model, computing model, communication model, and queue model of the wireless computing power network; The joint task-resource optimization problem is transformed into a Markov decision process and a deep reinforcement learning model to be trained is correspondingly constructed, and the deep reinforcement learning model to be trained is iteratively trained with the goal of maximizing the reward value to determine the trained deep reinforcement learning model; When the wireless computing power network receives a computing task to be processed, the target task offloading decision of the computing task to be processed is output through the trained deep reinforcement learning model.

2. The energy-saving wireless computing power network resource management method according to claim 1, wherein The task model includes: , ; When the computing task will be processed on the local computing node When the computing task will be transferred to other computing nodes for processing; In the formula, is a time slot , is a computing node of the wireless computing power network , is a computing node of the wireless computing power network , is the task generation rate, is a Poisson distribution, is the computing task of the computing node , is the task computing requirement of the computing node , is the task transmission requirement of the computing node , is the task push requirement of the computing node , is a time slot is the transmission decision variable; The computing model includes a local computing sub-model and other-node computing sub-models; The local computing sub-model includes: ; ; In the formula, For the time slot Compute Node The task processing time, For the time slot Compute Node The task computing requirements, For computing nodes The computing power of is the set of computing nodes of the wireless computing network, is a set of time slots, For the time slot Compute Node The task processing energy consumption, For computing nodes Relative to the energy coefficient of the chip architecture, For computing nodes The amount of computing resources required to complete the task; The other-node computing sub-models include: ; ; Among them, , ; Wherein, is the task processing time of the computing node in time slot ; is the computing resource allocation coefficient of the computing node in time slot ; is the computing power of the computing node ; is the task processing energy consumption of the computing node in time slot ; represents the energy consumption of the computing node for executing CPU / GPU cycles, is the total number of computing cycles required to complete the task, is the amount of computing resources required for the computing node to complete the task; The communication model includes: ; ; ; ; ; ; ; ; ; Wherein, is the bandwidth allocated from computing node to computing node in time slot . is the allocation coefficient of bandwidth resources, is the total bandwidth resource of computing node . is the orthogonal sub-channel . is the set of available sub-channels, is the signal-to-noise ratio when computing node transmits tasks to computing node . is the transmit power of computing node . is the channel gain between computing node and computing node in time slot . is the inter-cell interference when computing node allocates the orthogonal sub-channel to computing node for task transmission in time slot . is the noise spectral density, is the transmission rate when computing node transmits tasks to computing node in time slot . is the transmission delay when computing node transmits tasks to computing node in time slot . is the transmission energy consumption when computing node transmits tasks to computing node in time slot . is the signal-to-noise ratio when computing node pushes tasks to computing node . is the transmit power of computing node . is the transmission rate when computing node pushes tasks to computing node in time slot . is the task reception and processing delay when computing node pushes tasks to computing node in time slot . is when computing node in time slot The task push requirements For the time slot Computing node To the computing node The push energy consumption during task push The queue model includes a local queue model and other-node queue models: The local queue model includes: ; The other-node queue models include: ; Wherein, is the length of the task queue to be calculated on the computing node at time slot , is the length of the task queue to be calculated on the computing node at time slot .

3. The energy-saving wireless computing power network resource management method according to claim 1, wherein The joint task-resource optimization problem includes: ; Among them, , , , , , ; In the formula, is a time slot , is a computing node of the wireless computing power network , is a computing node of the wireless computing power network , is a time slot 's transmission decision variable, is the set of computing nodes of the wireless computing power network, is the set of time slots, is the transmission decision, is the computing resource, is at time slot 's computing node 's computing resource allocation coefficient, is the bandwidth resource, is the allocation coefficient of the bandwidth resource, is the time length of the time slot, is time slot 's total network energy consumption, is time slot computing node 's energy consumption for task completion, is at time slot 's total task delay, is at time slot computing node 's task processing time, is at time slot computing node 's task processing time, is at time slot computing node 's task transmission delay to computing node , is at time slot computing node 's task receiving and processing delay when pushing the task to computing node , is at time slot computing node 's task processing energy consumption, is at time slot computing node 's task transmission energy consumption when transmitting the task to computing node , is at time slot computing node 's task pushing energy consumption when pushing the task to computing node , is the maximum tolerable delay, is the time slot The length of the task queue, is the expectation.

4. The energy-saving wireless computing power network resource management method according to claim 1, wherein The Markov decision process includes: Defining the computing power network state space with the task generation rate, task queue length, channel gain, inter-cell interference, and the amount of computing resources required for the computing node to complete the task; Defining the computing power network action space with the transmission decision variable, the computing resource allocation coefficient of the computing node, and the transmission power of the computing node; Defining the reward function by synthesizing the total network energy consumption, total task delay, and task queue length; Quantifying the reward value of executing the computing power network action in the computing power network action space under the computing power network state in the computing power network state space based on the reward function.

5. The energy-saving wireless computing power network resource management method according to claim 4, wherein The reward function includes: ; Wherein, is the reward value, is the time slot of the total network energy consumption, is the total task delay in the time slot , is the time slot of the task queue length, is the energy consumption weight factor, is the delay weight factor, is the discount factor.

6. The energy-saving wireless computing power network resource management method according to claim 1, wherein When the wireless computing power network receives a computing task to be processed, outputting the target task offloading decision of the computing task to be processed through the trained deep reinforcement learning model, includes: When the wireless computing power network receives a computing task to be processed, determining the current computing power network state of the wireless computing power network; Inputting the current computing power network state into the trained deep reinforcement learning model and outputting the target task offloading decision of the computing task to be processed.

7. An energy-saving wireless computing power network resource management device, characterized in that, Including: An optimization problem determination module, which is used to construct a joint task-resource optimization problem with the goal of minimizing the total network energy consumption, based on the mutual transmission between computing nodes of the wireless computing power network and based on the task model, computing model, communication model, and queue model of the wireless computing power network; A reinforcement learning module, which is used to transform the joint task-resource optimization problem into a Markov decision process and correspondingly construct a deep reinforcement learning model to be trained, and iteratively train the deep reinforcement learning model to be trained with the goal of maximizing the reward value to determine the trained deep reinforcement learning model; A task-resource decision module, which is used to output the target task offloading decision of the computing task to be processed through the trained deep reinforcement learning model when the wireless computing power network receives a computing task to be processed.

8. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the energy-saving wireless computing power network resource management method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the energy-saving wireless computing power network resource management method according to any one of claims 1-6 are implemented.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the energy-saving wireless computing power network resource management method according to any one of claims 1-6 are implemented.