Task allocation method, device, equipment and program product

By combining the optimization objective function of communication, computing and energy consumption models with multi-agent reinforcement learning algorithms in mobile edge computing networks, task allocation is optimized, solving the problem that existing technologies cannot simultaneously achieve high-efficiency computing and low energy consumption, and improving task processing efficiency and energy utilization.

CN120929196APending Publication Date: 2025-11-11CHINA MOBILE GROUP JILIN BRANCH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410557631.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing task allocation methods cannot simultaneously provide efficient computing services and meet the requirements of low energy consumption. In particular, in mobile edge computing networks, existing methods lack consideration for the actual needs of users for multiple computing tasks and ignore network energy consumption.

Method used

An optimization objective function based on communication, computation, and energy consumption models is adopted, and a multi-agent reinforcement learning optimization algorithm is used to allocate tasks. By constructing formulas for upload and download carrier transmission rates, task processing time, and energy consumption models, the computation tasks, spectrum, and power allocation are optimized.

Benefits of technology

It achieves efficient computing services in mobile edge computing networks while reducing energy consumption, improving the efficiency of computing tasks and the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929196A_ABST
    Figure CN120929196A_ABST
Patent Text Reader

Abstract

The invention discloses a task allocation method and device, equipment and a program product, and relates to the technical field of communication networks, and the task allocation method comprises the steps: obtaining a to-be-allocated calculation task set; based on a preset optimization objective function, an optimization objective problem is determined according to the to-be-allocated calculation task set, and the optimization objective function is constructed based on a preset communication model, a calculation model and an energy consumption model; and based on a preset multi-agent reinforcement learning optimization algorithm, performing task allocation according to the optimization target problem and the to-be-allocated calculation task set, and obtaining a task allocation scheme. The task allocation based on the multi-agent reinforcement learning optimization algorithm is realized, the technical problem that the existing task allocation cannot meet the requirement of low energy consumption while giving consideration to efficient computing service is solved, and the energy consumption is reduced while the computing task processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication network technology, and in particular to a task allocation method, apparatus, device, and program product. Background Technology

[0002] With the rapid growth of data volume in user devices and applications, the demand for computing power and low latency in wireless communication networks has brought enormous challenges. To address this issue, edge computing has emerged, which can simultaneously meet users' requirements for efficient computing services and low transmission latency.

[0003] Both edge computing and traditional cloud computing distribute all or part of the tasks of dispersed users to centralized computer servers, and then send the data back to the users from the servers.

[0004] However, existing computing task allocation methods only consider the computing task requirements of a single agent for a certain type, or they only focus on improving the efficiency of network computing services while ignoring high energy consumption, and cannot meet the requirements of low energy consumption while providing efficient computing services.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main objective of this application is to provide a task allocation method, apparatus, device, and program product, which aims to solve the technical problem that existing task allocation methods cannot simultaneously provide high-efficiency computing services and meet the requirements of low energy consumption.

[0007] To achieve the above objectives, this application proposes a task allocation method, which includes:

[0008] Get the set of computation tasks to be assigned;

[0009] Based on a preset optimization objective function, the optimization objective problem is determined according to the set of computing tasks to be assigned. The optimization objective function is constructed based on a preset communication model, computing model, and energy consumption model.

[0010] Based on a preset multi-agent reinforcement learning optimization algorithm, tasks are allocated according to the optimization objective problem and the set of computational tasks to be assigned, and a task allocation scheme is obtained.

[0011] In one embodiment, before the step of determining the optimization objective problem based on the set of computational tasks to be assigned according to a preset optimization objective function, the method further includes:

[0012] A communication model is constructed based on the preset upload and download link carriers;

[0013] Based on the preset proportion of collaborative tasks, a computing model is constructed according to the preset edge computing tasks, local computing tasks, and collaborative computing tasks.

[0014] Based on the preset energy consumption of the computing cycle, an energy consumption model is constructed according to the edge computing task and the collaborative computing task;

[0015] Based on the aforementioned communication model, computation model, and energy consumption model, an optimization objective function is constructed.

[0016] In one embodiment, the step of constructing a communication model based on preset upload link carriers and download link carriers includes:

[0017] Based on the upload link carrier, the upload power loss rate, upload frequency allocation parameters, and upload transmit power are obtained, and an upload carrier transmission rate formula is generated based on the upload power loss rate, upload frequency allocation parameters, and upload transmit power.

[0018] Based on the download link carrier, obtain the download power loss rate, download frequency allocation parameters, and download transmit power, and generate a download carrier transmission rate formula based on the download power loss rate, download frequency allocation parameters, and download transmit power.

[0019] A communication model is constructed based on the formulas for the upload carrier transmission rate and the download carrier transmission rate.

[0020] In one embodiment, the step of constructing a computing model based on a preset collaborative task ratio and according to preset edge computing tasks, local computing tasks, and collaborative computing tasks includes:

[0021] Based on the preset task data volume and server task data processing ratio, server processing time parameters and server transmission time parameters are generated, and an edge task processing time formula is constructed based on the server processing time parameters and server transmission time parameters.

[0022] Based on the task data volume and the local device task data processing ratio, local device processing time parameters and local device transmission time parameters are generated, and a local task processing time formula is constructed based on the local device calculation time parameters and local device transmission time parameters.

[0023] Based on the collaborative task ratio, a collaborative task processing time formula is constructed according to the edge task processing time formula and the local task processing time formula.

[0024] A calculation model is constructed based on the edge task processing time formula, the local task processing time formula, and the collaborative task processing time formula.

[0025] In one embodiment, the step of obtaining a task allocation scheme by allocating tasks according to the optimization objective problem and the set of computational tasks to be assigned based on a preset multi-agent reinforcement learning optimization algorithm includes:

[0026] A mobile edge computing model is constructed based on the aforementioned optimization objective problem;

[0027] The set of computing tasks to be assigned is input into the mobile edge computing model for the following processing:

[0028] The set of computing tasks to be assigned is defined as a multi-agent system, which includes actions, environmental states, and reward values.

[0029] Based on a preset learning step value, the multi-agent is repeatedly subjected to action selection, environmental state update, and reward value calculation to iteratively optimize the multi-agent reinforcement learning optimization algorithm until a preset termination condition is met. The current action is then output, and a task allocation scheme is obtained based on the current action.

[0030] In one embodiment, the step of constructing a mobile edge computing model based on the optimization objective problem includes:

[0031] Based on the preset time energy consumption weight, environmental state elements are generated according to the calculation model and energy consumption model.

[0032] Based on the optimization objective problem, action elements are generated, and a reward function is constructed based on the environmental state elements and action elements.

[0033] A mobile edge computing model is constructed based on the environmental state elements, action elements, and reward functions.

[0034] In one embodiment, the task allocation method is applied to a task allocation platform, which includes an edge server and a user's local device;

[0035] The method based on a preset multi-agent reinforcement learning optimization algorithm, after the step of allocating tasks according to the optimization objective problem and the set of computational tasks to be assigned, and obtaining a task allocation scheme, further includes:

[0036] According to the task allocation scheme, the computing tasks in the set of computing tasks to be allocated are allocated to the edge server and / or the user's local device for task processing.

[0037] Furthermore, to achieve the above objectives, this application also proposes a task allocation device, which includes:

[0038] The task acquisition module is used to acquire a set of computing tasks to be assigned.

[0039] The problem determination module is used to determine the optimization objective problem based on the set of computing tasks to be assigned, according to a preset optimization objective function. The optimization objective function is constructed based on a preset communication model, computing model, and energy consumption model.

[0040] The task allocation module is used to allocate tasks based on a preset multi-agent reinforcement learning optimization algorithm, according to the optimization objective problem and the set of computational tasks to be allocated, and to obtain a task allocation scheme.

[0041] In addition, to achieve the above objectives, this application also proposes a task allocation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the task allocation method as described above.

[0042] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the task allocation method described above.

[0043] This application provides a task allocation method. First, it obtains a set of computational tasks to be allocated; then, it constructs an optimization objective function based on a communication model, a computational model, and an energy consumption model; it determines an optimization objective problem based on the set of computational tasks to be allocated; and finally, it determines the optimization objective of a multi-agent reinforcement learning optimization algorithm based on the optimization objective function constructed according to the combination of models. Based on the multi-agent reinforcement learning optimization algorithm, it allocates tasks according to the optimization objective problem and the set of computational tasks to be allocated, thereby obtaining a task allocation scheme that satisfies the optimization objective.

[0044] In summary, this application achieves task allocation based on multi-agent reinforcement learning optimization algorithm by combining the optimization objective function of communication model, computing model and energy consumption model with multi-agent reinforcement learning optimization algorithm. This achieves task allocation based on multi-agent reinforcement learning optimization algorithm, overcomes the technical defect of existing task allocation that cannot simultaneously achieve high-efficiency computing service and low-energy consumption, improves the processing efficiency of computing tasks and reduces energy consumption. Attached Figure Description

[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating the first embodiment of the task allocation method of this application.

[0048] Figure 2 This is a flowchart illustrating the second embodiment of the task allocation method of this application.

[0049] Figure 3 This is a flowchart illustrating the reinforcement learning optimization algorithm involved in Embodiment 3 of this application;

[0050] Figure 4 This is a schematic diagram of the module structure of the task allocation device according to an embodiment of this application;

[0051] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the task allocation method in the embodiments of this application.

[0052] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0053] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0054] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0055] The main solution of this application embodiment is: to obtain a set of computing tasks to be assigned; to determine an optimization objective problem based on a preset optimization objective function, wherein the optimization objective function is constructed based on a preset communication model, computing model and energy consumption model; and to allocate tasks according to the optimization objective problem and the set of computing tasks to be assigned based on a preset multi-agent reinforcement learning optimization algorithm, thereby obtaining a task allocation scheme.

[0056] With the rapid growth of data volume from 5G user mobile devices and applications, the demands for computing power and low latency in wireless communication networks have posed significant challenges. To address this issue, mobile edge computing has emerged, simultaneously meeting users' requirements for efficient computing services and low-latency transmission. Both edge computing and traditional cloud computing distribute all or part of the computing tasks of dispersed users to centralized computer servers, and then transmit the data back to the users, thus solving the problem of insufficient computing power for dispersed user application devices.

[0057] Deploying mobile edge computing networks requires comprehensive consideration of issues such as server and user computing task partitioning, spectrum allocation, and power settings to ensure the network can provide users with more efficient computing services and lower data transmission latency. Computing tasks in mobile edge computing networks are divided into three types: edge computing server tasks, user-local computing tasks, and collaborative computing tasks (computing tasks jointly performed by servers and user devices). Wireless communication between base stations and users utilizes orthogonal spectrum, which is divided into uplink orthogonal subcarriers and downlink orthogonal subcarriers. Furthermore, the transmit power of wireless communication between base stations and users is positively correlated with the data transmission rate; however, excessive transmit power undoubtedly leads to excessive energy consumption.

[0058] Therefore, enabling mobile edge computing networks to provide efficient computing services with low latency while minimizing energy consumption is a crucial challenge in their deployment. To achieve these goals, existing solutions primarily include:

[0059] (1) Model the communication and computing tasks in the network, and use the convex optimization method to allocate computing tasks and resources, so as to minimize the computing latency and data transmission latency in the network, and achieve the goal of efficient computing and low latency.

[0060] (2) The method of using reinforcement learning combined with Boltzmann greedy strategy is used to optimize the allocation of computing tasks, spectrum and transmit power configuration of mobile edge computing network, so as to achieve efficient computing and low latency data transmission.

[0061] While existing methods for optimizing computing tasks and resource allocation in mobile edge computing networks can achieve good optimization results, they generally lack attention to optimizing the overall attributes of network deployment. For example, they may only focus on energy consumption optimization, or ignore the actual needs of users for different types of computing tasks, only considering the needs of users for a certain type of computing task, or unilaterally focus on improving the efficiency of network computing services while ignoring the fact of high energy consumption. Therefore, there is still a lack of edge computing network task and resource allocation optimization methods that take into account both the actual needs of users for multiple computing tasks and the network energy consumption.

[0062] Furthermore, from the perspective of optimization methods, the following problems still exist:

[0063] (1) Conventional resource allocation convex optimization methods have high computational complexity, and once the network situation changes, such as adding a new user, the problem needs to be solved again to maintain the optimality of the result, which brings huge overhead.

[0064] (2) Reinforcement learning optimization methods can solve the defects of convex optimization. Even when the environment changes, new optimization results can still be obtained quickly based on historical experience, which has the advantage of low computational overhead. However, existing reinforcement learning network optimization is based on a single agent, that is, the optimization method is limited to individual nodes in the network, rather than multiple nodes cooperating. Because the allocation of resources by a node in the network often affects the external environment, thereby affecting the resource allocation results of other nodes, it will make it difficult for single agent reinforcement learning to converge or the performance after convergence cannot be guaranteed.

[0065] In summary, current mobile edge computing network optimization methods, from a business perspective, lack a task and resource allocation optimization method that considers both the actual needs of users for multiple computing tasks and the network's energy consumption. From a theoretical perspective, mobile edge computing networks with multiple users are multi-agent environments, and current optimization methods based on single-agent reinforcement learning are insufficient to solve the optimization problems in multi-agent environments.

[0066] This application optimizes the computational tasks, spectrum, and power allocation in the network by combining the optimization objective function of the communication model, computation model, and energy consumption model with a game theory-based multi-agent reinforcement learning method. Ultimately, it minimizes the computational consumption and data transmission time in the network, meeting the high standards required by users. It balances efficient computing services with low energy consumption, overcoming the technical shortcomings of existing task allocation methods that cannot simultaneously meet the requirements of efficient computing services and low energy consumption. This improves the efficiency of computing task processing while reducing energy consumption.

[0067] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or network device capable of performing the above functions. The following description uses a network device as an example to illustrate this embodiment and the subsequent embodiments.

[0068] Based on this, embodiments of this application provide a task allocation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the task allocation method of this application.

[0069] In this embodiment, the task allocation method includes steps S10 to S30:

[0070] Step S10: Obtain the set of computing tasks to be assigned;

[0071] It should be noted that the set of computing tasks to be assigned refers to the set of computing tasks to be assigned to a device for execution, including one or more computing tasks of the same or different types that need to be assigned, wherein the device executing the task can be an edge server in an edge computing network or a user's local device.

[0072] It is understandable that in order to provide basic data for subsequent task allocation, it is necessary to obtain the set of computing tasks to be allocated. Therefore, step S10 is performed to avoid resource waste and uneven allocation, thereby improving the system's load balancing and overall execution efficiency.

[0073] Step S20: Based on the preset optimization objective function, determine the optimization objective problem according to the set of computing tasks to be assigned. The optimization objective function is constructed based on the preset communication model, computing model and energy consumption model.

[0074] It should be noted that the optimization objective function is constructed from a pre-defined communication model, computation model, and energy consumption model. The specific construction method may involve combining algorithms with these models, such as greedy algorithms, genetic algorithms, or neural networks. The communication model is the communication-related model in the optimization objective, the computation model is the computation-processing-related model in the optimization objective, and the energy consumption model is the energy-related model in the optimization objective.

[0075] It is understandable that, in order to clarify the direction of subsequent reinforcement learning optimization, the optimization target problem is determined based on the set of computational tasks to be assigned using a pre-set optimization objective function. Therefore, step S20 is performed to optimize computational efficiency and energy consumption, thereby improving the overall optimization effect.

[0076] Step S30: Based on a preset multi-agent reinforcement learning optimization algorithm, tasks are allocated according to the optimization target problem and the set of computational tasks to be allocated, and a task allocation scheme is obtained.

[0077] It should be noted that the multi-agent reinforcement learning optimization algorithm is a reinforcement learning method based on multiple agents to solve complex optimization problems. Reinforcement learning is a machine learning method used to enable agents to learn how to make decisions to maximize cumulative rewards through interaction with the environment. In reinforcement learning, agents learn how to make the optimal sequence of actions in a specific environment by trying different actions and observing the environment's feedback (reward signal). Reinforcement learning algorithms typically include the following elements:

[0078] Intelligent agent: that is, a learner, who learns how to make optimal decisions through interaction with the environment.

[0079] Environment: The external environment in which the agent exists, whose state may change over time and affect the agent's actions.

[0080] State: Describes a specific situation or configuration of the agent's environment; it can be discrete or continuous.

[0081] Action: The action or decision that an agent can take in a specific state.

[0082] Reward: The immediate feedback given by the environment after the agent performs an action, used to evaluate the quality of the action.

[0083] Policy: The method by which an agent selects an action based on the current state. It can be deterministic or probabilistic, or it can be an algorithmic choice based on the current state, such as a greedy algorithm.

[0084] The goal of reinforcement learning is to find an optimal policy that maximizes the cumulative reward of an agent during its interaction with the environment. Common reinforcement learning algorithms include Q-learning, SARSA, Deep Q Network (DQN), and Policy Gradient. This application employs a multi-agent reinforcement learning algorithm, capable of handling situations where multiple agents cooperate or compete in the same environment. By learning and optimizing their respective policies, agents can achieve mutually beneficial cooperation or adapt to the competitive strategies of their opponents, aligning with game theory principles. Agents can adjust their policies and behaviors based on changes in the environment and the actions of other agents, exhibiting a degree of adaptability and flexibility. They can adapt to complex, dynamic, and uncertain environments, learning adaptive policies to cope with various challenges and changes.

[0085] It is understandable that, since task allocation in edge computing networks involves many complex factors such as communication, computing and energy consumption, using a pre-defined multi-agent reinforcement learning optimization algorithm, combined with the optimization objective problem and the set of computing tasks to be allocated, to allocate tasks and obtain the optimal task allocation scheme, step S30 can avoid the uneven resource utilization caused by a simple task allocation scheme, thereby improving the overall task execution efficiency.

[0086] In this embodiment, task allocation is performed by combining the objective function optimization and the multi-agent reinforcement learning optimization algorithm to achieve effective utilization of system resources and improve task execution efficiency. This can effectively solve the task allocation problem in edge computing networks and optimize the overall performance and reliability of the system.

[0087] In one feasible implementation, steps S201 to S204 may be included before step S20:

[0088] Step S201: Construct a communication model based on the preset upload link carrier and download link carrier;

[0089] To predict and optimize communication performance between nodes and ensure high efficiency and reliability of data transmission, a communication model is established based on the upload and download link carriers to estimate key performance indicators of communication.

[0090] It should be noted that the upload link carrier and download link carrier refer to the carrier signals used to transmit data during communication. In this embodiment, they can refer to the communication transmission between the user's local device and the edge server. The communication model is a mathematical model used to describe the communication behavior between nodes, and can represent indicators such as communication transmission latency, bandwidth, packet loss rate, and transmission rate.

[0091] It is understandable that, since communication plays an important role in a multi-node computing environment, step S201 can construct an accurate communication model based on the upload link carrier and the download link carrier, thereby avoiding the problem of inefficient data transmission caused by communication latency and insufficient bandwidth in subsequent task allocation, and improving the overall efficiency of computing task processing.

[0092] Step S202: Based on the preset collaborative task ratio, construct a computing model according to the preset edge computing tasks, local computing tasks, and collaborative computing tasks;

[0093] It should be noted that the collaborative task ratio represents the proportion of edge computing tasks, local computing tasks, and collaborative computing tasks among different agents. The computing model represents the indicators of different types and proportions of computing tasks, such as computing power and processing time, providing a basis for using multi-agent reinforcement learning optimization algorithms to achieve task allocation, enabling tasks to be executed in the best way on different agents and improving the system's computing efficiency.

[0094] It is understandable that different types of computing tasks are processed and transmitted between edge servers and user local devices. Different task allocations will lead to different overall processing efficiency. Therefore, step S202 can construct an accurate computing model based on the preset edge computing tasks, local computing tasks and collaborative computing tasks, so as to serve as one of the optimization targets for subsequent multi-agent reinforcement learning, thereby avoiding resource waste and low task execution efficiency, and thus improving the overall computing efficiency of the system.

[0095] Step S203: Based on the preset computing cycle energy consumption, construct an energy consumption model according to the edge computing task and the collaborative computing task;

[0096] It should be noted that the energy consumption per computing cycle refers to the energy consumed by a computing node within one computing cycle. The energy consumption model describes the energy consumption model of the edge server during task execution and may include indicators such as power consumption and energy efficiency.

[0097] Understandably, since energy consumption is a significant cost of edge server computing, step S203 can construct an accurate energy consumption model based on the preset computing cycle energy consumption and the types of computing tasks related to the edge server. This avoids increased costs and environmental burden caused by excessive energy consumption, thereby improving the overall energy efficiency of the system.

[0098] Step S204: Based on the communication model, computing model and energy consumption model, construct an optimization objective function.

[0099] It should be noted that the optimization objective function is a mathematical function used to measure system performance and efficiency. It may include indicators such as data transmission efficiency, computing time and energy consumption, providing guidance for task scheduling and resource allocation, so that the system can reach an optimal state under multiple indicators.

[0100] It is understandable that since the optimization objective function directly affects the task allocation result, step S204 can construct an accurate optimization objective function based on the indicators of the communication model, computing model and energy consumption model, so that the task allocation is more in line with the overall performance of the system and the requirements for reducing energy consumption, thereby improving the overall optimization level of the system.

[0101] In this embodiment, an optimization target algorithm is constructed by comprehensively considering indicators such as communication, computing and energy consumption, so as to formulate a suitable task scheduling strategy, optimize the overall performance and efficiency of the system, solve technical problems such as resource allocation and energy management during task allocation, and improve the overall competitiveness and sustainable development capability of the edge computing network.

[0102] In one possible implementation, step S201 may include steps S2011 to S2013:

[0103] Step S2011: Based on the upload link carrier, obtain the upload power loss rate, upload frequency allocation parameters and upload transmit power, and generate the upload carrier transmission rate formula based on the upload power loss rate, upload frequency allocation parameters and upload transmit power.

[0104] Step S2012: Based on the download link carrier, obtain the download power loss rate, download frequency allocation parameters, and download transmit power, and generate a download carrier transmission rate formula based on the download power loss rate, download frequency allocation parameters, and download transmit power.

[0105] Step S2013: Based on the upload carrier transmission rate formula and the download carrier transmission rate formula, construct a communication model.

[0106] It should be noted that the power loss rate refers to the power loss rate of the uplink or downlink carrier during transmission, affecting the strength and stability of data transmission. The frequency allocation parameter refers to the carrier's allocation in the spectrum, affecting the carrier's frequency bandwidth and transmission rate. The transmit power refers to the carrier's transmit power, directly affecting the signal transmission distance and quality. The carrier transmission rate formula can be constructed by combining the above parameters with parameters generated during communication, such as communication distance, attenuation parameters, and communication noise.

[0107] It is understandable that, due to the different characteristics of the upload link carrier and the download link carrier, performing steps S2011, S2012 and S2013 can avoid inaccurate transmission rates or degraded communication quality caused by ignoring communication-related characteristics during the communication process, thereby improving the performance stability and reliability of communication.

[0108] In this implementation, these parameters are determined using mathematical modeling or experimental measurement methods based on actual data and communication standards. Then, upload and download rate formulas are constructed. Based on these formulas, along with other communication parameters, a comprehensive communication model is built to predict the communication performance between the user's local device and the edge server. By integrating the upload and download carrier rate formulas, along with other communication parameters, the complete communication model can more comprehensively describe the communication behavior and performance between nodes. This allows for more accurate prediction and optimization of the communication performance between the device and the server, avoiding resource waste and performance degradation caused by inaccurate communication parameters, thereby improving the overall communication efficiency and performance of the system.

[0109] In one possible implementation, step S202 may include steps S2021 to S2024:

[0110] Step S2021: Based on the preset task data volume and server task data processing ratio, generate server processing time parameters and server transmission time parameters, and construct an edge task processing time formula based on the server processing time parameters and server transmission time parameters.

[0111] It should be noted that the task data volume refers to the amount of data contained in a task to be processed, the server task data processing ratio refers to the proportion of the task set processed on the server, the server processing time parameter refers to the parameter of the time required for the server to process the task, and the server transmission time parameter is the parameter of the time required for the data to be transmitted from the server to the user's local device.

[0112] It is understandable that, due to the uncertainty of edge task processing time on the server and the subsequent data transmission, step S2021 can avoid insufficient processing time for tasks on the server, thereby improving system stability and performance.

[0113] Step S2022: Based on the task data volume and the local device task data processing ratio, generate local device processing time parameters and local device transmission time parameters, and construct a local task processing time formula based on the local device calculation time parameters and local device transmission time parameters.

[0114] It should be noted that the local device task data processing ratio refers to the proportion of tasks processed on the user's local device, the local device processing time parameter describes the time required for the user's local device to process the task, and the local device transmission time parameter is the time required for data to be transmitted from the user's local device to the server.

[0115] It is understandable that, due to the uncertainty of the processing process on the user's local device and the subsequent data transmission, step S2022 can avoid the situation where the task takes too long to process on the local device or the transmission delay is high, resulting in low overall processing efficiency, thereby improving the execution efficiency and response speed of the local task.

[0116] Step S2023: Based on the collaborative task ratio, construct the collaborative task processing time formula according to the edge task processing time formula and the local task processing time formula;

[0117] It should be noted that the collaborative task ratio refers to the proportion of tasks processed by the edge server and the user's local device, and the collaborative task processing time formula describes the time required for the edge server and the user's local device to process tasks together.

[0118] It is understandable that, since the proportion of tasks processed by edge servers and user local devices is uncertain and computing tasks can be processed simultaneously on servers or devices, the processing time of tasks affects each other between servers and devices. Therefore, performing step S2023 can avoid resource waste or excessive task execution time caused by unreasonable task allocation, thereby improving the overall efficiency of task processing and resource utilization.

[0119] Step S2024: Construct a calculation model based on the edge task processing time formula, the local task processing time formula, and the collaborative task processing time formula.

[0120] It should be noted that the calculation model described is a mathematical model that describes the task processing time and resource allocation of different task allocation schemes, and uses task processing time as the evaluation index.

[0121] Understandably, due to the lack of a comprehensive task processing time estimation and resource allocation scheme, step S2024 is necessary to combine the remaining models during subsequent task allocation, evaluate the comprehensive impact of task processing time and other model evaluation indicators on task allocation, avoid improper task allocation leading to overall performance degradation or resource waste, and thus improve the overall performance and efficiency of computing task processing.

[0122] In this embodiment, by combining edge task processing time, local task processing time, and collaborative task processing time, and utilizing the computing resources of edge servers and user local devices, a comprehensive task processing time estimate is provided. This helps determine the allocation of computing tasks between edge servers and user local devices, optimizes task scheduling and resource allocation, and improves the overall performance and efficiency of task processing.

[0123] In one feasible implementation, step S203 may include step S2031:

[0124] Step S2031: Based on the collaborative task ratio and computation cycle energy consumption, construct an energy consumption model according to the task quantity parameters of the edge computing task and the collaborative computing task.

[0125] It is understandable that step S2031 is necessary because the energy consumption of edge computing tasks and collaborative computing tasks needs to be considered. This avoids ignoring the energy consumption impact of edge computing tasks and collaborative computing tasks when constructing the energy consumption model, thereby improving the accuracy and applicability of the energy consumption model.

[0126] In this embodiment, an energy consumption model is constructed based on the proportion of collaborative tasks and the energy consumption of the computing cycle, according to the task quantity parameters of edge computing tasks and collaborative computing tasks. The model takes into account the energy consumption of the edge server when executing tasks, as well as the proportion of computing tasks between different allocation schemes. Thus, the energy consumption model is incorporated into the comprehensive task processing time estimation and resource allocation scheme, which can more comprehensively consider the energy consumption during task execution, thereby better optimizing the energy consumption and performance of task allocation.

[0127] The above is only one feasible implementation of step S203 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S203.

[0128] In one feasible implementation, the task allocation method is applied to a task allocation platform, which includes an edge server and user local devices. Step S30 may be followed by step S40:

[0129] Step S40: According to the task allocation scheme, the computing tasks in the set of computing tasks to be allocated are allocated to the edge server and / or the user's local device for task processing.

[0130] It is understandable that since the task allocation scheme needs to be applied to the task allocation platform, step S40 is performed after step S30. This can effectively allocate the computing tasks in the set of computing tasks to be allocated to the edge server and / or the user's local device, thereby achieving effective allocation and execution of task processing.

[0131] It should be noted that the network device described as the execution subject in this embodiment can be a network base station or a server with task allocation function, which can establish communication connections and transmit data with edge servers and user local devices.

[0132] In this embodiment, the task distribution process is realized through a task allocation platform. Tasks can be dynamically allocated to appropriate devices for processing according to specific task allocation schemes, so as to achieve full utilization of system resources and efficient execution of task processing.

[0133] This embodiment provides a task allocation method that combines the optimization objective function of a communication model, a computing model, and an energy consumption model with a multi-agent reinforcement learning optimization algorithm to allocate tasks. This method balances efficient computing services with low energy consumption, achieving task allocation based on a multi-agent reinforcement learning optimization algorithm. It overcomes the technical deficiency of existing task allocation methods that cannot simultaneously meet the requirements of efficient computing services and low energy consumption, thereby improving the efficiency of computing task processing while reducing energy consumption.

[0134] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S30 includes steps S301 to S304:

[0135] Step S301: Construct a mobile edge computing model based on the optimization objective problem;

[0136] It should be noted that the optimization objective problem refers to the objective problem that needs to be solved or optimized, determined by the objective function based on the set of computing tasks to be assigned, such as the task allocation problem in a mobile edge computing environment. The mobile edge computing model is a model describing task allocation in a mobile edge computing environment. The construction process of this model is based on adjusting existing problems to adapt to specific mobile edge computing scenarios and abstracting the practical problem into a mathematical model.

[0137] Understandably, due to the complexity of the optimization problem, step S301, which transforms the optimization target problem into a manageable model, avoids simplistic treatment of the task allocation problem, allows for a better understanding of the essence of the problem, and provides a foundation for subsequent iterative optimization, thereby improving the accuracy of the problem analysis.

[0138] Step S302: The set of computing tasks to be assigned is input into the mobile edge computing model for the following processing:

[0139] By combining the actual set of computing tasks with the mobile edge computing model, the task set can be analyzed and processed within the mobile edge computing model in order to find the optimal task allocation scheme.

[0140] Step S303: Define the set of computing tasks to be assigned as a multi-agent, wherein the multi-agent includes actions, environmental states and reward values;

[0141] It should be noted that the term "multi-agent" refers to multiple agent entities for computational tasks in the algorithm, which can learn and make decisions independently. In the multi-agent reinforcement learning optimization algorithm, defining the set of computational tasks to be assigned as multiple agents is to model the task allocation problem as a multi-agent system, so that each user corresponding to a task has its own agent, thereby enabling multiple agents to achieve the overall task allocation optimization goal through interaction and learning.

[0142] Furthermore, it's important to note that in reinforcement learning, actions are the behaviors taken by the agent. Therefore, defining actions, environmental states, and reward values ​​can avoid ambiguous agent behavior, thereby improving the accuracy and interpretability of the system's decisions. The environmental state reflects the situation and conditions in which the agent exists. Changing actions during reinforcement learning alters the environmental state, preventing decisions based on outdated information and improving the agent's understanding and adaptability to the environment. The reward value is the feedback the agent receives based on its actions, preventing blind exploration and ineffective behaviors, thus improving learning efficiency and the effectiveness of the results.

[0143] Understandably, due to the problem of uneven computational tasks in task allocation, step S303 is performed to transform the task allocation problem into an optimization objective of multiple agents. The task allocation process is optimized using reinforcement learning algorithms. By cooperating and learning among agents, the task allocation effect is optimized, which can avoid the system performance from being overloaded by some devices, thereby improving the overall efficiency and stability of the system.

[0144] Step S304: Based on a preset learning step value, repeatedly perform action selection, environmental state update, and reward value calculation on the multi-agent to iteratively optimize the multi-agent reinforcement learning optimization algorithm until a preset termination condition is met, output the current action, and obtain a task allocation scheme based on the current action.

[0145] It should be noted that the learning step value is a parameter that controls the speed and stability of the learning process. The iterative optimization is a task allocation process that combines the algorithmic ideas of reinforcement learning and game theory. By repeatedly adjusting the actions of the multi-agent, updating the environment and state of the multi-agent based on the actions, and calculating the reward value based on the selected actions and the updated environment state, the iterative optimization of the multi-agent reinforcement learning algorithm is completed by updating the parameters of the multi-agent reinforcement learning optimization algorithm. When the termination condition of the iterative optimization is met, the current action is output to obtain the optimized task allocation scheme. The preset termination condition is the condition that specifies when the iterative optimization process should terminate, such as reaching a preset threshold or the number of iterations.

[0146] For example, when the reinforcement learning algorithm is the Q-learning algorithm, during the iterative optimization process, the Q value of the current iteration can be calculated based on the selected action, the updated environment state, and the calculated reward value, combined with the Q value obtained in the previous iteration. The iteration can be terminated if the Q value meets a preset threshold. When the iteration termination condition is met, the action selected in this iteration is used as the optimal allocation strategy to obtain an optimized task allocation scheme.

[0147] It is understandable that in order to find the optimal task allocation scheme that balances task processing efficiency and energy consumption, step S304 is performed. This step involves continuously adjusting the task allocation scheme through an iterative optimization algorithm, gradually bringing it closer to the optimal solution, i.e., the action at the end of the iterative optimization, so as to obtain the best task allocation scheme that satisfies the optimization objective.

[0148] In one feasible implementation, step 301 includes steps S3011 to S3013:

[0149] Step S3011: Based on the preset time energy consumption weight, generate environmental state elements according to the calculation model and energy consumption model;

[0150] It should be noted that the time and energy consumption weights are the weights of the computation model and the energy consumption model in the environmental state elements. They reflect the importance of various influencing factors in the task allocation process, including the processing time requirements of the computation task and the energy consumption characteristics of the equipment. This provides a basis for the agent's decision-making, enabling it to take into account the processing efficiency requirements of the computation task and energy consumption factors, so as to make appropriate task allocation decisions.

[0151] It is understandable that since it is impossible to balance task processing efficiency and equipment energy consumption, step S3011 is performed to avoid excessive energy consumption of certain equipment leading to energy waste, while improving task processing efficiency and thus improving the overall energy utilization rate of computing tasks.

[0152] Step S3012: Generate action elements according to the optimization target problem, and construct a reward function according to the environmental state elements and action elements;

[0153] It should be noted that the optimization objective problem is the objective that needs to be optimized in task allocation, such as minimizing energy consumption, minimizing task processing time (i.e. maximizing task completion efficiency), etc. The action element is the action that the agent can take, such as adjusting the proportion of carrier spectrum, power allocation and cooperative task allocation. The reward function is used to evaluate the quality of each action of the agent and guide the agent's learning and optimization process.

[0154] Understandably, since the task allocation strategy is not intelligent enough, step S3012 is performed to avoid low task completion efficiency caused by unreasonable task allocation, thereby improving the overall efficiency and performance of the system.

[0155] Step S3013: Based on the environmental state elements, action elements, and reward function, construct a mobile edge computing model.

[0156] It is understandable that since the environment state element, action element, and reward function together constitute the key components of the reinforcement learning algorithm framework, performing step S3013 can avoid information loss and inconsistency in the model building process, thereby improving the modeling quality and performance of the mobile edge computing model.

[0157] In this embodiment, based on the concepts of reinforcement learning and mobile edge computing, the optimization objective problem is combined with the environmental state, actions and rewards in reinforcement learning to achieve intelligent task allocation and optimization, thereby improving the performance and efficiency of the system.

[0158] In this embodiment, reinforcement learning algorithms are used to optimize the task allocation process. By cooperating and learning among agents, the task allocation effect can be optimized, which can solve problems such as uneven task allocation, energy waste, and low task completion efficiency, thereby improving the overall performance and reliability of the system.

[0159] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, the task allocation method may include:

[0160] Edge computing networks provide end users with low-latency and high-efficiency computing services. Because the network's computing servers are located at the cellular base stations provided by operators, the computing service is coupled with wireless transmission. Therefore, efficiently configuring edge computing networks is more complex than cloud computing, requiring consideration of issues such as compute partitioning, wireless spectrum allocation, and power settings. How to enable edge computing networks to provide high-efficiency computing services while minimizing energy consumption is a highly challenging issue.

[0161] This embodiment proposes an edge computing collaborative optimization method based on game theory multi-agent reinforcement learning. It constructs a network communication model, a computing model, and an energy consumption model. Using these network models, the user is treated as a single agent, and the computational task partitioning, spectrum, and power allocation in the network are optimized. Ultimately, the computational consumption and data transmission time in the network are minimized to meet the user's high-standard requirements.

[0162] To simplify the optimization problem in this proposal, the edge computing network is divided into unit networks centered on a single base station, and the optimization problem within the unit networks is addressed. The number of users within a unit network is set to M, with each user having one computing task. Each base station connects to an edge computing server, providing offline computing power services to users and ultimately transmitting data to them via the base station. This includes the following steps:

[0163] Step 1: Constructing a communication model

[0164] Each network element is configured with U = {1, 2, ..., u} orthogonal subcarriers for uplink links and D = {1, 2, ..., d} orthogonal subcarriers for downlink links, and the bandwidth of each uplink subcarrier is W. U The download link subcarrier bandwidth is W. D Then, the data transmission rates of user m on upload link carrier u and download link carrier d are respectively:

[0165]

[0166]

[0167] in, and These refer to the transmit power of user m on the uplink carrier u and the downlink carrier d, respectively. and These refer to the power loss rates of user m on the upload link carrier u and the download link carrier d, respectively. and It is the Rayleigh fading parameter, g m This refers to the distance between user m and the base station. It is a common type of Gaussian white noise. It uploads subcarrier frequency allocation parameters, when When, the uplink carrier representing user m is u, when At that time, the uplink subcarrier u was not allocated to user m. Similarly, when When, the download link carrier representing user m is d, when At that time, the download link subcarrier d was not allocated to user m.

[0168] Step 2: Constructing the computational model

[0169] Each user is assigned a computational task with a data volume of γ. m In edge computing networks, computing tasks are divided into three categories: edge computing tasks, user-local computing tasks, and collaborative computing tasks. The processing times for these three types of tasks are as follows:

[0170] 1. Edge computing tasks: These tasks are all handled by edge computing servers. After processing the tasks, the edge servers send the data back to user m. The time required is:

[0171]

[0172] Where F is the CPU clock frequency of the edge server, ω is the number of CPU cycles the server processes per unit of data, and α is the proportion of data obtained after the task data has been processed by the server, which is a constant. This refers to the processing time of the task on the server. It is the time it takes for the processed data to be transmitted to user m.

[0173] 2. User Local Computing Tasks: Local tasks are processed only by the user's local device. For example, compressing files on the local device before transmitting them to the base station takes approximately:

[0174]

[0175] Among them, f m ω is the CPU clock frequency of the user's local device. m β is the number of CPU cycles required to process a unit of data on the user's local device, and β is the proportion of task data obtained after processing by the local device, which is a constant. This refers to the processing time of the task on the local device. It is the time it takes for the processed data to be transmitted to the server.

[0176] 3. Collaborative computing tasks: These tasks can be randomly assigned to edge computing servers and user m's local device, allowing them to work together to complete the task. The required time is:

[0177]

[0178] Where, ε m ∈[0,1] represents the proportion of collaborative tasks assigned to user m's local device. This refers to the computation time of the collaborative task on the user's local device. It is the time it takes for a task assigned to the edge server to be transmitted from the user device to the server. The computation time allocated to edge server tasks. This is the time it takes for the server to send the data back to the user after processing the task. Since the server and the user can process collaborative computing tasks simultaneously, the total time used for task computation is determined by the maximum time between local computation time and the time used for data transmission.

[0179] Step 3: Construct an energy consumption model

[0180] The number of edge computing tasks, local computing tasks, and collaborative computing tasks in the unit network are set as follows: and Since energy consumption in edge computing networks mainly occurs during server operation, the energy consumption model in the cell network is as follows:

[0181]

[0182] Where P is the energy consumption per unit cycle of server operation.

[0183] Step 4: Use multi-agent reinforcement learning to optimize task allocation and subcarrier spectrum allocation

[0184] Based on the modeling process from step one to step three, the final optimization problem is to minimize the time required for task processing and the energy consumption of the edge computing server by allocating carrier spectrum, power, and collaborative task partitioning.

[0185] Reinforcement learning algorithms consist of three parts: state χ, action χ, and action χ. And the reward R. In this proposal, users in the network unit are treated as multiple agents, and the reinforcement learning optimization steps for each agent are as follows:

[0186] State: The environment state element x∈χ is defined as the weighted sum of the processing time and energy consumption of all users' computational tasks.

[0187]

[0188] Where K t and K e These are the weighted values ​​of task processing time and energy consumption, respectively.

[0189] Actions: Given the subcarrier spectrum to be optimized, the power allocation and cooperative task allocation ratios, the agent's actions... It consists of three parts: (i) Subcarrier spectrum allocation in This refers to the upload link where the spectrum of the i-th upload carrier is allocated to user m. (ii) The download link for user m where the spectrum of the j-th download carrier is allocated; in This is the maximum power setting. The power allocated to the i-th upload carrier is p. U , The power allocated to the j-th download carrier is p. D ;(iii)∈=ε m This refers to the proportion allocated to local devices in collaborative tasks.

[0190] Reward: The goal of optimization is to minimize the individual's task processing time and the energy consumed in processing the task, based on state x and action. The reward function is set as follows:

[0191]

[0192] Where x is the current state, which is the sum of time and energy consumption of the network unit, and t m ∈ That is, the time required to process user tasks, E m =aω m γ m P+b(1-ε m )ω m γ m P represents the energy consumed by the user in processing the task, where a = 1 indicates that the user's task is an edge computing task, otherwise a = 0; b = 1 indicates that the user's task is a collaborative computing task, otherwise b = 0.

[0193] Step 5: Update the Q-value table for multi-agent Q-learning.

[0194] According to step four, user m takes an action based on the current environmental state x. Receive reward points This is used to update the Q-value in the state-action table. It's worth noting that this proposal does not use the general reinforcement learning method for updating Q-values. Instead, based on the fact that edge computing networks are multi-agent environments, it employs a game theory-based multi-agent reinforcement learning method to update the Q-values.

[0195]

[0196] Where α is the learning step. For example... Figure 3 As shown, Figure 3 This is a flowchart illustrating the reinforcement learning optimization algorithm involved in Embodiment 3 of this application. The action selection employs a greedy algorithm, that is, selecting the action a corresponding to the maximum Q(x) value in state x. When the Q value does not meet the set threshold, the process returns to step four to obtain a new state and action, and calculates the reward value. When the Q value meets the set threshold or reaches the maximum number of iterations, the iteration terminates, and the action is returned.

[0197] Through the above steps, the optimization problem of task, resource, and power allocation in edge computing networks is modeled using multi-agent reinforcement learning, ultimately obtaining the optimal allocation scheme to meet users' needs for efficient computing services and low latency, while achieving optimal energy consumption.

[0198] In this embodiment, a collaborative optimization method for edge computing based on game theory multi-agent reinforcement learning is proposed. The decisions of user agents in edge computing networks often affect the decisions of other agents. Therefore, when the reinforcement learning method designed by a single agent is directly applied to a multi-agent environment, it is easy to cause the learning process to be difficult to converge or the performance after convergence to be difficult to guarantee. In this embodiment, the idea of ​​game theory is used to solve the multi-agent optimization decision problem, and finally the reward function after all agents in the environment play the game reaches the equilibrium point.

[0199] This application also provides a task allocation device, please refer to... Figure 4 The task allocation device includes:

[0200] Task acquisition module 10 is used to acquire a set of computing tasks to be assigned;

[0201] Problem determination module 20 is used to determine the optimization objective problem based on the set of computing tasks to be assigned, according to a preset optimization objective function. The optimization objective function is constructed based on a preset communication model, computing model and energy consumption model.

[0202] The task allocation module 30 is used to allocate tasks based on the preset multi-agent reinforcement learning optimization algorithm, according to the optimization target problem and the set of computational tasks to be allocated, and to obtain a task allocation scheme.

[0203] Optionally, the problem determination module 20 is further configured to:

[0204] A communication model is constructed based on the preset upload and download link carriers;

[0205] Based on the preset proportion of collaborative tasks, a computing model is constructed according to the preset edge computing tasks, local computing tasks, and collaborative computing tasks.

[0206] Based on the preset energy consumption of the computing cycle, an energy consumption model is constructed according to the edge computing task and the collaborative computing task;

[0207] Based on the aforementioned communication model, computation model, and energy consumption model, an optimization objective function is constructed.

[0208] Optionally, the problem determination module 20 is further configured to:

[0209] Based on the upload link carrier, the upload power loss rate, upload frequency allocation parameters, and upload transmit power are obtained, and an upload carrier transmission rate formula is generated based on the upload power loss rate, upload frequency allocation parameters, and upload transmit power.

[0210] Based on the download link carrier, obtain the download power loss rate, download frequency allocation parameters, and download transmit power, and generate a download carrier transmission rate formula based on the download power loss rate, download frequency allocation parameters, and download transmit power.

[0211] A communication model is constructed based on the formulas for the upload carrier transmission rate and the download carrier transmission rate.

[0212] Optionally, the problem determination module 20 is further configured to:

[0213] Based on the preset task data volume and server task data processing ratio, server processing time parameters and server transmission time parameters are generated, and an edge task processing time formula is constructed based on the server processing time parameters and server transmission time parameters.

[0214] Based on the task data volume and the local device task data processing ratio, local device processing time parameters and local device transmission time parameters are generated, and a local task processing time formula is constructed based on the local device calculation time parameters and local device transmission time parameters.

[0215] Based on the collaborative task ratio, a collaborative task processing time formula is constructed according to the edge task processing time formula and the local task processing time formula.

[0216] A calculation model is constructed based on the edge task processing time formula, the local task processing time formula, and the collaborative task processing time formula.

[0217] Optionally, the problem determination module 20 is further configured to:

[0218] A mobile edge computing model is constructed based on the aforementioned optimization objective problem;

[0219] The set of computing tasks to be assigned is input into the mobile edge computing model for the following processing:

[0220] The set of computing tasks to be assigned is defined as a multi-agent system, which includes actions, environmental states, and reward values.

[0221] Based on a preset learning step value, the multi-agent is repeatedly subjected to action selection, environmental state update, and reward value calculation to iteratively optimize the multi-agent reinforcement learning optimization algorithm until a preset termination condition is met. The current action is then output, and a task allocation scheme is obtained based on the current action.

[0222] Optionally, the problem determination module 20 is further configured to:

[0223] Based on the preset time energy consumption weight, environmental state elements are generated according to the calculation model and energy consumption model.

[0224] Based on the optimization objective problem, action elements are generated, and a reward function is constructed based on the environmental state elements and action elements.

[0225] A mobile edge computing model is constructed based on the environmental state elements, action elements, and reward functions.

[0226] Optionally, the task allocation module 30 is further configured to:

[0227] According to the task allocation scheme, the computing tasks in the set of computing tasks to be allocated are allocated to edge servers and / or user local devices for task processing.

[0228] The task allocation device provided in this application, employing the task allocation method described in the above embodiments, can solve the technical problem that existing task allocation methods cannot simultaneously provide high-efficiency computing services and meet the requirement of low energy consumption. Compared with the prior art, the beneficial effects of the task allocation device provided in this application are the same as those of the task allocation method described in the above embodiments, and other technical features in the task allocation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0229] This application provides a task allocation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the task allocation method described in the first embodiment above.

[0230] The following is for reference. Figure 5The diagram illustrates a structural schematic of a task allocation device suitable for implementing embodiments of this application. The task allocation device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The task allocation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0231] like Figure 5 As shown, the task allocation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the task allocation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the task allocation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows task allocation devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0232] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0233] The task allocation device provided in this application, employing the task allocation method described in the above embodiments, can solve the technical problem that existing task allocation methods cannot simultaneously provide high-efficiency computing services and meet the requirement of low energy consumption. Compared with the prior art, the beneficial effects of the task allocation device provided in this application are the same as those of the task allocation method described in the above embodiments, and other technical features of this task allocation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0234] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0235] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0236] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0237] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0238] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0239] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the task allocation method described above.

[0240] The computer program product provided in this application can solve the technical problem that existing task allocation methods cannot simultaneously achieve high-efficiency computing services and meet the requirement of low energy consumption. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the task allocation method provided in the above embodiments, and will not be repeated here.

[0241] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A task allocation method, characterized in that, The method includes: Get the set of computation tasks to be assigned; Based on a preset optimization objective function, the optimization objective problem is determined according to the set of computing tasks to be assigned. The optimization objective function is constructed based on a preset communication model, computing model, and energy consumption model. Based on a preset multi-agent reinforcement learning optimization algorithm, tasks are allocated according to the optimization objective problem and the set of computational tasks to be assigned, and a task allocation scheme is obtained.

2. The method as described in claim 1, characterized in that, Before the step of determining the optimization objective problem based on the set of computational tasks to be assigned according to the preset optimization objective function, the method further includes: A communication model is constructed based on the preset upload and download link carriers; Based on the preset proportion of collaborative tasks, a computing model is constructed according to the preset edge computing tasks, local computing tasks, and collaborative computing tasks. Based on the preset energy consumption of the computing cycle, an energy consumption model is constructed according to the edge computing task and the collaborative computing task; Based on the aforementioned communication model, computation model, and energy consumption model, an optimization objective function is constructed.

3. The method as described in claim 2, characterized in that, The step of constructing a communication model based on preset upload link carriers and download link carriers includes: Based on the upload link carrier, the upload power loss rate, upload frequency allocation parameters, and upload transmit power are obtained, and an upload carrier transmission rate formula is generated based on the upload power loss rate, upload frequency allocation parameters, and upload transmit power. Based on the download link carrier, obtain the download power loss rate, download frequency allocation parameters, and download transmit power, and generate a download carrier transmission rate formula based on the download power loss rate, download frequency allocation parameters, and download transmit power. A communication model is constructed based on the formulas for the upload carrier transmission rate and the download carrier transmission rate.

4. The method as described in claim 2, characterized in that, The step of constructing a computing model based on a preset collaborative task ratio and according to preset edge computing tasks, local computing tasks, and collaborative computing tasks includes: Based on the preset task data volume and server task data processing ratio, server processing time parameters and server transmission time parameters are generated, and an edge task processing time formula is constructed based on the server processing time parameters and server transmission time parameters. Based on the task data volume and the local device task data processing ratio, local device processing time parameters and local device transmission time parameters are generated, and a local task processing time formula is constructed based on the local device calculation time parameters and local device transmission time parameters. Based on the collaborative task ratio, a collaborative task processing time formula is constructed according to the edge task processing time formula and the local task processing time formula. A calculation model is constructed based on the aforementioned edge task processing time formula, local task processing time formula, and collaborative task processing time formula.

5. The method as described in claim 1, characterized in that, The steps of the preset multi-agent reinforcement learning optimization algorithm, which allocates tasks according to the optimization objective problem and the set of computational tasks to be assigned, and obtains a task allocation scheme, include: A mobile edge computing model is constructed based on the aforementioned optimization objective problem; The set of computing tasks to be assigned is input into the mobile edge computing model for the following processing: The set of computing tasks to be assigned is defined as a multi-agent system, which includes actions, environmental states, and reward values. Based on a preset learning step value, the multi-agent is repeatedly subjected to action selection, environmental state update, and reward value calculation to iteratively optimize the multi-agent reinforcement learning optimization algorithm until a preset termination condition is met. The current action is then output, and a task allocation scheme is obtained based on the current action.

6. The method as described in claim 5, characterized in that, The step of constructing a mobile edge computing model based on the optimization objective problem includes: Based on the preset time energy consumption weight, environmental state elements are generated according to the calculation model and energy consumption model. Based on the optimization objective problem, action elements are generated, and a reward function is constructed based on the environmental state elements and action elements. A mobile edge computing model is constructed based on the environmental state elements, action elements, and reward functions.

7. The method according to any one of claims 1 to 6, characterized in that, The task allocation method is applied to a task allocation platform, which includes an edge server and user local devices. The method based on a preset multi-agent reinforcement learning optimization algorithm, after the step of allocating tasks according to the optimization objective problem and the set of computational tasks to be assigned, and obtaining a task allocation scheme, further includes: According to the task allocation scheme, the computing tasks in the set of computing tasks to be allocated are allocated to the edge server and / or the user's local device for task processing.

8. A task allocation device, characterized in that, The device includes: The task acquisition module is used to acquire a set of computing tasks to be assigned. The problem determination module is used to determine the optimization objective problem based on the set of computing tasks to be assigned, according to a preset optimization objective function. The optimization objective function is constructed based on a preset communication model, computing model, and energy consumption model. The task allocation module is used to allocate tasks based on a preset multi-agent reinforcement learning optimization algorithm, according to the optimization objective problem and the set of computational tasks to be allocated, and to obtain a task allocation scheme.

9. A task allocation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the task allocation method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the task allocation method as described in any one of claims 1 to 7.