Method and apparatus for task offloading and communication management in mobile edge computing, and medium
By constructing a Q-table based on reinforcement learning, the decision-making process for task offloading and communication management is optimized, solving the problem of low task success rate in IoT systems and achieving efficient task offloading and communication management.
Patent Information
- Application Number
- PCT/CN2024/104521
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2026-01-15
AI Technical Summary
Existing technologies struggle to effectively optimize task offloading and communication management decisions in IoT systems, resulting in low task success rates.
By acquiring the current state information of the network environment and mobile devices, a first Q-table and a second Q-table are constructed using reinforcement learning methods. The first decision action is determined for task unloading and computing resource allocation, and the second decision action is determined for communication resource allocation, speed control, and communication mode, thereby optimizing the decision-making process.
It improves the success rate of tasks in IoT systems and enables optimized decision-making for task offloading and communication management.
Smart Images

Figure CN2024104521_15012026_PF_FP_ABST
Abstract
Description
Methods, apparatus, and media for task offloading and communication management in mobile edge computing Technical Field
[0001] This disclosure relates to the field of communications, and in particular to a method, apparatus, and medium for task offloading and communication management in mobile edge computing. Background Technology
[0002] Driven by technologies such as the Internet of Things (IoT) and 5G, edge computing is gradually becoming one of the key technologies for future computing and data processing. Edge computing allows computing power to be located closer to the network edge, which can greatly shorten the data transmission path and improve data processing efficiency and speed.
[0003] Task offloading decision-making, as one of the key technologies in edge computing, can use intelligent algorithms to scientifically determine which tasks to keep locally on the device and which tasks to offload to edge nodes for processing, thereby optimizing system performance, saving network resources, and ensuring that data processing needs can be met.
[0004] Summary of the Invention
[0005] To optimize task offloading and communication management decisions in edge computing and improve task success rates in IoT systems, this disclosure provides a method, apparatus, and medium for task offloading and communication management in mobile edge computing.
[0006] According to a first aspect of the present disclosure, a method for task offloading and communication management in mobile edge computing is provided, the method comprising:
[0007] Obtain first status information, which is used to indicate the current status of the network environment and the current status of the mobile device;
[0008] Based on the first state information, multiple first decision actions, along with the benefit information and probability distribution information of each first decision action, are determined from a pre-established first Q table. Multiple second decision actions, along with the benefit information and probability distribution information of each second decision action, are determined from a pre-established second Q table. The first decision actions are used to indicate the action decision results of task unloading and computational resource allocation, and the second decision actions are used to indicate the action decision results of communication resource allocation, speed control, and communication mode. The first Q table and the second Q table are obtained through reinforcement learning.
[0009] Based on the benefit information and probability distribution information of each first decision action, a first target action is determined from the plurality of first decision actions; based on the benefit information and probability distribution information of each second decision action, a second target action is determined from the plurality of second decision actions.
[0010] According to a second aspect of the present disclosure, a communication device is provided, comprising:
[0011] The processing module is configured to acquire first status information, which is used to indicate the current status of the network environment and the current status of the mobile device.
[0012] The processing module is further configured to determine multiple first decision actions, as well as the benefit information and probability distribution information of each first decision action, from a pre-established first Q table based on the first state information; and to determine multiple second decision actions, as well as the benefit information and probability distribution information of each second decision action, from a pre-established second Q table. The first decision actions are used to indicate the action decision results of task unloading and computing resource allocation, and the second decision actions are used to indicate the action decision results of communication resource allocation, speed control, and communication mode. The first Q table and the second Q table are obtained through reinforcement learning.
[0013] The processing module is further configured to determine a first target action from the plurality of first decision actions based on the benefit information and probability distribution information of each first decision action, and to determine a second target action from the plurality of second decision actions based on the benefit information and probability distribution information of each second decision action.
[0014] According to a third aspect of the present disclosure, a communication device is provided, comprising:
[0015] One or more processors;
[0016] The communication device is used to execute the task offloading and communication management method in mobile edge computing as described in the first aspect.
[0017] According to a fourth aspect of the present disclosure, a storage medium is provided that stores instructions that, when executed on a communication device, cause the communication device to perform the task offloading and communication management method in mobile edge computing as described in the first aspect.
[0018] In this embodiment, by acquiring first state information indicating the current state of the network environment and the current state of the mobile device, multiple first decision actions, along with the benefit information and probability distribution information of each first decision action, are determined from a first Q-table obtained through reinforcement learning based on the first state information. Similarly, multiple second decision actions, along with the benefit information and probability distribution information of each second decision action, are determined from a pre-established second Q-table. Furthermore, based on the benefit information and probability distribution information of each first decision action, a first target action is determined from the multiple first decision actions. Finally, based on the benefit information and probability distribution information of each second decision action, a second target action is determined from the multiple second decision actions. The first decision actions indicate the decision results for task offloading and computing resource allocation, while the second decision actions indicate the decision results for communication resource allocation, speed control, and communication mode. This optimizes task offloading and communication management decisions, improving the task success rate in the Internet of Things (IoT) system.
[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0021] Figure 1A is a schematic diagram of a system architecture of an edge computing network according to an embodiment of the present disclosure.
[0022] Figure 1B is a schematic diagram of a system architecture of a mobile edge computing network according to an embodiment of the present disclosure.
[0023] Figure 2 is an interactive schematic diagram of a task offloading and communication management method in mobile edge computing according to an embodiment of the present disclosure.
[0024] Figure 3A is a schematic flowchart illustrating the construction process of a first Q table according to an embodiment of the present disclosure.
[0025] Figure 3B is a flowchart illustrating the construction process of a second Q table according to an embodiment of the present disclosure.
[0026] Figure 4 is a schematic diagram illustrating the principle of a multi-agent Q-learning algorithm according to an embodiment of the present disclosure.
[0027] Figure 5 is a schematic diagram of the structure of the communication device proposed in the embodiments of this disclosure.
[0028] Figure 6A is a schematic diagram of the structure of the communication device 6100 proposed in an embodiment of this disclosure.
[0029] Figure 6B is a schematic diagram of the structure of the chip 6200 proposed in the embodiment of this disclosure. Detailed Implementation
[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0031] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of at least one associated listed item.
[0032] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various messages, these messages should not be limited to these terms. These terms are used only to distinguish messages of the same type from one another. For example, without departing from the scope of this disclosure, a first message may also be referred to as a second message, and similarly, a second message may also be referred to as a first message. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0033] This disclosure presents a method, apparatus, and medium for task offloading and communication management in mobile edge computing.
[0034] In a first aspect, embodiments of this disclosure propose a method for task offloading and communication management in mobile edge computing, the method comprising:
[0035] Obtain first status information, which is used to indicate the current status of the network environment and the current status of the mobile device;
[0036] Based on the first state information, multiple first decision actions, along with the benefit information and probability distribution information of each first decision action, are determined from a pre-established first Q table. Multiple second decision actions, along with the benefit information and probability distribution information of each second decision action, are determined from a pre-established second Q table. The first decision actions are used to indicate the action decision results of task unloading and computational resource allocation, and the second decision actions are used to indicate the action decision results of communication resource allocation, speed control, and communication mode. The first Q table and the second Q table are obtained through reinforcement learning.
[0037] Based on the benefit information and probability distribution information of each first decision action, a first target action is determined from the plurality of first decision actions; based on the benefit information and probability distribution information of each second decision action, a second target action is determined from the plurality of second decision actions.
[0038] In the above embodiments, by acquiring first state information indicating the current state of the network environment and the current state of the mobile device, multiple first decision actions, along with the benefit information and probability distribution information of each first decision action, are determined from a first Q-table obtained in advance through reinforcement learning based on the first state information. Similarly, multiple second decision actions, along with the benefit information and probability distribution information of each second decision action, are determined from a pre-established second Q-table. Furthermore, based on the benefit information and probability distribution information of each first decision action, a first target action is determined from the multiple first decision actions. Finally, based on the benefit information and probability distribution information of each second decision action, a second target action is determined from the multiple second decision actions. The first decision actions indicate the decision results for task offloading and computing resource allocation, while the second decision actions indicate the decision results for communication resource allocation, speed control, and communication mode, thereby optimizing task offloading and communication management decisions and improving the task success rate in the IoT system.
[0039] In conjunction with some embodiments of the first aspect, in some embodiments, the first Q table construction process includes:
[0040] Based on the first state information of multiple mobile devices, determine the first decision action for each mobile device;
[0041] Determine the reward function and probability distribution information for the first decision action of each mobile device;
[0042] Based on the reward function and probability distribution information of the first decision action of each mobile device, update the reward information and probability distribution information of each first decision action in the first Q table to be built.
[0043] In the above embodiments, by providing possible implementation methods for constructing the first Q-table, the purpose of constructing the first Q-table is achieved through reinforcement learning methods, so as to ensure the smooth progress of the task unloading decision and computing resource allocation decision process.
[0044] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first decision action for each mobile device based on the first state information of multiple mobile devices includes:
[0045] Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the first decision action of each mobile device.
[0046] In the above embodiments, a method for determining the first decision action based on the first state information of the mobile device is provided, so as to achieve the purpose of making task unloading decisions and computing resource allocation decisions based on the first state information, so as to ensure the smooth progress of subsequent processes.
[0047] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first decision action of each mobile device based on the first state information of the plurality of mobile devices using a greedy algorithm includes:
[0048] Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the first decision action of each mobile device, and each first decision action corresponds to a probability distribution information.
[0049] Based on the first decision action of the plurality of mobile devices, the communication resources to be allocated to the plurality of mobile devices are determined;
[0050] Based on the communication resources to be allocated to the plurality of mobile devices and the first communication resources of the network, it is determined whether to adopt the first decision action determined by the greedy algorithm, wherein the first communication resources are used to indicate the available communication resources of the mobile edge computing (MEC) device in the network environment.
[0051] In the above embodiments, a possible implementation method is provided to determine the first decision action based on the first state information using a greedy algorithm, so as to achieve the purpose of task unloading decision and computing resource allocation decision through the greedy algorithm, so as to ensure the smooth progress of subsequent processes.
[0052] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether to adopt the determined first decision action based on the communication resources to be allocated to the plurality of mobile devices and the first communication resources of the network includes any one of the following:
[0053] If the sum of the communication resources to be allocated to the plurality of mobile devices is less than or equal to the first communication resource, the first decision action determined by the greedy algorithm is adopted.
[0054] If the total amount of communication resources to be allocated to the plurality of mobile devices is greater than the first communication resource, the reward function of the first decision action of each mobile device is set to the first value, and the first decision action of each mobile device is redetermined through a greedy algorithm.
[0055] In the above embodiments, by determining whether to adopt the first decision action determined by the greedy algorithm, the rationality and legality of the adopted first decision action are guaranteed, and the security and legality of the task unloading decision and computing resource allocation decision process are guaranteed.
[0056] In conjunction with some embodiments of the first aspect, in some embodiments, updating the reward information and probability distribution information of each first decision action in the first Q table to be established based on the reward function and probability distribution information of the first decision action of each mobile device includes:
[0057] Based on the reward function of the first decision action of each mobile device, as well as the pre-set learning rate and discount factor, update the revenue information of each first decision action in the first Q table to be built;
[0058] Based on the probability distribution information of the first decision action of each mobile device, update the probability distribution information of each first decision action in the first Q table to be built.
[0059] In the above embodiments, possible implementation methods are provided for updating the benefit information and probability distribution information of each first decision action in the first Q table, so as to ensure the timeliness and accuracy of the benefit information and probability distribution information recorded in the first Q table.
[0060] In conjunction with some embodiments of the first aspect, in some embodiments, updating the revenue information of each first decision action in the first Q table to be built based on the reward function of the first decision action of each mobile device, and a pre-set learning rate and discount factor, includes:
[0061] Update the payoff information for each first decision action in the first Q table to be built according to the following formula:
[0062] Among them, Q u,v,1 (s u,v,m,n (t),a u,v,m,n (t) represents the payoff information after the first decision action, Q. u,v,1 (s u,v,m,n (t′),a u,v,m,n (t′)) represents the payoff information before the first decision action is updated, r u,v,m,n (t) represents the reward function for the first decision action, λ u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the coverage area of the m-th access point (AP), and n represent the n-th MEC device.
[0063] In the above embodiments, possible implementation methods are provided for updating the revenue information of each first decision action in the first Q table, so as to ensure that the revenue information of each first decision action in the first Q table can be updated, thereby ensuring the timeliness and accuracy of the revenue information recorded in the first Q table.
[0064] In conjunction with some embodiments of the first aspect, in some embodiments, the first state information includes at least one of the following:
[0065] Current AP coverage area;
[0066] The availability of current time-slot wireless communication;
[0067] The current data size generated by the mobile device in the current time slot;
[0068] The size of the calculation result for the current time slot mobile device;
[0069] The available computing resources of the current time-slot mobile device;
[0070] The speed of the mobile device in the current time slot;
[0071] The MEC device selected by the mobile device in the previous time slot;
[0072] Available computing resources for all MEC devices in the current time slot.
[0073] In the above embodiments, by providing the types of information that the first state information may include, task unloading decisions and computing resource allocation decisions can be made based on multiple types of first state information, thereby improving the flexibility of the task unloading decision process and the computing resource allocation decision process.
[0074] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0075] Determine second state information, which is used to indicate a partial state of the network environment in the next time slot and a partial state of the mobile device in the next time slot.
[0076] In the above embodiments, by determining second state information indicating a portion of the network environment in the next time slot and a portion of the mobile device in the next time slot, decisions on subsequent task offloading and computing resource allocation can be made based on the second state information of each mobile device.
[0077] In conjunction with some embodiments of the first aspect, in some embodiments, the second state information includes at least one of the following:
[0078] The next time slot AP coverage area;
[0079] Availability of next time-slot wireless communication;
[0080] The size of the data generated by the mobile device in the next time slot;
[0081] The size of the calculation result for the next time slot mobile device;
[0082] Available computing resources for the next time slot mobile device;
[0083] The speed of mobile devices in the next time slot;
[0084] The MEC device selected by the mobile device in the current time slot;
[0085] Available computing resources for all MEC devices in the next time slot.
[0086] In the above embodiments, by providing the types of information that the second state information may include, multiple possible states of the network and mobile devices in the next time slot can be determined after the task offloading decision and computing resource allocation decision are completed in the current time slot, so as to improve the flexibility and diversity of the state prediction process.
[0087] In conjunction with some embodiments of the first aspect, in some embodiments, the second Q-table construction process includes:
[0088] If the AP connected to any of the multiple mobile devices changes, a second decision action is determined for each mobile device based on the first state information of the multiple mobile devices.
[0089] Determine the payoff information and probability distribution information for the second decision action of each mobile device;
[0090] Based on the payoff information and probability distribution information of the second decision action for each mobile device, update the payoff information and probability distribution information of each second decision action in the second Q table to be built.
[0091] In the above embodiments, by providing possible implementation methods for constructing the second Q-table, the purpose of constructing the second Q-table is achieved through reinforcement learning methods, so as to ensure the smooth progress of communication resource allocation decision, speed control decision and communication mode decision processes.
[0092] In conjunction with some embodiments of the first aspect, in some embodiments, determining the second decision action for each mobile device based on the first state information of the plurality of mobile devices includes:
[0093] Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the second decision action of each mobile device.
[0094] In the above embodiments, a method for determining the second decision action based on the second state information of the mobile device is provided, so as to achieve the purpose of communication resource allocation decision, speed control decision and communication mode decision based on the second state information, so as to ensure the smooth progress of subsequent processes.
[0095] In conjunction with some embodiments of the first aspect, in some embodiments, determining the second decision action of each mobile device based on the first state information of the plurality of mobile devices using a greedy algorithm includes:
[0096] Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the second decision action of each mobile device, and each second decision action corresponds to a probability distribution information.
[0097] Based on the second decision action of the plurality of mobile devices, the communication bandwidth to be allocated to the plurality of mobile devices is determined;
[0098] Based on the communication bandwidth to be allocated to the plurality of mobile devices and the first communication bandwidth of the network, it is determined whether to adopt the second decision action determined by the greedy algorithm, wherein the communication bandwidth is used to indicate the available communication bandwidth of relay devices and / or MEC devices in the network environment.
[0099] In the above embodiments, a possible implementation method is provided to determine the second decision action based on the second state information using a greedy algorithm, so as to achieve the purpose of communication resource allocation decision, speed control decision and communication mode decision through the greedy algorithm, so as to ensure the smooth progress of subsequent processes.
[0100] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether to adopt the second decision action determined by the greedy algorithm based on the communication bandwidth to be allocated to the plurality of mobile devices and the first communication bandwidth of the network includes any one of the following:
[0101] If the combined result of the communication bandwidth to be allocated to the multiple mobile devices does not exceed the first communication bandwidth, then the second decision action determined by the greedy algorithm is adopted.
[0102] If the combined result of the communication bandwidth to be allocated to the multiple mobile devices exceeds the first communication bandwidth, the reward function of the second decision action of each mobile device is set to the first value, and the second decision action of each mobile device is redetermined through a greedy algorithm.
[0103] In the above embodiments, by determining whether to adopt the second decision action determined by the greedy algorithm, the rationality and legality of the adopted second decision action are ensured, thereby guaranteeing the security and legality of the communication resource allocation decision, speed control decision, and communication mode decision process.
[0104] In conjunction with some embodiments of the first aspect, in some embodiments, updating the payoff information and probability distribution information of each second decision action in the second Q table to be established based on the payoff information and probability distribution information of each second decision action of each mobile device includes:
[0105] Based on the reward function of the second decision action for each mobile device, as well as the pre-set learning rate and discount factor, update the revenue information of each second decision action in the second Q table to be built;
[0106] Based on the probability distribution information of the second decision action of each mobile device, update the probability distribution information of each first decision action in the second Q table to be established.
[0107] In the above embodiments, possible implementation methods are provided for updating the benefit information and probability distribution information of each second decision action in the second Q table, so as to ensure the timeliness and accuracy of the benefit information and probability distribution information recorded in the second Q table.
[0108] In conjunction with some embodiments of the first aspect, in some embodiments, updating the revenue information of each second decision action in the second Q table to be built based on the reward function of the second decision action of each mobile device, and a pre-set learning rate and discount factor, includes:
[0109] Update the payoff information for each second decision action in the second Q table to be built according to the following formula:
[0110] Among them, Q u,v,2 (s u,v,m ,a u,v,m Q represents the updated payoff information after the second decision action. u,v,2 (s u,v,m′ ,a u,v,m′ () represents the payoff information before the second decision action is updated, r m Let λ represent the reward function for the second decision action. u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the m-th access point (AP) coverage area, and m' represent the m'-th access point (AP) coverage area.
[0111] In the above embodiments, possible implementation methods are provided for updating the benefit information of each second decision action in the second Q table, so as to ensure that the benefit information of each first decision action in the second Q table can be updated, thereby ensuring the timeliness and accuracy of the benefit information recorded in the second Q table.
[0112] In conjunction with some embodiments of the first aspect, in some embodiments, the first state information includes at least one of the following:
[0113] The coverage area of the previous time slot AP;
[0114] Current AP coverage area;
[0115] The initial speed of the mobile device within the current AP coverage area;
[0116] The available bandwidth of the current time-slot network device;
[0117] The available bandwidth of the current time slot relay device.
[0118] In the above embodiments, by providing the information types that the first state information may include, communication resource allocation decisions, speed control decisions, and communication mode decisions can be made based on multiple types of first state information, thereby improving the flexibility of the communication resource allocation decision, speed control decision, and communication mode decision process.
[0119] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0120] A third state information is determined, which is used to indicate a partial state of the network environment in the next time slot and a partial state of the mobile device in the next time slot.
[0121] In the above embodiments, by determining third state information indicating the partial state of the network environment in the next time slot and the partial state of the mobile device in the next time slot, subsequent communication resource allocation, speed control and communication mode decision-making processes can be carried out based on the third state information of each mobile device.
[0122] In conjunction with some embodiments of the first aspect, in some embodiments, the third state information includes at least one of the following:
[0123] Current timeslot AP coverage area;
[0124] The next time slot AP coverage area;
[0125] The initial speed of the mobile device within the AP coverage area in the next time slot;
[0126] Available bandwidth for network devices in the next time slot;
[0127] Available bandwidth for the next time slot relay device.
[0128] In the above embodiments, by providing the types of information that the third state information may include, multiple possible states of the network and mobile devices in the next time slot can be determined after the communication resource allocation decision, speed control decision and communication mode decision are completed in the current time slot, so as to improve the flexibility and diversity of the state prediction process.
[0129] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0130] An edge computing model is constructed, which is used to train a first agent and a second agent. The first agent is used to determine a first decision action based on a first Q-table, and the second agent is used to determine a second decision action based on a second Q-table.
[0131] In the above embodiments, by constructing an edge computing model including two agents, the determination of the first decision action and the second decision action can be realized by the first agent and the second agent respectively, so as to ensure that the determination process of the first decision action and the second decision action will not affect each other, thereby ensuring the accuracy of the decision result.
[0132] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0133] Determine the reward function and probability distribution information of the first target action;
[0134] Based on the reward function and probability distribution information of the first target action, update the reward information and probability distribution information of each first decision action in the first Q table.
[0135] In the above embodiments, after implementing the task unloading decision and computing resource allocation decision, the first Q table is further updated based on the determined first target action to further ensure the implementation and accuracy of the first Q table.
[0136] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0137] Determine the reward function and probability distribution information for the second target action;
[0138] Based on the reward function and probability distribution information of the second target action, update the reward information and probability distribution information of each item in the second Q table.
[0139] In the above embodiments, after implementing communication resource allocation decisions, speed control decisions, and communication mode decisions, the second Q table is further updated based on the determined second target action to further ensure the implementation and accuracy of the second Q table.
[0140] In a second aspect, embodiments of this disclosure provide a communication device, comprising:
[0141] The processing module is configured to acquire first status information, which is used to indicate the current status of the network environment and the current status of the mobile device.
[0142] The processing module is further configured to determine multiple first decision actions, as well as the benefit information and probability distribution information of each first decision action, from a pre-established first Q table based on the first state information; and to determine multiple second decision actions, as well as the benefit information and probability distribution information of each second decision action, from a pre-established second Q table. The first decision actions are used to indicate the action decision results of task unloading and computing resource allocation, and the second decision actions are used to indicate the action decision results of communication resource allocation, speed control, and communication mode. The first Q table and the second Q table are obtained through reinforcement learning.
[0143] The processing module is further configured to determine a first target action from the plurality of first decision actions based on the benefit information and probability distribution information of each first decision action, and to determine a second target action from the plurality of second decision actions based on the benefit information and probability distribution information of each second decision action.
[0144] Thirdly, embodiments of this disclosure provide a communication device, including:
[0145] One or more processors;
[0146] The communication device is used to execute the task offloading and communication management method in mobile edge computing as described in the first aspect and any embodiment of the first aspect.
[0147] Fourthly, embodiments of this disclosure provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the task offloading and communication management method in mobile edge computing as described in the first aspect and any embodiment of the first aspect.
[0148] Fifthly, embodiments of this disclosure provide a communication system comprising: a plurality of mobile devices, relay devices, and network devices; wherein at least one of the mobile devices, the relay devices, and the network devices is configured to perform the task offloading and communication management method in mobile edge computing as described in the first aspect and any embodiment of the first aspect.
[0149] In a sixth aspect, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the task offloading and communication management method in mobile edge computing as described in the first aspect and any embodiment of the first aspect.
[0150] In a seventh aspect, embodiments of this disclosure provide a computer program that, when run on a communication device, causes the communication device to perform the task offloading and communication management method in mobile edge computing as described in the first aspect and any embodiment of the first aspect.
[0151] Eighthly, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the task offloading and communication management methods in mobile edge computing as described in the first aspect and any embodiment of the first aspect.
[0152] It is understood that the aforementioned communication devices, storage media, communication systems, program products, computer programs, chips, or chip systems are all used to execute the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0153] This disclosure provides a method, apparatus, and medium for task offloading and communication management in mobile edge computing. In some embodiments, the terms "task offloading and communication management method in mobile edge computing" can be used interchangeably with "information processing method," "communication method," and "mobile edge computing method," and the terms "task offloading and communication management apparatus in mobile edge computing" can be used interchangeably with "information processing apparatus," "communication apparatus," and "mobile edge computing apparatus," and the terms "information processing system" and "communication system" can be used interchangeably.
[0154] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0155] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0156] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.
[0157] In this embodiment of the disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular expression or a plural expression.
[0158] In the embodiments disclosed herein, "multiple" refers to two or more.
[0159] In some embodiments, the terms “at least one of”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.
[0160] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of B); in some embodiments, B (execute B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). The same applies when there are more branches such as A, B, C, etc.
[0161] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execution of A regardless of B); in some embodiments, B (execution of B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, C, etc.
[0162] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0163] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0164] In some embodiments, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc., can be used interchangeably.
[0165] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.
[0166] In some embodiments, the apparatus and device may be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they may also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "body", etc.
[0167] In some embodiments, "network" can be interpreted as devices included in the network, such as access network devices, core network devices, etc.
[0168] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)," "base station (BS)," "radio base station," or "fixed station." In some embodiments, it may also be understood as "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," or "bandwidth part (BWP)."
[0169] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.
[0170] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.
[0171] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0172] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.
[0173] In traditional cloud computing, computing resources and services are typically concentrated in large data centers, and users can only access these resources and services across enterprises, resulting in low business processing efficiency. Among related technologies, edge computing technology is being heavily considered for development to help improve business message processing efficiency.
[0174] In some embodiments, edge computing technology brings computing services closer to the network edge or to network edge devices, enabling users to receive faster and more reliable services, and allowing businesses to process data and support applications more quickly, significantly reducing business latency.
[0175] In some embodiments, network edge devices can be a type of physical hardware, such as IoT gateways, industrial controllers, smart displays, point-of-sale terminals, vending machines, robots, and drones, but are not limited thereto. These devices can all be located at the network edge and have sufficient memory, processing power, and computing resources to collect, process, and execute data in near real-time with the help of other components in the network.
[0176] In some embodiments, an enterprise can deploy thousands of network edge devices in its network architecture and manage these thousands of network edge devices from a central location to improve the efficiency and speed of business processing.
[0177] Edge computing enables businesses to maintain operations by deploying only essential services and functions in a central location, reducing deployment costs and bandwidth usage. Furthermore, businesses can continue operating and maintain remote resilience even if a network edge device loses connection to the core data center or cloud.
[0178] Referring to Figure 1A, Figure 1A is a schematic diagram of a system architecture of an edge computing network according to an embodiment of the present disclosure. As shown in Figure 1A, the edge computing network may include a mobile device 101, a relay device 102, and a network device 103.
[0179] In some embodiments, the mobile device 101 includes, for example, a terminal, but is not limited thereto.
[0180] In some embodiments, the terminal includes, but is not limited to, at least one of the following: a mobile robot, a mobile phone, a wearable device, an Internet of Things device, a car with communication capabilities, a smart car, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, and a wireless terminal device in a smart home.
[0181] In some embodiments, the relay device 102 includes, for example, a mobile communication repeater, a satellite repeater, a fiber optic repeater, a Wireless Fidelity (Wi-Fi) repeater, a power line communication repeater, etc., but is not limited thereto.
[0182] In some embodiments, mobile communication repeaters include, but are not limited to, drones, mobile base station vehicles, emergency communication boxes, etc.
[0183] In some embodiments, satellite repeaters include, but are not limited to, satellite telephone equipment, mobile satellite communication vehicles, etc.
[0184] In some embodiments, fiber optic repeaters include, for example, fiber optic amplifiers, fiber optic switching devices, etc., but are not limited thereto.
[0185] In some embodiments, a Wi-Fi repeater may include, for example, a mobile Wi-Fi repeater device, but is not limited thereto.
[0186] In some embodiments, the power line communication repeater includes, for example, a power line communication adapter, but is not limited thereto.
[0187] In some embodiments, network device 103 may include, for example, access network device, core network device, etc., but is not limited thereto.
[0188] In some embodiments, the access network device is, for example, a node or device that connects a terminal to a wireless network. The access network device may include, but is not limited to, a satellite in a satellite communication system, an evolved Node B (eNB), a next-generation eNB (ng-eNB), a next-generation Node B (gNB), a node B (NB), a home node B (HNB), a home evolved node B (HeNB) in a 5G communication system, a wireless backhaul device, a radio network controller (RNC), a base station controller (BSC), a base transceiver station (BTS), a base band unit (BBU), a mobile switching center, a base station in a 6G communication system, an open RAN, a cloud RAN, a base station in other communication systems, and an access node in a Wi-Fi system.
[0189] In some embodiments, the technical solutions of this disclosure can be applied to the Open RAN architecture. In this case, the interfaces between or within access network devices involved in the embodiments of this disclosure can be transformed into internal interfaces of Open RAN. The processes and information interactions between these internal interfaces can be implemented by software or programs.
[0190] In some embodiments, the access network device may be composed of a central unit (CU) and a distributed unit (DU). The CU may also be called a control unit. The CU-DU structure can separate the protocol layer of the access network device. Some of the protocol layer functions are centrally controlled by the CU, while the remaining part or all of the protocol layer functions are distributed in the DU and centrally controlled by the CU. However, this is not the only possibility.
[0191] In some embodiments, the core network equipment may be a single device comprising multiple network elements, or it may be multiple devices or a group of devices, each comprising all or part of the multiple network elements. Network elements may be virtual or physical. The core network may include, for example, at least one of the Evolved Packet Core (EPC), 5G Core Network (5GCN), and Next Generation Core (NGC).
[0192] In some embodiments, the core network equipment may include a first network element, such as an Access and Mobility Management Function (AMF).
[0193] In some embodiments, the first network element is used for user access management and mobility management, but is not limited thereto.
[0194] In some embodiments, the core network device may include a second network element, such as a Session Management Function (SMF).
[0195] In some embodiments, the second network element is used for session management of the control plane and user plane, but is not limited thereto.
[0196] In some embodiments, the core network device may include a third network element, such as a User Plane Function (UPF).
[0197] In some embodiments, the third network element is used for user plane data forwarding, traffic statistics, Quality of Service (QoS) management, etc., but is not limited to these.
[0198] In some embodiments, the core network device may include a fourth network element, such as a Policy Control Function (PCF).
[0199] In some embodiments, the fourth network element is used to implement user control policy management, including but not limited to QoS control, service access control, etc.
[0200] In some embodiments, the core network equipment may include a fifth network element, such as a unified data management function (UDM).
[0201] In some embodiments, the fifth network element is used to implement user subscription data management, roaming control, etc., but is not limited to these.
[0202] In some embodiments, the core network device may include a sixth network element, such as an Authentication Server Function (AUSF).
[0203] In some embodiments, the sixth network element is used to implement user authentication, but is not limited thereto.
[0204] In some embodiments, each of the above network elements can be independent of the core network equipment.
[0205] In some embodiments, each of the above network elements may be part of the core network equipment.
[0206] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.
[0207] The following embodiments of this disclosure can be applied to the system architecture or some of the main components shown in FIG1A, but are not limited thereto. The main components shown in FIG1A are illustrative. The edge computing network may include all or some of the main components in FIG1A, or may include other main components outside of FIG1A. The number and form of each main component are arbitrary. Each main component may be physical or virtual. The connection relationship between the main components is illustrative. The main components may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.
[0208] The embodiments disclosed herein can be applied to satellite communication systems, Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other communication methods, and next-generation systems built upon them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).
[0209] Currently, with the continuous development of communication technology and the continuous improvement of economic levels, the construction of cellular network infrastructure has gradually begun to spread to remote, dangerous, or sparsely populated areas. However, the cost of deploying cellular network infrastructure in these areas is relatively high. To reduce the cost of deploying cellular network infrastructure in these areas, satellite networks are gradually becoming an effective solution.
[0210] In some embodiments, connected robots are becoming an important component of satellite networks because they can collaboratively complete complex tasks in remote, dangerous, or sparsely populated areas without human intervention.
[0211] In some embodiments, satellite networks can provide a cost-effective deployment strategy for wide-area coverage and information exchange, enabling robotic teams to autonomously complete tasks in remote, dangerous, or sparsely populated areas. However, the high propagation latency and severe path loss of satellite networks pose challenges to the execution of tasks by robots in remote, dangerous, or sparsely populated areas.
[0212] To address these challenges, hybrid satellite-drone networks have emerged. These networks provide seamless and on-demand connectivity for multiple robots with diverse mission requirements. Furthermore, for latency-sensitive and compute-intensive robotic applications, edge computing servers can be deployed on drones or satellites, enabling robots to send data to the network edge for processing. This allows robotic teams to support compute-intensive and latency-sensitive services through offloading based on Mobile Edge Computing (MEC).
[0213] Referring to Figure 1B, Figure 1B is a schematic diagram of a system architecture of a mobile edge computing network according to an embodiment of the present disclosure. As shown in Figure 1B, the hybrid satellite-UAV network based on mobile edge computing may include satellites, unmanned aerial vehicles (UAVs), gateway stations, and mobile robots.
[0214] In some embodiments, edge computing devices may be deployed on the drone so that the mobile robot can send data to the edge computing devices deployed on the drone for processing, thereby enabling mobile edge computing.
[0215] In some embodiments, an edge computing device may be deployed on the gateway station so that the mobile robot can send data to the edge computing device deployed on the gateway station for processing, thereby enabling mobile edge computing.
[0216] In some embodiments, the edge computing device may be a mobile edge computing server (MEC server), but is not limited thereto.
[0217] In some embodiments, service migration can be performed between satellites and drones.
[0218] In some embodiments, the satellite may send offloading flow to the gateway station, which can then deliver results to the satellite based on the task offloading status.
[0219] In some embodiments, the drone can send unloading traffic to the mobile robot, which can then deliver results to the drone based on the task unloading status.
[0220] In some embodiments, there may be U mobile robots, which can be used to perform V tasks, each of which can be completed collaboratively by the U mobile robots.
[0221] In some embodiments, the mobile robot can move according to business needs to achieve business migration.
[0222] However, in order to complete tasks within a limited time, the rapid collective movement of mobile robots may lead to frequent business migrations, and in satellite and drone communications, a large number of robots clustered together may compete for limited bandwidth resources. Therefore, offloading latency may increase significantly.
[0223] In view of this, this disclosure aims to provide a method for task offloading and communication management in mobile edge computing to maximize the mission success rate of satellite-drone service IoT systems.
[0224] It should be noted that the task offloading and communication management method in mobile edge computing provided in this disclosure embodiment can be executed by a communication device. The communication device can be at least one of a mobile device, a relay device, and a network device, but is not limited thereto. The communication device can also be other devices in the network environment that need to make task offloading decisions and communication management decisions.
[0225] Figure 2 is an interactive schematic diagram of a task offloading and communication management method in mobile edge computing according to an embodiment of the present disclosure. As shown in Figure 2, the embodiments of the present disclosure relate to a task offloading and communication management method in mobile edge computing, the method including:
[0226] Step S2101: Obtain the first state information.
[0227] In some embodiments, the first state information is used to indicate the current state of the network environment and the current state of the mobile device.
[0228] In some embodiments, the first state information used to indicate the current state of the network environment may include, but is not limited to, the coverage area of the current access point (AP), the availability of wireless communication in the current time slot, the available computing resources of all MEC devices in the current time slot, the coverage area of the AP in the previous time slot, the available bandwidth of network devices (such as satellites), the available bandwidth of relay devices (such as drones) in the current time slot, etc.
[0229] In some embodiments, the first state information used to indicate the current state of the mobile device may include, but is not limited to, the data size generated by the mobile device in the current time slot, the size of the calculation result of the mobile device in the current time slot, the available computing resources of the mobile device in the current time slot, the speed of the mobile device in the current time slot, the MEC device selected by the mobile device in the previous time slot, the initial speed of the mobile device in the current AP coverage area, etc.
[0230] In some embodiments, the terms “slot”, “sub-slot”, “mini-slot”, “frame”, “radio frame”, “subframe”, “symbol”, “symbol”, and “transmission time interval (TTI)” can be used interchangeably.
[0231] In some embodiments, the terms “resource,” “resource set,” “resource group,” “precoding,” “precoder,” “weight,” “precoding weight,” “quasi-co-location (QCL),” “transmission configuration indication (TCI) status,” “spatial relation,” “spatial domain filter,” “transmission power,” “phase rotation,” “antenna port,” “antenna port group,” “layer,” “the number of layers,” “rank,” “beam,” “beam width,” “beam angular degree,” “antenna,” “antenna element,” and “panel” can be used interchangeably.
[0232] In some embodiments, the name of the first state information is not limited, and it may be, for example, "first state".
[0233] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0234] Step S2102: Based on the first state information, determine multiple first decision actions, as well as the payoff information and probability distribution information of each first decision action, from the pre-established first Q table; determine multiple second decision actions, as well as the payoff information and probability distribution information of each second decision action, from the pre-established second Q table.
[0235] In some embodiments, the first Q-table and the second Q-table can be obtained through reinforcement learning.
[0236] In some embodiments, an edge computing model can be constructed, and reinforcement learning can be used to obtain the first Q-table and the second Q-table from the edge computing model. The specific training process will be detailed below and will not be repeated here.
[0237] In some embodiments, the edge computing model can be constructed with minimizing the average completion time of task unloading for all mobile devices in the network environment as the optimization objective.
[0238] In some embodiments, the edge computing model can be used to train two agents, namely a first agent and a second agent, wherein the first agent can be used to determine a first decision action based on a first Q-table, and the second agent can be used to determine a second decision action based on a second Q-table.
[0239] In some embodiments, for mobile devices in a network environment, each mobile device may correspond to a set of dual agents, that is, each mobile device may correspond to a first agent and a second agent.
[0240] In some embodiments, the name of the agent is not limited, and it may be, for example, "sub-agent".
[0241] In some embodiments, the name of the first agent is not limited, and may be, for example, "first sub-agent", "unload agent", "unload sub-agent", etc.
[0242] In some embodiments, the name of the second agent is not limited, and it may be, for example, "second sub-agent", "speed control agent", "speed control sub-agent", etc.
[0243] In some embodiments, based on the first state information, a first intelligent agent can determine multiple first decision actions, as well as the reward information and probability distribution information of each first decision action, from a pre-established first Q table.
[0244] It should be noted that the first Q table can maintain possible first decision actions corresponding to different first state information, and the first Q table can maintain the payoff information of the first decision actions under different first state information. In addition, the payoff information corresponding to different first state information can also be associated with probability distribution information.
[0245] In some embodiments, a first intelligent agent can search for multiple first decision actions corresponding to the first state information from a pre-established first Q table, and determine the benefit information and probability distribution information of each first decision action.
[0246] In some embodiments, the benefit information is used to indicate the benefit value of the first decision action under the optimization objective. It should be noted that the benefit information can be inversely proportional to the average task unloading completion time; that is, the shorter the average task unloading completion time, the greater the benefit, and vice versa.
[0247] In some embodiments, the name of the revenue information is not limited, and it may be, for example, "Q value".
[0248] In some embodiments, the first state information used to determine the first decision action may include, but is not limited to, the current AP coverage area, the availability of wireless communication in the current time slot, the data size generated by the mobile device in the current time slot, the size of the calculation result of the mobile device in the current time slot, the available computing resources of the mobile device in the current time slot, the speed of the mobile device in the current time slot, the MEC device selected by the mobile device in the previous time slot, and the available computing resources of all MEC devices in the current time slot.
[0249] In some embodiments, the first decision action is used to indicate the decision result of the action for task unloading and computing resource allocation.
[0250] In some embodiments, the first decision action may include offloading decisions and the computing resources allocated by the MEC server.
[0251] In some embodiments, the name of the first decision action is not limited, and may be, for example, "task unloading decision action", "unloading decision action", "computing resource allocation decision action", etc.
[0252] In some embodiments, based on the first state information, a second agent can determine multiple second decision actions, as well as the payoff information and probability distribution information of each second decision action, from a pre-established second Q table.
[0253] It should be noted that the second Q table can maintain possible second decision actions corresponding to different first state information, and the second Q table can maintain the payoff information of the second decision actions under different first state information. In addition, the payoff information corresponding to different second state information can also be associated with probability distribution information.
[0254] In some embodiments, a second agent can search for multiple second decision actions corresponding to the first state information from a pre-established second Q table, and determine the benefit information and probability distribution information of each second decision action.
[0255] In some embodiments, the benefit information is used to indicate the benefit value of the second decision action under the optimization objective. It should be noted that the benefit information can be inversely proportional to the average task unloading completion time; that is, the shorter the average task unloading completion time, the greater the benefit, and vice versa.
[0256] In some embodiments, the first state information used to determine the second decision action may include, but is not limited to, the coverage area of the previous time slot AP, the coverage area of the current AP, the initial speed of the mobile device in the coverage area of the current AP, the available bandwidth of the current time slot network device, the available bandwidth of the current time slot relay device, etc.
[0257] In some embodiments, the second decision action is used to indicate the action decision results of communication resource allocation, speed control, and communication mode.
[0258] In some embodiments, the second decision action may include the target moving speed of the mobile device, the communication mode, the bandwidth allocation of network devices (such as satellites), and the bandwidth allocation of relay devices (such as drones).
[0259] In some embodiments, the name of the second decision action is not limited, and may be, for example, "speed control decision action", "communication mode decision action", "bandwidth resource allocation decision action", etc.
[0260] Step S2103: Based on the benefit information and probability distribution information of each first decision action, determine the first target action from multiple first decision actions; based on the benefit information and probability distribution information of each second decision action, determine the second target action from multiple second decision actions.
[0261] In some embodiments, the mobile device may determine a first target action from a plurality of first decision actions based on the benefit information and probability distribution information of each first decision action, according to business needs, and determine a second target action from a plurality of second decision actions based on the benefit information and probability distribution information of each second decision action.
[0262] In some embodiments, the business requirement may be to select the decision action that minimizes the average completion time of task unloading, or the business requirement may be to select the decision action that minimizes the average completion time of task unloading to the expected duration and whose probability indicated by the corresponding probability distribution information is greater than the expected probability, and so on, but is not limited thereto.
[0263] In some embodiments, the first target action is used to indicate the action decision results of task unloading and computing resource allocation.
[0264] In some embodiments, the first target action may include offloading decisions and computing resources allocated to the MEC server.
[0265] In some embodiments, the second target action is used to indicate the action decision results of communication resource allocation, speed control, and communication mode.
[0266] In some embodiments, the second target action may include the target moving speed of the mobile device, the communication mode, the bandwidth allocation of network devices (such as satellites), and the bandwidth allocation of relay devices (such as drones).
[0267] Step S2104: Determine the reward function and probability distribution information of the first target action, and determine the reward function and probability distribution information of the second target action.
[0268] In some embodiments, after the first target action and the second target action are determined, the reward function and probability distribution information of the first target action and the reward function and probability distribution information of the second target action can be determined respectively.
[0269] In some embodiments, an unloading reward and a penalty greater than the movement time can be used as the instantaneous reward for the first target action, so that the reward function for achieving the first target action can be determined based on the instantaneous reward.
[0270] In some embodiments, the reward function for the second target action may be the average of the cumulative rewards of all mobile devices, wherein the cumulative reward may be the sum of all instantaneous rewards of each mobile device within its corresponding current AP coverage area.
[0271] Step S2105: Based on the reward function and probability distribution information of the first target action, update the revenue information and probability distribution information of the first decision action corresponding to the first target action in the first Q table; based on the reward function and probability distribution information of the second target action, update the revenue information and probability distribution information of the second target action corresponding to the second target action in the second Q table.
[0272] In some embodiments, the revenue information of the first decision action corresponding to the first target action in the first Q table can be updated based on the reward function of the first target action and the pre-set learning rate and discount factor.
[0273] In some embodiments, the revenue information of the first decision action corresponding to the first target action in the first Q table can be updated according to the following formula (1):
[0274] Among them, Q u,v,1 (s u,v,m,n (t),a u,v,m,n (t) represents the benefit information after the first decision action is updated, corresponding to the first objective action. u,v,1 (s u,v,m,n (t′),a u,v,m,n (t′)) represents the revenue information before the first decision action was updated, r u,v,m,n (t) represents the reward function for the first decision action, λ u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the coverage area of the m-th access point (AP), and n represent the n-th MEC device.
[0275] In some embodiments, the probability distribution information of the first decision action corresponding to the first target action in the first Q table can be updated based on the probability distribution information of the first target action. For example, the probability distribution information of the first decision action corresponding to the first target action in the first Q table can be replaced with the probability distribution information of the first target action.
[0276] In some embodiments, the reward information of the second decision action corresponding to the second objective action in the second Q table can be updated using the reward function of the second objective action and the pre-set learning rate and discount factor.
[0277] In some embodiments, the payoff information of the second decision action corresponding to the second objective action in the second Q table can be updated according to the following formula (2):
[0278] Among them, Q u,v,2 (s u,v,m ,a u,v,m ) represents the payoff information after the second decision action is updated, corresponding to the second objective action; Q,v,2(s,vv,,n′,a,v,m′) represents the payoff information before the second decision action is updated; r m Let λ represent the reward function for the second decision action. u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the m-th access point (AP) coverage area, and m' represent the m'-th access point (AP) coverage area.
[0279] In some embodiments, the probability distribution information of the second decision action corresponding to the second target action in the second Q table can be updated based on the probability distribution information of the second target action. For example, the probability distribution information of the second decision action corresponding to the second target action in the second Q table can be replaced with the probability distribution information of the second target action.
[0280] Step S2106: Determine the second state information and the third state information.
[0281] In some embodiments, after implementing task offloading decisions and computing resource allocation decisions based on first state information, second state information of the mobile device can be determined. The second state information is used to indicate a partial state of the network environment in the next time slot and a partial state of the mobile device in the next time slot.
[0282] In some embodiments, the second state information includes, but is not limited to, the coverage area of the AP in the next time slot, the availability of wireless communication in the next time slot, the data size generated by the mobile device in the next time slot, the size of the calculation result of the mobile device in the next time slot, the available computing resources of the mobile device in the next time slot, the speed of the mobile device in the next time slot, the MEC device selected by the mobile device in the current time slot, and the available computing resources of all MEC devices in the next time slot.
[0283] In some embodiments, after making communication resource allocation decisions, speed control decisions, and communication mode decisions based on first state information, third state information of the mobile device can be determined. The third state information is used to indicate a partial state of the network environment in the next time slot and a partial state of the mobile device in the next time slot.
[0284] In some embodiments, the third state information includes the coverage area of the current time slot AP, the coverage area of the next time slot AP, the initial speed of the mobile device in the AP coverage area of the next time slot, the available bandwidth of the network device in the next time slot, the available bandwidth of the relay device in the next time slot, etc., but is not limited to these.
[0285] In some embodiments, “get,” “obtain,” “receive,” “transmit,” “bidirectional transmission,” and “send and / or receive” can be used interchangeably and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining through self-processing, or autonomous implementation, among other meanings.
[0286] In some embodiments, terms such as “send,” “transmit,” “report,” “distribute,” “transfer,” “bidirectional transmission,” “send and / or receive” can be used interchangeably.
[0287] In some embodiments, terms such as "certain," "preset," "default," "set," "indicated," "a certain," "any," and "first" can be used interchangeably. "Certain A," "preset A," "default A," "set A," "indicated A," "a certain A," "any A," and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.
[0288] In some embodiments, the determination or judgment can be made by a value represented by 1 bit (0 or 1), or by a true or false value (boolean), or by a comparison of numerical values (e.g., a comparison with a predetermined value), but is not limited thereto.
[0289] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or as not performing subsequent processing on the data after receiving it; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the receiver to respond to the sent content.
[0290] The communication method involved in the embodiments of this disclosure may include at least one of steps S2101 to S2106. For example, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, steps S2101+S2102 can be implemented as an independent embodiment, steps S2101+S2103 can be implemented as an independent embodiment, steps S2102+S2103 can be implemented as an independent embodiment, steps S2102+S2106 can be implemented as an independent embodiment, steps S2103+S2104 can be implemented as an independent embodiment, steps S2103+S2105 can be implemented as an independent embodiment, and steps S2103+S2106 can be implemented as an independent embodiment. The embodiments are implemented as follows: steps S2101+S2102+S2103 can be implemented as independent embodiments; steps S2101+S2102+S2106 can be implemented as independent embodiments; steps S2101+S2103+S2104 can be implemented as independent embodiments; steps S2101+S2103+S2105 can be implemented as independent embodiments; steps S2101+S2103+S2106 can be implemented as independent embodiments; steps S2102+S2103+S2104 can be implemented as independent embodiments; steps S2102+S2103+S2105 can be implemented as independent embodiments. As can be implemented as an independent embodiment, steps S2102+S2103+S2106 can be implemented as an independent embodiment, steps S2103+S2104+S2105 can be implemented as an independent embodiment, steps S2103+S2104+S2106 can be implemented as an independent embodiment, steps S2101+S2102+S2103+S2104 can be implemented as an independent embodiment, steps S2101+S2102+S2103+S2105 can be implemented as an independent embodiment, and steps S2101+S2102+S2103+S2106 can be implemented as an independent embodiment. The steps S2102+S2103+S2104+S2105, S2102+S2103+S2104+S2106, S2101+S2102+S2103+S2104+S2105, S2101+S2102+S2103+S2104+S2106, and S21022+S2103+S2104+S2105+S2106 can be implemented as independent embodiments, but are not limited thereto.
[0291] In some embodiments, steps S2105 and S2106 may be performed in an alternate order or simultaneously.
[0292] In some embodiments, steps S2101, S2102, S2104, S2105, and S2106 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0293] In some embodiments, steps S2101, S2103, S2104, S2105, and S2106 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0294] In some embodiments, other optional implementations may be described before or after the specification corresponding to FIG2.
[0295] Figure 3A is a flowchart illustrating a construction process of a first Q-table according to an embodiment of the present disclosure. As shown in Figure 3A, the embodiment of the present disclosure relates to a construction process of a first Q-table, the process including:
[0296] Step S3101: Based on the first state information of multiple mobile devices, determine the first decision action for each mobile device.
[0297] In some embodiments, the first state information is used to indicate the current state of the network environment and the current state of the mobile device.
[0298] In some embodiments, the first state information includes, but is not limited to, the current AP coverage area, the availability of wireless communication in the current time slot, the data size generated by the mobile device in the current time slot, the size of the calculation result of the mobile device in the current time slot, the available computing resources of the mobile device in the current time slot, the speed of the mobile device in the current time slot, the MEC device selected by the mobile device in the previous time slot, and the available computing resources of all MEC devices in the current time slot.
[0299] In some embodiments, the first decision action of each mobile device can be determined by a greedy algorithm based on the first state information of multiple mobile devices.
[0300] In some embodiments, a first decision action for each mobile device can be determined using a greedy algorithm based on first state information of multiple mobile devices, with each first decision action corresponding to a probability distribution information; communication resources to be allocated to the multiple mobile devices can be determined based on the first decision actions of the multiple mobile devices; and whether to adopt the first decision action determined by the greedy algorithm can be determined based on the communication resources to be allocated to the multiple mobile devices and the first communication resources of the network, with the first communication resources used to indicate the available communication resources of the mobile edge computing (MEC) device in the network environment.
[0301] In some embodiments, when determining whether to adopt a determined first decision action based on communication resources to be allocated to multiple mobile devices and the first communication resources of the network, if the sum of communication resources to be allocated to multiple mobile devices is less than or equal to the first communication resources, the first decision action determined by the greedy algorithm can be adopted; otherwise, if the sum of communication resources to be allocated to multiple mobile devices is greater than the first communication resources, the reward function of the first decision action of each mobile device can be set to a first value, and the first decision action of each mobile device can be re-determined by the greedy algorithm.
[0302] Step S3102: Determine the reward function and probability distribution information for the first decision action of each mobile device.
[0303] In some embodiments, the reward function for the first decision action of each first mobile device can be determined using an uninstallation reward and a penalty greater than the time spent moving as instantaneous rewards.
[0304] Step S3103: Based on the reward function and probability distribution information of the first decision action of each mobile device, update the revenue information and probability distribution information of each first decision action in the first Q table to be established.
[0305] In some embodiments, the revenue information of each first decision action in the first Q table to be built can be updated based on the reward function of the first decision action of each mobile device, as well as the pre-set learning rate and discount factor.
[0306] In some embodiments, the revenue information of each first decision action in the first Q table to be established can be updated according to the following formula (3):
[0307] Among them, Q u,v,1 (s u,v,m,n (t),a u,v,m,n (t) represents the payoff information after the first decision action, Q. u,v,1 (s u,v,m,n (t′),a u,v,m,n (t′)) represents the payoff information before the first decision action is updated, r u,v,m,n (t) represents the reward function for the first decision action, λ u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the coverage area of the m-th access point (AP), and n represent the n-th MEC device.
[0308] In some embodiments, the probability distribution information of each first decision action in the first Q table to be established can be updated based on the probability distribution information of the first decision action of each mobile device.
[0309] Step S3104: Determine the second state information.
[0310] In some embodiments, the second state information is used to indicate a portion of the state of the network environment in the next time slot and a portion of the state of the mobile device in the next time slot.
[0311] In some embodiments, the second state information includes, but is not limited to, the coverage area of the AP in the next time slot, the availability of wireless communication in the next time slot, the data size generated by the mobile device in the next time slot, the size of the calculation result of the mobile device in the next time slot, the available computing resources of the mobile device in the next time slot, the speed of the mobile device in the next time slot, the MEC device selected by the mobile device in the current time slot, and the available computing resources of all MEC devices in the next time slot.
[0312] The communication method involved in the embodiments of this disclosure may include at least one of steps S3101 to S3104. For example, step S3103 may be implemented as an independent embodiment, step S3101+S3103 may be implemented as an independent embodiment, step S3102+S3103 may be implemented as an independent embodiment, step S3103+S3104 may be implemented as an independent embodiment, step S3101+S3102+S3103 may be implemented as an independent embodiment, step S3101+S3103+S3104 may be implemented as an independent embodiment, and step S3102+S3103+S3104 may be implemented as an independent embodiment, but is not limited thereto.
[0313] In some embodiments, steps S3103 and S3104 may be performed in an alternate order or simultaneously.
[0314] In some embodiments, steps S3101, S3102, and S3104 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0315] Figure 3B is a flowchart illustrating a construction process of a second Q-table according to an embodiment of the present disclosure. As shown in Figure 3B, the embodiment of the present disclosure relates to a construction process of a second Q-table, the process including:
[0316] In step S3201, the AP connected to any of the multiple mobile devices changes. Based on the first state information of the multiple mobile devices, a second decision action is determined for each mobile device.
[0317] In some embodiments, the first state information is used to indicate the current state of the network environment and the current state of the mobile device.
[0318] In some embodiments, the first state information includes, but is not limited to, the coverage area of the previous time slot AP, the coverage area of the current AP, the initial speed of the mobile device in the coverage area of the current AP, the available bandwidth of the current time slot network device (such as a satellite), the available bandwidth of the current time slot relay device (such as a drone), etc.
[0319] In some embodiments, if the AP connected to any of the multiple mobile devices changes, a second decision action for each mobile device can be determined by a greedy algorithm based on the first state information of the multiple mobile devices.
[0320] In some embodiments, a second decision action for each mobile device can be determined using a greedy algorithm based on first state information of multiple mobile devices, with each second decision action corresponding to a probability distribution information; based on the second decision actions of multiple mobile devices, a communication bandwidth to be allocated to multiple mobile devices is determined; based on the communication bandwidth to be allocated to multiple mobile devices and the first communication bandwidth of the network, it is determined whether to adopt the second decision action determined by the greedy algorithm, whereby the communication bandwidth is used to indicate the available communication bandwidth of relay devices and / or MEC devices in the network environment.
[0321] In some embodiments, when determining whether to adopt a second decision action determined by a greedy algorithm based on the communication bandwidth to be allocated to multiple mobile devices and the first communication bandwidth of the network, if the combined result of the communication bandwidth to be allocated to multiple mobile devices does not exceed the first communication bandwidth, the second decision action determined by the greedy algorithm can be adopted; otherwise, if the combined result of the communication bandwidth to be allocated to multiple mobile devices exceeds the first communication bandwidth, the reward function of the second decision action of each mobile device can be set to a first value, and the second decision action of each mobile device can be re-determined by the greedy algorithm.
[0322] Step S3202: Determine the reward function and probability distribution information for the second decision action of each mobile device.
[0323] In some embodiments, the average of the cumulative rewards of all mobile devices can be used as the reward function for the second decision action, wherein the cumulative reward is defined as the sum of all instantaneous rewards of the mobile device within its current AP coverage area.
[0324] Step S3203: Based on the reward function and probability distribution information of the second decision action of each mobile device, update the reward information and probability distribution information of each second decision action in the second Q table to be established.
[0325] In some embodiments, the revenue information of each second decision action in the second Q table to be built can be updated based on the reward function of the second decision action of each mobile device, as well as the pre-set learning rate and discount factor.
[0326] In some embodiments, the payoff information for each second decision action in the second Q table to be established can be updated according to the following formula (4):
[0327] Among them, Q u,v,2 (s u,v,m a u,v,m Q represents the updated payoff information after the second decision action. u,v,2 (s u,v,m′ a u,v,m′ () represents the payoff information before the second decision action is updated, r m Let λ represent the reward function for the second decision action. u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the m-th access point (AP) coverage area, and m' represent the m'-th access point (AP) coverage area.
[0328] In some embodiments, the probability distribution information of each first decision action in the second Q table to be established can be updated based on the probability distribution information of the second decision action of each mobile device.
[0329] Step S3204: Determine the third state information.
[0330] In some embodiments, the third state information is used to indicate a portion of the state of the network environment in the next time slot and a portion of the state of the mobile device in the next time slot.
[0331] In some embodiments, the third state information includes, but is not limited to, the coverage area of the current time slot AP, the coverage area of the next time slot AP, the initial speed of the mobile device in the coverage area of the next time slot AP, the available bandwidth of the next time slot network device (such as a satellite), and the available bandwidth of the next time slot relay device (such as a drone).
[0332] The communication method involved in the embodiments of this disclosure may include at least one of steps S3201 to S3204. For example, step S3203 may be implemented as an independent embodiment, step S3201+S3203 may be implemented as an independent embodiment, step S3202+S3203 may be implemented as an independent embodiment, step S3203+S3204 may be implemented as an independent embodiment, step S3201+S3202+S3203 may be implemented as an independent embodiment, step S3201+S3203+S3204 may be implemented as an independent embodiment, and step S3202+S3203+S3204 may be implemented as an independent embodiment, but is not limited thereto.
[0333] In some embodiments, steps S3203 and S3204 may be performed in an alternate order or simultaneously.
[0334] In some embodiments, steps S3201, S3202, and S3204 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0335] In the embodiments disclosed herein, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations in other embodiments.
[0336] According to the solution provided in this disclosure, an optimization problem is proposed for multi-robot (i.e., mobile device) task offloading, speed control, and resource allocation to minimize the average completion time of offloading. The solution provided in this disclosure can transform an optimization problem with long-term constraints into an optimization problem with short-term constraints, and can decompose the overall robot optimization problem into an optimization problem of the robot within the coverage area of a single access point (AP).
[0337] This disclosure proposes a multi-agent Q-learning algorithm that utilizes multiple sets of dual-agent Q-learning to solve the proposed problem, while taking into account observed wireless communication availability, reduced computational resource states, and global rewards. In the dual-agent Q-learning framework, one agent is responsible for offloading decision-making and computational resource allocation, while the other agent is responsible for speed control, communication mode determination, and communication resource allocation.
[0338] In some embodiments, each pair of dual agents corresponds to one robot. One agent can perform time-slot-based unloading decisions and computing resource allocation, while the other agent can perform speed control, communication mode determination, and communication resource allocation based on the AP coverage area.
[0339] Alternatively, according to the solution provided in the embodiments of this disclosure, referring to Figure 4, which is a schematic diagram of the principle of a multi-agent Q-learning algorithm according to the embodiments of this disclosure, as shown in Figure 4, two different sub-agents can be provided for each network environment, namely, the unloaded sub-agent Agent. u,v,1 and speed control sub-agent u,v,2 To uninstall the sub-agent u,v,1 and speed control sub-agent u,v,2 Each robot makes decisions on different actions.
[0340] The states, rewards, and behaviors of these two sub-agents can be defined based on Markov Decision Process (MDP).
[0341] In some embodiments, for uninstalling the sub-agent u,v,1 Its status can include: the current AP coverage area, the availability of wireless communication, the size of the generated data, the size of the calculation result, the available computing resources of the mobile robot, the speed of the mobile robot, the MEC server selected in the previous slot, and the available computing resources of all MEC servers.
[0342] In some embodiments, for uninstalling the sub-agent u,v,1 Its actions can include: offloading decisions and allocating computing resources to the MEC server.
[0343] In some embodiments, for uninstalling the sub-agent u,v,1 Its instantaneous rewards can include: offloading rewards and penalties exceeding the movement time.
[0344] In some embodiments, when a mobile robot switches between adjacent AP coverage areas, the speed control sub-agent... u,v, 2. Speed control strategy, communication method strategy and bandwidth resource allocation strategy are obtained through Q-learning.
[0345] In some embodiments, for the speed control sub-agent u,v,2 Its status can include: the coverage area of the previous time slot AP, the coverage area of the current AP, the initial speed in the current AP coverage area, and the available bandwidth for satellite communication and drone communication.
[0346] In some embodiments, for the speed control sub-agent u,v,2 Its actions can include: target speed, communication method, and bandwidth allocation for satellite and drone communications.
[0347] In some embodiments, for the speed control sub-agent u,v,2 The reward can be the average of the cumulative rewards of all robots, where the cumulative reward is defined as the sum of all instantaneous rewards of the mobile robot within its current AP coverage area.
[0348] In some embodiments, uninstalling the sub-agent Agent u,v,1 and speed control sub-agent u,v,2 The Q-learning process can be implemented using the following algorithm:
[0349] Algorithm 1: Joint offloading and speed control based on multiple sets of dual-agent Q-learning algorithms:
[0350] For Algorithm 1, its input can include the reward value Q. u,v,1 (s,a) = 0 and the reward value Q u,v,2 (s,a)=0, velocity v u,v,m (t), distance traveled c m Available satellite communication bandwidth B s,m Available UAV communication bandwidth B c,m Available MEC computing resources F n (t), local computing resources f local,u,v (t), learning rate λ u,v Greed factor ε u,v Discount factor γ u,v Its output can include unloading decision α u,v,n (t) Calculate resource allocation f u,v,n (t), target velocity v * u,v,m Communication mode β u,v,m Bandwidth allocation W u,v,m .
[0351] In some embodiments, Algorithm 1 can be implemented using the following code:
[0352] Algorithm 2 is used to implement behavioral decisions for bandwidth resource allocation. The output of Algorithm 2 may include communication mode β. u,v,m and bandwidth allocation W u,v,m .
[0353] In some embodiments, Algorithm 2 can be implemented using the following code:
[0354] Algorithm 3 is used to implement behavioral decisions regarding unloading and computational resource allocation. The output of Algorithm 3 may include the unloading decision α. u,v,n (t) and computational resource allocation f u,v,n (t).
[0355] In some embodiments, Algorithm 3 can be implemented using the following code:
[0356] This disclosure also proposes an apparatus for implementing any of the above methods. For example, an apparatus is proposed that includes units or modules for implementing the steps performed by the communication device in any of the above methods.
[0357] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0358] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).
[0359] Figure 5 is a schematic diagram of the structure of the communication device proposed in an embodiment of this disclosure. As shown in Figure 5, the communication device 5100 may include at least a processing module 5101. In some embodiments, the processing module 5101 is configured to acquire first state information, which indicates the current state of the network environment and the current state of the mobile device; the processing module 5101 is further configured to, based on the first state information, determine a plurality of first decision actions from a pre-established first Q table, and the benefit information and probability distribution information of each first decision action, and determine a plurality of second decision actions from a pre-established second Q table, and the benefit information and probability distribution information of each second decision action, wherein the first decision actions indicate the action decision results of task unloading and computing resource allocation, and the second decision actions indicate the action decision results of communication resource allocation, speed control, and communication mode, wherein the first Q table and the second Q table are obtained through reinforcement learning; the processing module 5101 is further configured to, based on the benefit information and probability distribution information of each first decision action, determine a first target action from the plurality of first decision actions, and based on the benefit information and probability distribution information of each second decision action, determine a second target action from the plurality of second decision actions. Optionally, the processing module 5101 is used to perform at least one of the other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, but not limited thereto) performed by the communication device in any of the above methods, which will not be elaborated here. In some embodiments, the communication device 5100 may further include a transceiver module. Optionally, the transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the communication device in any of the above methods, which will not be elaborated here.
[0360] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, which may be separate or integrated. Optionally, the transceiver module may be interchangeable with a transceiver.
[0361] In some embodiments, the processing module may be a single module or may include multiple sub-modules. Optionally, the multiple sub-modules may each perform all or part of the steps required by the processing module. Optionally, the processing module may be interchangeable with a processor.
[0362] Figure 6A is a schematic diagram of the structure of the communication device 6100 proposed in an embodiment of this disclosure. The communication device 6100 can be a mobile device, a relay device, a network device, or a chip, chip system, or processor that supports the implementation of any of the above methods in a mobile device, a relay device, or a network device. The communication device 6100 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0363] As shown in Figure 6A, the communication device 6100 includes one or more processors 6101. The processor 6101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process program data. The communication device 6100 is used to execute any of the above methods.
[0364] In some embodiments, the communication device 6100 further includes one or more memories 6102 for storing instructions. Optionally, all or part of the memories 6102 may also be located outside the communication device 6100.
[0365] In some embodiments, the communication device 6100 further includes one or more transceivers 6103. When the communication device 6100 includes one or more transceivers 6103, the transceivers 6103 perform at least one of the communication steps such as sending and / or receiving in the above method, and the processor 6101 performs at least one of other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, but not limited thereto).
[0366] In some embodiments, a transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, etc., may be used interchangeably; the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc., may be used interchangeably; and the terms receiver, receiving unit, receiver, receiving circuit, etc., may be used interchangeably.
[0367] In some embodiments, the communication device 6100 may include one or more interface circuits 6104. Optionally, the interface circuit 6104 is connected to the memory 6102, and the interface circuit 6104 can be used to receive signals from the memory 6102 or other devices, and can be used to send signals to the memory 6102 or other devices. For example, the interface circuit 6104 can read instructions stored in the memory 6102 and send the instructions to the processor 6101.
[0368] The communication device 6100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 6100 described in this disclosure is not limited thereto, and the structure of the communication device 6100 may not be limited by FIG. 6A. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0369] Figure 6B is a schematic diagram of the structure of chip 6200 according to an embodiment of this disclosure. For cases where the communication device 6100 can be a chip or a chip system, please refer to the schematic diagram of chip 6200 shown in Figure 6B, but it is not limited thereto.
[0370] Chip 6200 includes one or more processors 6201, which are used to perform any of the above methods.
[0371] In some embodiments, chip 6200 further includes one or more interface circuits 6202. Optionally, the interface circuit 6202 is connected to memory 6203, and the interface circuit 6202 can be used to receive signals from memory 6203 or other devices, and the interface circuit 6202 can be used to send signals to memory 6203 or other devices. For example, the interface circuit 6202 can read instructions stored in memory 6203 and send the instructions to processor 6201.
[0372] In some embodiments, the interface circuit 6202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 6201 performs at least one of other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, but not limited thereto).
[0373] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0374] In some embodiments, chip 6200 further includes one or more memories 6203 for storing instructions. Optionally, all or part of the memories 6203 may be located outside of chip 6200.
[0375] This disclosure also proposes a storage medium storing instructions that, when executed on the communication device 6100, cause the communication device 6100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.
[0376] This disclosure also provides a program product that, when executed by the communication device 6100, causes the communication device 6100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0377] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.
[0378] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0379] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for task offloading and communication management in mobile edge computing, characterized in that, The method includes: Obtain first status information, which is used to indicate the current status of the network environment and the current status of the mobile device; Based on the first state information, multiple first decision actions, along with the benefit information and probability distribution information of each first decision action, are determined from a pre-established first Q table. Multiple second decision actions, along with the benefit information and probability distribution information of each second decision action, are determined from a pre-established second Q table. The first decision actions are used to indicate the action decision results of task unloading and computational resource allocation, and the second decision actions are used to indicate the action decision results of communication resource allocation, speed control, and communication mode. The first Q table and the second Q table are obtained through reinforcement learning. Based on the benefit information and probability distribution information of each first decision action, a first target action is determined from the plurality of first decision actions; based on the benefit information and probability distribution information of each second decision action, a second target action is determined from the plurality of second decision actions.
2. The method according to claim 1, characterized in that, The first Q table construction process includes: Based on the first state information of multiple mobile devices, determine the first decision action for each mobile device; Determine the reward function and probability distribution information for the first decision action of each mobile device; Based on the reward function and probability distribution information of the first decision action of each mobile device, update the reward information and probability distribution information of each first decision action in the first Q table to be built.
3. The method according to claim 2, characterized in that, The determination of the first decision action for each mobile device based on the first state information of multiple mobile devices includes: Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the first decision action of each mobile device.
4. The method according to claim 3, characterized in that, The step of determining the first decision action for each mobile device based on the first state information of the multiple mobile devices using a greedy algorithm includes: Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the first decision action of each mobile device, and each first decision action corresponds to a probability distribution information. Based on the first decision action of the plurality of mobile devices, the communication resources to be allocated to the plurality of mobile devices are determined; Based on the communication resources to be allocated to the plurality of mobile devices and the first communication resources of the network, it is determined whether to adopt the first decision action determined by the greedy algorithm, wherein the first communication resources are used to indicate the available communication resources of the mobile edge computing (MEC) device in the network environment.
5. The method according to claim 4, characterized in that, The determination of whether to adopt the determined first decision action based on the communication resources to be allocated to the plurality of mobile devices and the first communication resources of the network includes any one of the following: If the sum of the communication resources to be allocated to the plurality of mobile devices is less than or equal to the first communication resource, the first decision action determined by the greedy algorithm is adopted. If the total amount of communication resources to be allocated to the plurality of mobile devices is greater than the first communication resource, the reward function of the first decision action of each mobile device is set to the first value, and the first decision action of each mobile device is redetermined through a greedy algorithm.
6. The method according to claim 2, characterized in that, The reward function and probability distribution information of the first decision action based on each mobile device are used to update the payoff information and probability distribution information of each first decision action in the first Q table to be established, including: Based on the reward function of the first decision action of each mobile device, as well as the pre-set learning rate and discount factor, update the revenue information of each first decision action in the first Q table to be built; Based on the probability distribution information of the first decision action of each mobile device, update the probability distribution information of each first decision action in the first Q table to be built.
7. The method according to claim 6, characterized in that, The reward function based on the first decision action of each mobile device, along with the pre-set learning rate and discount factor, updates the revenue information of each first decision action in the first Q table to be built, including: Update the payoff information for each first decision action in the first Q table to be built according to the following formula: Among them, Q u,v,1 (s u,v,m,n (t), a u,v,m,n (t) represents the payoff information after the first decision action, Q. u,v,1 (s u,v,m,n (t′), a u,v,m,n (t′)) represents the payoff information before the first decision action is updated, r u,v,m,n (t) represents the first The reward function for the decision action, λ u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the coverage area of the m-th access point (AP), and n represent the n-th MEC device.
8. The method according to any one of claims 2 to 7, characterized in that, The first status information includes at least one of the following: Current AP coverage area; The availability of current time-slot wireless communication; The current data size generated by the mobile device in the current time slot; The size of the calculation result for the current time slot mobile device; The available computing resources of the current time-slot mobile device; The speed of the mobile device in the current time slot; The MEC device selected by the mobile device in the previous time slot; Available computing resources for all MEC devices in the current time slot.
9. The method according to any one of claims 2 to 8, characterized in that, The method further includes: Determine second state information, which is used to indicate a partial state of the network environment in the next time slot and a partial state of the mobile device in the next time slot.
10. The method according to claim 9, characterized in that, The second status information includes at least one of the following: The next time slot AP coverage area; Availability of next time-slot wireless communication; The size of the data generated by the mobile device in the next time slot; The size of the calculation result for the next time slot mobile device; Available computing resources for the next time slot mobile device; The speed of mobile devices in the next time slot; The MEC device selected by the mobile device in the current time slot; Available computing resources for all MEC devices in the next time slot.
11. The method according to any one of claims 1 to 10, characterized in that, The second Q table construction process includes: If the AP connected to any of the multiple mobile devices changes, a second decision action is determined for each mobile device based on the first state information of the multiple mobile devices. Determine the payoff information and probability distribution information for the second decision action of each mobile device; Based on the payoff information and probability distribution information of the second decision action for each mobile device, update the payoff information and probability distribution information of each second decision action in the second Q table to be built.
12. The method according to claim 11, characterized in that, The step of determining a second decision action for each mobile device based on the first state information of the plurality of mobile devices includes: Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the second decision action of each mobile device.
13. The method according to claim 12, characterized in that, The step of determining the second decision action for each mobile device based on the first state information of the plurality of mobile devices using a greedy algorithm includes: Based on the first state information of the multiple mobile devices, a greedy algorithm is used to determine the second decision action of each mobile device, and each second decision action corresponds to a probability distribution information. Based on the second decision action of the plurality of mobile devices, the communication bandwidth to be allocated to the plurality of mobile devices is determined; Based on the communication bandwidth to be allocated to the plurality of mobile devices and the first communication bandwidth of the network, it is determined whether to adopt the second decision action determined by the greedy algorithm, wherein the communication bandwidth is used to indicate the available communication bandwidth of relay devices and / or MEC devices in the network environment.
14. The method according to claim 13, characterized in that, The step of determining whether to adopt the second decision action determined by the greedy algorithm based on the communication bandwidth to be allocated to the plurality of mobile devices and the first communication bandwidth of the network includes any one of the following: If the combined result of the communication bandwidth to be allocated to the multiple mobile devices does not exceed the first communication bandwidth, then the second decision action determined by the greedy algorithm is adopted. If the combined result of the communication bandwidth to be allocated to the multiple mobile devices exceeds the first communication bandwidth, the reward function of the second decision action of each mobile device is set to the first value, and the second decision action of each mobile device is redetermined through a greedy algorithm.
15. The method according to claim 11, characterized in that, The method of updating the payoff information and probability distribution information of each second decision action in the second Q table to be established based on the payoff information and probability distribution information of each second decision action based on the payoff information and probability distribution information of each second decision action based on the second decision action of each mobile device includes: Based on the reward function of the second decision action for each mobile device, as well as the pre-set learning rate and discount factor, update the revenue information of each second decision action in the second Q table to be built; Based on the probability distribution information of the second decision action of each mobile device, update the probability distribution information of each first decision action in the second Q table to be established.
16. The method according to claim 15, characterized in that, The reward function for each second decision action based on the mobile device, along with the pre-set learning rate and discount factor, updates the revenue information for each second decision action in the second Q table to be built, including: Update the payoff information for each second decision action in the second Q table to be built according to the following formula: Among them, Q u,v,2 (s u,v,m ,a u,v,m Q represents the updated payoff information after the second decision action. u,v,2 (s u,v,m′ a u,v,m′ () represents the payoff information before the second decision action is updated, r m Let λ represent the reward function for the second decision action. u,v γ represents the learning rate. u,v Let u represent the discount factor, v represent the v-th mobile device, m represent the m-th access point (AP) coverage area, and m' represent the m'-th access point (AP) coverage area.
17. The method according to any one of claims 11 to 16, characterized in that, The first status information includes at least one of the following: The coverage area of the previous time slot AP; Current AP coverage area; The initial speed of the mobile device within the current AP coverage area; The available bandwidth of the current time-slot network device; The available bandwidth of the current time slot relay device.
18. The method according to any one of claims 11 to 17, characterized in that, The method further includes: A third state information is determined, which is used to indicate a partial state of the network environment in the next time slot and a partial state of the mobile device in the next time slot.
19. The method according to claim 18, characterized in that, The third state information includes at least one of the following: Current timeslot AP coverage area; The next time slot AP coverage area; The initial speed of the mobile device within the AP coverage area in the next time slot; Available bandwidth for network devices in the next time slot; Available bandwidth for the next time slot relay device.
20. The method according to any one of claims 1 to 19, characterized in that, The method further includes: An edge computing model is constructed, which is used to train a first agent and a second agent. The first agent is used to determine a first decision action based on a first Q-table, and the second agent is used to determine a second decision action based on a second Q-table.
21. The method according to any one of claims 1 to 20, characterized in that, The method further includes: Determine the reward function and probability distribution information of the first target action; Based on the reward function and probability distribution information of the first target action, update the reward information and probability distribution information of each first decision action in the first Q table.
22. The method according to any one of claims 1 to 21, characterized in that, The method further includes: Determine the reward function and probability distribution information for the second target action; Based on the reward function and probability distribution information of the second target action, update the reward information and probability distribution information of each item in the second Q table.
23. A mobile device, characterized in that, include: The processing module is configured to acquire first status information, which is used to indicate the current status of the network environment and the current status of the mobile device. The processing module is further configured to determine multiple first decision actions, as well as the benefit information and probability distribution information of each first decision action, from a pre-established first Q table based on the first state information; and to determine multiple second decision actions, as well as the benefit information and probability distribution information of each second decision action, from a pre-established second Q table. The first decision actions are used to indicate the action decision results of task unloading and computing resource allocation, and the second decision actions are used to indicate the action decision results of communication resource allocation, speed control, and communication mode. The first Q table and the second Q table are obtained through reinforcement learning. The processing module is further configured to determine a first target action from the plurality of first decision actions based on the benefit information and probability distribution information of each first decision action, and to determine a second target action from the plurality of second decision actions based on the benefit information and probability distribution information of each second decision action.
24. A communication device, characterized in that, include: One or more processors; The communication device is used to execute the task offloading and communication management method in mobile edge computing according to any one of claims 1-22.
25. A storage medium storing instructions, characterized in that, When the instruction is executed on the communication device, the communication device performs the task offloading and communication management method in mobile edge computing as described in any one of claims 1-22.
Citation Information
Patent Citations
Low-complexity mobile edge computing resource allocation method
CN114281527A
Fog computing task unloading method based on fuzzy logic strategy
CN114637552A
Calculation and communication resource collaborative allocation method in machine network
CN117202262A
Computation offloading method and communication apparatus
US20230081937A1
Cited By
Distributed task unloading and service caching joint optimization method and device
CN122054237A