A method, apparatus, medium, and equipment for training task offloading in an edge computing system.

By introducing an exploration and collaborative offloading method into the edge computing system, and utilizing maximum entropy and mutual information to update the parameters of the offloader and commenter, the problem of low task offloading efficiency in the prior art is solved, and efficient task offloading and system operation are achieved.

CN119669758BActive Publication Date: 2025-12-02HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411750089.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-12-02
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

In existing technologies, mobile devices in edge computing systems only receive guidance from the policy commentator network by offloading costs, ignoring the exploratory and collaborative nature between systems, resulting in low task offloading efficiency.

Method used

An exploration and collaborative unloading method was designed. By introducing exploration and collaboration metrics with maximum entropy into the unloading decision of mobile devices, and combining mutual information as a reward, the parameters of the unloader and commenter are updated to improve the efficiency of task unloading.

Benefits of technology

The improved offloading strategy enhances the task offloading efficiency and system operation accuracy of the edge computing system, ensuring an efficient task offloading process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669758B_ABST
    Figure CN119669758B_ABST
Patent Text Reader

Abstract

This invention proposes a task offloading training method, apparatus, medium, and device for edge computing systems. In the nth time slice of the ka-th training round, tasks are sent to each node device, enabling each node device to complete task offloading under the control of its internal offloader, thus gaining task offloading experience. Based on the task offloading experience of the node devices in the nth time slice, and combining the commentator loss function and the offloader queue function, the offloader parameters within each node device and the corresponding commentator parameters are updated. This task offloading training method for edge computing systems ensures the accuracy of the offloader parameters within the node devices and the corresponding commentator parameters, facilitating efficient task offloading in the edge computing system and guaranteeing the overall system efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of edge computing, and more specifically, to a method, apparatus, medium, and device for training task offloading in an edge computing system. Background Technology

[0002] In recent years, multi-access edge computing (MEC) systems have provided mobile devices with satisfactory computing resource services and lower task latency by offloading tasks to nearby edge servers. Utilizing multi-agent reinforcement learning (MARL) to design decentralized offloading methods is a popular trend, primarily based on a centralized training and decentralized execution model. However, current MD systems only accept guidance from a policy critic network (also known as a commentator) based on offloading costs for task offloading, neglecting the exploratory and collaborative aspects across the entire edge computing system. Summary of the Invention

[0003] The purpose of this invention is to provide a method, apparatus, medium, and device for training task offloading in an edge computing system, so as to improve the above-mentioned problems.

[0004] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:

[0005] In a first aspect, embodiments of the present invention provide a task offloading training method for an edge computing system, the method comprising:

[0006] In the nth time slice of the kath training round, tasks are sent to each node device so that each node device can complete task unloading under the control of its internal unloader to gain task unloading experience.

[0007] Based on the task unloading experience of the node device in the nth time slice, and combining the commentator loss function and the unloader echelon function, the unloader parameters inside each node device and the commentator parameters corresponding to each node device are updated.

[0008] In a second aspect, embodiments of the present invention provide an edge computing system task offloading training apparatus, the apparatus comprising:

[0009] The first processing unit is used to send tasks to each node device in the nth time slice of the kath training round, so that each node device can complete the task unloading under the control of its internal unloader and obtain task unloading experience.

[0010] The second processing unit is used to update the unloader parameters inside each node device and the commentator parameters corresponding to each node device based on the task unloading experience of the node device in the nth time slice, combined with the commentator loss function and the unloader echelon function.

[0011] Thirdly, embodiments of the present invention provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.

[0012] Fourthly, embodiments of the present invention provide an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store one or more programs; when the one or more programs are executed by the processor, the above-described method is implemented.

[0013] Compared to existing technologies, the task offloading training method, apparatus, medium, and device for edge computing systems provided in this invention send tasks to each node device in the nth time slice of the ka-th training round, enabling each node device to complete task offloading under the control of its internal offloader, thereby gaining task offloading experience. Based on the task offloading experience of the node devices in the nth time slice, and combining the commentator loss function and the offloader echelon function, the offloader parameters inside each node device and the commentator parameters corresponding to each node device are updated. This task offloading training method for edge computing systems ensures the accuracy of the offloader parameters inside the node devices and the commentator parameters corresponding to the node devices, which is beneficial for efficient task offloading in the edge computing system and ensures the overall system operating efficiency.

[0014] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0017] Figure 2 This is one of the flowcharts illustrating the task offloading training method for an edge computing system provided in an embodiment of the present invention.

[0018] Figure 3This is the second flowchart illustrating the task offloading training method for an edge computing system provided in an embodiment of the present invention.

[0019] Figure 4 This is the third flowchart illustrating the task offloading training method for an edge computing system provided in an embodiment of the present invention.

[0020] Figure 5 This is a schematic diagram of a task offloading training device for an edge computing system provided in an embodiment of the present invention.

[0021] In the diagram: 10-Processor; 11-Memory; 12-Bus; 13-Communication interface; 501-First processing unit; 502-Second processing unit. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0023] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0024] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0025] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0026] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0027] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0028] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0029] This invention provides a task offloading training method for edge computing systems, proposing an Explorative and Collaborative Offloading (ExplabOff) method. By consciously utilizing the exploratory and collaborative information implicit in the current state of mobile devices (MDs) and their offloading decisions, it achieves stronger task offloading benefits. Specifically, two additional offloading strategy learning metrics are designed: an exploratory metric based on the maximum entropy of joint offloading actions by MDs, and a collaborative metric based on the degree of influence of one MD on the offloading decisions of other MDs. These two metrics are then combined into a new evaluation criterion, defined as the mutual information (MI) between MD states and actions, and used as an additional reward during intensive training, besides the reward obtained from reducing offloading costs. Furthermore, MI is differentiated into superior and inferior offloading, with distinct enhancements and weakenings applied. Therefore, the task offloading training method for edge computing systems provided by this invention exhibits better performance.

[0030] This invention provides an electronic device, which can be a server device or a computer device, serving as a training terminal in an edge computing system. Please refer to... Figure 1 This is a schematic diagram of the structure of an electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected via the bus 12. The processor 10 is used to execute executable modules, such as computer programs, stored in the memory 11.

[0031] Processor 10 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the edge computing system task offloading training method can be completed through integrated logic circuits in the hardware or software instructions within processor 10. Processor 10 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0032] The memory 11 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0033] Bus 12 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. Figure 1 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus 12 or one type of bus 12.

[0034] The memory 11 is used to store programs, such as programs corresponding to an edge computing system task offloading training device. The edge computing system task offloading training device includes at least one software functional module that can be stored in the memory 11 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device. Upon receiving an execution instruction, the processor 10 executes the program to implement the edge computing system task offloading training method.

[0035] The electronic device provided in this embodiment of the invention may further include a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.

[0036] It should be understood that, Figure 1 The structure shown is only a partial schematic diagram of the electronic device; the electronic device may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0037] The task offloading training method for an edge computing system provided in this embodiment of the invention can be applied to, but is not limited to, [various applications]. Figure 1 For the specific process of the electronic devices shown, please refer to [link / reference]. Figure 2 The training methods for task offloading in edge computing systems include S110 and S120, as detailed below.

[0038] S110, in the nth time slice of the ka-th training round, sends tasks to each node device so that each node device can complete task unloading under the control of its internal unloader, thereby gaining task unloading experience.

[0039] It should be noted that the edge computing system includes M node devices and E edge servers. In the edge computing system, each node device communicates and connects to E edge servers, enabling task offloading. Node devices can be, but are not limited to, mobile devices, such as wristbands or mobile phones.

[0040] The task sent to the m-th node device in the n-th time slice can be represented as:

[0041]

[0042] in, This represents the task sent from the nth time slice to the mth node device, where c represents the data processing speed, the number of CPU cycles required to process each bit of data, and tmax represents the preset maximum tolerable latency.

[0043] S120, based on the task offloading experience of the node device in the nth time slice, and combined with the commentator loss function and the offloader echelon function, updates the offloader parameters inside each node device and the commentator parameters corresponding to each node device.

[0044] Optionally, the task unloading experience is {sn, an, rn, s′n+1}, where sn represents the system state data of the edge computing system in the nth time slice, an represents the task unloading data of the edge computing system in the nth time slice, rn represents the reward value of the edge computing system in the nth time slice, and s′n+1 represents the estimated system state data of the edge computing system in the nth time slice. The experience pool... The experience of storing task unloading is {sn,an,rn,s′n+1}.

[0045]

[0046] Where Lθ(Qm) represents the commenter loss function, Let Qm(si,ai) represent the unloading echelon function, where Qm(si,ai) represents the value estimate of the unloading state-action pair (si,ai) for the i-th time slice of the m-th node device drawn from the experience pool β using a neural estimator. To represent the expected value of the extracted samples using a symbolic notation. Indicates from the experience pool Extract the task unloading experience from the i-th time slice. Indicates from the experience pool The unloading state action pair for the i-th time slice is extracted. This represents the unloader function for the m-th node device. Find the gradient with respect to the parameter φ. This indicates that the device at the known m-th node has adopted... Under the condition of the unloading action, the evaluator corresponding to the m-th node device. Find the gradient of the action am of the m-th node device, where γ represents the pre-set reward discount factor;

[0047] This indicates that the unloading method used for all M known node devices was... Under the condition of the unloading action, the target evaluator corresponding to the m-th node device. The value estimate is given by θ′, which represents the parameters of the target evaluator, φ′, which represents the parameters of the target unloader of the corresponding node device, r(s,a)=rn=r(sn,an), which represents the reward value of the edge computing system in the nth time slice, INCE(s;a) represents the estimate of the lower limit of mutual information obtained based on the first neural estimator, IL1Out(s;a) represents the estimate of the upper limit of mutual information obtained based on the second neural estimator, μ represents the weight of the first neural estimator set in advance, ν represents the weight of the second neural estimator set in advance, si represents the system state data of the edge computing system in the i-th time slice of the ka-th training round, ai represents the task unloading data of the edge computing system in the i-th time slice of the ka-th training round, and si′+1 represents the system state prediction data of the edge computing system in the i-th time slice of the ka-th training round.

[0048] It should be noted that the parameters θ′ of the target evaluator and φ′ of the target unloader of the node device can also be updated according to the second period interval. Both the target evaluator and the target unloader are deployed within the electronic device. The update method can be to determine the unloader parameters inside the node device after the current time slice update as the parameters of the node device's target unloader, and to determine the evaluator parameters corresponding to the node device after the current time slice update as the parameters of the node device's target evaluator.

[0049] Please refer to Figure 3 In an optional implementation, the edge computing system task offloading training method further includes S130 and S140, which are described in detail below.

[0050] S130 determines whether the commenter loss function and the unloader echelon function have converged. If both have converged, training ends; if either has not converged, then S140 is executed, setting n = n + 1.

[0051] S140, let n = n + 1.

[0052] After S140, S110 is executed again, and in the nth time slice of the ka-th training round, the task is sent to each node device.

[0053] It should be noted that after setting n = n + 1, we can also check whether n > N is true. If it is true, it means that the training round ka has ended. When the training round ka ends, set ka = ka + 1, set n = 1, and repeat S110. In the nth time slice of the training round ka, send the task to each node device.

[0054] Please refer to Figure 4 In an optional implementation, during the last time slice of the ka-th training round, the edge computing system task offloading training method further includes: S210, S220, S230, S240, and S250, which are described in detail below.

[0055] S210, obtain the comprehensive reward for the kath training round.

[0056] The formula for calculating the overall reward return is:

[0057] rΓ=∑γn-1[r(s,a)+μINCE(s;a)-νIL1Out(s;a)]

[0058] Where rΓ represents the overall reward return.

[0059] S220: Determine whether the overall reward return is greater than the reward threshold. If yes, proceed to S230; otherwise, proceed to S240.

[0060] S230, add the unloaded state action pair (s,a) of training round ka to the first buffer.

[0061] The unloading state action pair (s,a) of the ka-th training round includes the system state data and task unloading data of each node device in each time slice of the ka-th training round.

[0062] S240, add the unloaded state action pair (s,a) of training round ka to the second buffer.

[0063] S250, update the parameters of the neural estimator according to the first period interval, including: updating the parameters of the first neural estimator according to the samples in the first buffer, and updating the parameters of the second neural estimator according to the samples in the second buffer.

[0064] Optionally, the formulas for the first neural estimator and the second neural estimator are:

[0065] INCE(s; a) = log(K) - LNCE

[0066]

[0067]

[0068] I(s;a)≤IL1Out(s;a), I(s;a)≥INCE(s;a)

[0069] Where K represents the expected value. The number of samples, where exp represents the exponential function. Indicates the first buffer Calculate the expected value of the unloading state action pair extracted in the i-th time slice i with respect to the expression within the square brackets. Indicates the first buffer Calculate the mathematical expectation of the subsequent formula for the unloading state action of the j-th time slice extracted from the data. This represents a fractional function based on parameters ψN, where ψN represents the parameters of the first neural estimator. Indicates the second buffer Calculate the expected value of the unloading state action pair extracted in the i-th time slice with respect to the expression within the square brackets. This represents a fractional function based on parameter ψL1, where ψL1 represents the parameters of the second neural estimator. Indicates the second buffer Calculate the mathematical expectation of the subsequent formula for the unloading state action of the j-th time slice extracted from the data. Let i represent a fractional function based on parameter ψL1, where i is not equal to j.

[0070] Optionally, This represents the amount of task data for the m-th node device in the nth time slice. This represents the length of the task queue of the e-th edge server in the nth time slice. This represents the available energy of the m-th node device in the n-th time slice. ε represents the number of node devices in the edge computing system task, and ε represents the number of edge servers in the edge computing system task.

[0071]

[0072] in, This represents the unloader for the m-th node device in the n-th time slice. The input of the unloader is the state. This represents the task unloading data of the m-th node device in the n-th time slice. This represents the task offloading ratio of the m-th node device in the n-th time slice. This represents the task unloading object of the m-th node device in the nth time slice.

[0073]

[0074]

[0075] Where rn represents the reward value of the edge computing system in the nth time slice, A and B are set positive real numbers, and Ω(sn,an) is the penalty term. tmax represents the task latency of the m-th node device in the n-th time slice, and tmax represents the preset maximum tolerable latency.

[0076]

[0077] in, Let represent the average system cost over the first N time slices of training round ka, and cn represent the average system cost over the nth time slice of training round ka. M represents the execution cost of the m-th node device in the n-th time slice. η represents the number of node devices in an edge computing system task, and η represents the time cost weight. This represents the task delay of the m-th node device in the n-th time slice. This represents the energy cost of the m-th node device in the n-th time slice.

[0078]

[0079] in, This represents the local task latency of the m-th node device in the n-th time slice. This represents the edge computing latency of the m-th node device in the nth time slice. This represents the local computing power consumption of the m-th node device in the n-th time slice. This represents the edge computing power consumption of the m-th node device in the nth time slice;

[0080]

[0081] in, This represents the time when the m-th node device in the n-th time slice transmits data to its task offloading object. This represents the edge computing latency of the m-th node device in the nth time slice. This represents the start time of the first execution of the task corresponding to the m-th node device in the n-th time slice. This represents the edge computing execution time of the m-th node device in the nth time slice. Indicates satisfaction The indicator function, ptran, represents the transmit power of the node device. This represents the rate at which the m-th node device in the n-th time slice transmits data to its task offloading object. This represents the task offloading ratio of the m-th node device in the n-th time slice. c represents the amount of task data of the m-th node device in the n-th time slice, and c represents the data processing speed. This represents the computing power of the m-th node device. ξ represents the computing power of the second edge server, and ξ represents the energy efficiency coefficient of the m-th node device; This represents the set of node devices that offload tasks to the same edge server before the m-th node device in the nth time slice.

[0082] Optionally, Where B′ represents the system bandwidth, σ 2 Indicates Gaussian noise. This represents the link channel interference between the m-th node device in the n-th time slice and the e-th edge server. This represents the channel gain between the m-th node device in the n-th time slice and the e-th edge server.

[0083] Optionally, in, This represents the channel gain between the m′-th node device in the nth time slice and the e-th edge server.

[0084] Optionally,

[0085] in, and These represent the offloading links from the m-th node device to the e-th edge server in the n-th time slice. Small-scale fading and large-scale fading.

[0086] in, and It is an independent and identically distributed circularly symmetric complex Gaussian random variable with unit variance, and its correlation coefficient is expressed as ψ=J0(2πfdT), where J0(·) is the zero-order Bessel function of the first kind, and fd is the maximum Doppler frequency.

[0087] in, It is an offload link The Euclidean distance is given by lgz, which is a log-normal variable.

[0088] Please see Figure 5 , Figure 5 An edge computing system task offloading training device is provided as an embodiment of the present invention. Optionally, the edge computing system task offloading training device is applied to the electronic device described above.

[0089] The edge computing system task offloading training device includes: a first processing unit 501 and a second processing unit 502.

[0090] The first processing unit 501 is used to send tasks to each node device in the nth time slice of the kath training round, so that each node device can complete task unloading under the control of its internal unloader to obtain task unloading experience.

[0091] The second processing unit 502 is used to update the unloader parameters inside each node device and the commentator parameters corresponding to each node device based on the task unloading experience of the node device in the nth time slice, combined with the commentator loss function and the unloader echelon function.

[0092] Optionally, the second processing unit 502 may execute S120 as described above, and the first processing unit 501 may execute the other steps described above.

[0093] It should be noted that the edge computing system task offloading training device provided in this embodiment can execute the method flow shown in the above method flow embodiment to achieve the corresponding technical effects. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments.

[0094] This invention also provides a storage medium storing computer instructions and programs, which, when read and executed, perform the edge computing system task offloading training method described above. The storage medium may include memory, flash memory, registers, or a combination thereof.

[0095] The following provides an electronic device, which can be a server device or a computer device, serving as a training terminal in an edge computing system. This electronic device, as follows... Figure 1 As shown, the edge computing system task offloading training method described above can be implemented. Specifically, the electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs, and when one or more programs are executed by the processor 10, the edge computing system task offloading training method of the above embodiment is executed.

[0096] In summary, the task offloading training method, apparatus, medium, and device for an edge computing system provided by this invention sends tasks to each node device in the nth time slice of the ka-th training round, enabling each node device to complete task offloading under the control of its internal offloader, thereby gaining task offloading experience. Based on the task offloading experience of the node devices in the nth time slice, and combining the commentator loss function and the offloader echelon function, the offloader parameters inside each node device and the commentator parameters corresponding to each node device are updated. The task offloading training method for an edge computing system provided by this invention can ensure the accuracy of the offloader parameters inside the node devices and the commentator parameters corresponding to the node devices, which is beneficial for efficient task offloading in the edge computing system and ensures the overall system operating efficiency.

[0097] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0098] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A task offloading training method for an edge computing system, characterized in that, The method includes: In the nth time slice of the kath training round, tasks are sent to each node device so that each node device can complete task unloading under the control of its internal unloader to gain task unloading experience. Based on the task unloading experience of the node device in the nth time slice, and combining the commentator loss function and the unloader echelon function, update the unloader parameters inside each node device and the commentator parameters corresponding to each node device. The experience of uninstalling the task is as follows ,in, This represents the system state data of the edge computing system in the nth time slice. This indicates that the edge computing system is unloading data from its tasks in the nth time slice. This represents the reward value of the edge computing system in the nth time slice. The experience pool represents the system state prediction data of the edge computing system in the nth time slice. Used to store the experience of unloading the task. ; ; in, This represents the commenter loss function. This represents the unloader queue functions. This indicates the use of a neural estimator on the experience pool. The unloading status action pair of the m-th node device in the i-th time slice is extracted ( The estimated value of ) To represent the expected value of the extracted samples using a symbolic notation. Indicates from the experience pool Extract the task unloading experience from the i-th time slice. Indicates from the experience pool The unloading state action pair for the i-th time slice is extracted. This represents the unloader function for the m-th node device. Please provide information about the parameters. For the gradient of the variable, This indicates that the device at the known m-th node has adopted... Under the condition of the unloading action, the evaluator corresponding to the m-th node device. Find the actions of the m-th node device. For the gradient of the variable, This represents the pre-set reward discount factor; This indicates that the unloading method used for all M known node devices was... Under the condition of the unloading action, the target evaluator corresponding to the m-th node device. The estimated value, This represents the parameters of the target evaluator. This indicates the parameters of the target unloader for the corresponding node device. , This represents the estimated value of the mutual information lower bound obtained based on the first neural estimator. This represents the estimated upper bound of mutual information obtained based on the second neural estimator. This represents the pre-defined weight of the first neural estimator. This indicates the pre-defined weight of the second neural estimator. This represents the system state data of the edge computing system in the i-th time slice of the ka-th training round. This indicates that the edge computing system unloads data from the task in the i-th time slice of the ka-th training round. This represents the system state prediction data of the edge computing system in the i-th time slice of the ka-th training round.

2. The edge computing system task offloading training method as described in claim 1, characterized in that, The method further includes: Determine whether the commenter loss function and the unloader echelon function converge; Training ends when both convergences are achieved. In any case of non-convergence, let n=n+1, and repeat the task to each node device in the nth time slice of the kath training round.

3. The edge computing system task offloading training method as described in claim 1, characterized in that, In the last time slice of the ka-th training round, the method further includes: Obtain the overall reward for training round ka; Determine whether the overall reward return is greater than the reward threshold; If so, then the unloading state action of training round ka will be... Add to the first buffer; If not, then the unloading state action of training round ka will be... Add to the second buffer; The parameters of the neural estimator are updated according to the first period interval, including: updating the parameters of the first neural estimator according to the samples in the first buffer, and updating the parameters of the second neural estimator according to the samples in the second buffer.

4. The edge computing system task offloading training method as described in claim 3, characterized in that, The formula for calculating the overall reward return is: in, This indicates a comprehensive reward and return.

5. The edge computing system task offloading training method as described in claim 3, characterized in that, The formulas for the first neural estimator and the second neural estimator are as follows: Where K represents the expected value. The number of samples, where exp represents the exponential function. Indicates the first buffer The i-th time slice extracted from The unloading state action is related to the mathematical expectation of the expression within the square brackets. Indicates the first buffer Calculate the mathematical expectation of the subsequent formula for the unloading state action of the j-th time slice extracted from the data. Indicates based on parameters fractional function, This represents the parameters of the first neural estimator. Indicates the second buffer Calculate the expected value of the unloading state action pair extracted in the i-th time slice with respect to the expression within the square brackets. Indicates based on parameters fractional function, This represents the parameters of the second neural estimator. Indicates the second buffer Calculate the mathematical expectation of the subsequent formula for the unloading state action of the j-th time slice extracted from the data. Indicates based on parameters The fractional function, where i is not equal to j.

6. The edge computing system task offloading training method as described in claim 1, characterized in that, , This represents the amount of task data for the m-th node device in the n-th time slice. This represents the length of the task queue of the e-th edge server in the nth time slice. This represents the available energy of the m-th node device in the n-th time slice. This indicates the number of node devices in an edge computing system task. This indicates the number of edge servers in the edge computing system task. ; ; in, This represents the unloader for the m-th node device in the nth time slice. The input of the unloader is the state. , This represents the task unloading data of the m-th node device in the nth time slice. This represents the task offloading ratio of the m-th node device in the n-th time slice. This represents the task offloading object of the m-th node device in the nth time slice; ; ; ; in, Let A and B represent the reward value of the edge computing system in the nth time slice, where A and B are set positive real numbers. As a penalty item, This represents the task delay of the m-th node device in the n-th time slice. This indicates the preset maximum tolerable delay. This represents the average system cost over the first N time slices in the ka-th training round.

7. A task offloading training device for an edge computing system, characterized in that, The device includes: The first processing unit is used to send tasks to each node device in the nth time slice of the kath training round, so that each node device can complete the task unloading under the control of its internal unloader and obtain task unloading experience. The second processing unit is used to update the unloader parameters inside each node device and the commentator parameters corresponding to each node device based on the task unloading experience of the node device in the nth time slice, combined with the commentator loss function and the unloader echelon function. The experience of uninstalling the task is as follows ,in, This represents the system state data of the edge computing system in the nth time slice. This indicates that the edge computing system is unloading data from its tasks in the nth time slice. This represents the reward value of the edge computing system in the nth time slice. The experience pool represents the system state prediction data of the edge computing system in the nth time slice. Used to store the experience of unloading the task. ; ; in, This represents the commenter loss function. This represents the unloader queue functions. This indicates the use of a neural estimator on the experience pool. The unloading status action pair of the m-th node device in the i-th time slice is extracted ( The estimated value of ) To represent the expected value of the extracted samples using a symbolic notation. Indicates from the experience pool Extract the task unloading experience from the i-th time slice. Indicates from the experience pool The unloading state action pair for the i-th time slice is extracted. This represents the unloader function for the m-th node device. Please provide information about the parameters. For the gradient of the variable, This indicates that the device at the known m-th node has adopted... Under the condition of the unloading action, the evaluator corresponding to the m-th node device. Find the actions of the m-th node device. For the gradient of the variable, This represents the pre-set reward discount factor; This indicates that the unloading method used for all M known node devices was... Under the condition of the unloading action, the target evaluator corresponding to the m-th node device. The estimated value, This represents the parameters of the target evaluator. This indicates the parameters of the target unloader for the corresponding node device. , This represents the estimated value of the mutual information lower bound obtained based on the first neural estimator. This represents the estimated upper bound of mutual information obtained based on the second neural estimator. This represents the pre-defined weight of the first neural estimator. This indicates the pre-defined weight of the second neural estimator. This represents the system state data of the edge computing system in the i-th time slice of the ka-th training round. This indicates that the edge computing system unloads data from the task in the i-th time slice of the ka-th training round. This represents the system state prediction data of the edge computing system in the i-th time slice of the ka-th training round.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

9. An electronic device, characterized in that, include: Processor and memory, the memory being used to store one or more programs; When the one or more programs are executed by the processor, the method as described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Task unloading method for multi-view reasoning application in edge computing environment

    CN116166336A

  • End-side cloud collaborative unloading method based on deep reinforcement learning

    CN118283708A