A Resource Optimization Method and System Based on Mobile Edge Computing

By combining radio maps and deep reinforcement learning in smart factories to optimize AGV trajectory design, the complex MIP problem of AGV mobile edge computing system in smart factories is solved, and the system's computing efficiency and resource utilization are improved.

CN116192635BActive Publication Date: 2025-07-25NANCHANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310124613.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-16
Publication Date
2025-07-25
Estimated Expiration
2043-02-16

AI Technical Summary

Technical Problem

The prior art In the AGV-based mobile edge computing system in smart factories, there are complex MIP problems and high computing complexity, and the deep reinforcement learning method converges slowly in high-dimensional action space, and it is impossible to effectively handle high-concurrency and large-scale data communication needs.

Method used

AGV trajectory design is combined with radio maps, and resource allocation is optimized using deep reinforcement learning, unloading decisions are generated through neural network models, and tasks are offloaded to the mobile edge computing server equipped with AGV, optimizing resource allocation for user unloading time and charging time.

Benefits of technology

It effectively reduces the computational complexity, improves the overall computing rate of the system, is suitable for continuous state space, solves the problems of dimensional curse and slow convergence speed, and achieves near-optimal performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116192635B_ABST
    Figure CN116192635B_ABST
Patent Text Reader

Abstract

The present application discloses a resource optimization method and system based on mobile edge computing. The resource optimization method based on mobile edge computing specifically includes the following steps: simulating the radio map of an intelligent factory; setting the initial trajectory of an automatic guided vehicle and the positions of each user, and generating the channel state information corresponding to the initial trajectory; optimizing the movement trajectory of the automatic guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automatic guided vehicle, and generating the channel state information corresponding to the optimized stack; building a neural network model and using the channel state information to generate offloading decisions; performing resource allocation to obtain the best offloading action; updating the offloading strategy according to the obtained best offloading action to obtain the optimal offloading strategy. The present application completes the trajectory optimization of the AGV, and at the same time can design the trajectory of the AGV in combination with the radio map in the intelligent factory, effectively improving the overall computing rate of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a resource optimization method and system based on mobile edge computing. Background Art

[0002] Continuous breakthroughs in science and technology have brought rapid changes to people's lifestyles, and the intelligent scenarios they are in have also become increasingly complex, such as vehicle networking, AR / VR, smart homes, wireless communications, etc. As a large number of tasks are connected to the network, traditional storage methods and computing methods can no longer meet the computing and processing of massive data and the increasing demands of users for communication. Currently, industrial digitization has become an important part of the digital economy. With the advancement of industrial digitalization and intelligent transformation, a large number of high-concurrency and large-scale data communication and processing services have also emerged in intelligent factories. This article focuses on the mobile edge computing (MEC) system based on AGV (Automated Guided Vehicle) in intelligent factories, and uses mobile edge computing technology to offload users' computationally intensive tasks to the mobile edge computing server carried by the AGV, so as to alleviate the overloaded tasks in the entire system. Regarding the offloading decision and resource allocation problems in edge computing, many related works jointly model the computing mode decision problem and resource allocation problem in the MEC network as a mixed integer programming (MIP) problem. For example, the coordinate descent (CD) method can be used to perform a search and optimization for only one variable dimension each time. A heuristic search method can also be used to iteratively adjust the binary offloading decision. Another widely used heuristic method is through convex relaxation. For example, the 0-1 integer variable is relaxed to a continuous variable between 0 and 1, or quadratic constraints are used to approximate binary constraints. However, the above methods all have certain defects. First, the low-complexity heuristic algorithms cannot guarantee the optimality of the solution. Second, the methods based on search and convex relaxation both require a large number of iterations to achieve a relatively good local optimal effect, and it is not applicable to fast-fading channels.

[0003] The method of this patent is inspired by deep reinforcement learning in dealing with reinforcement learning problems in large state spaces and action spaces. This method learns from training data samples through a deep neural network (DNN) and finally generates an optimal mapping from the state space to the action space. Currently, there is limited work on mobile edge computing (MEC) network offloading based on deep reinforcement learning. There is a study on an online computing offloading strategy under random task arrivals based on Deep Q-Network (DQN), which uses discrete channel gains as the input state vector. When high channel quantization accuracy is required, problems such as the curse of dimensionality and slow convergence speed exist. In addition, due to the exhaustive search nature of DQN when selecting actions in each iteration, it is not suitable for dealing with problems in high-dimensional action spaces.

[0004] Based on this, therefore, how to provide a resource optimization method that can completely solve complex Mixed Integer Programming (MIP) problems and reduce computational complexity is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] This application proposes a resource optimization method for a mobile edge computing system based on deep reinforcement learning in the scenario of Automated Guided Vehicle (AGV) transportation in an intelligent factory. This solution learns from past offloading experiences under various wireless fading conditions and automatically improves its action generation strategy. Specifically, it combines a radio map to design the trajectory of an Automated Guided Vehicle (AGV) and uses the method of deep reinforcement learning to optimize resource allocation. Specifically, for the AGV transportation scenario in an intelligent factory, it comprehensively considers channel quality, AGV movement distance, and user location to optimize the design of the AGV trajectory, and uses deep reinforcement learning to generate offloading decisions and optimize resource allocation for user offloading time and charging time. Therefore, it completely eliminates the need to solve complex MIP problems, and the computational complexity does not explode as the network scale increases. In addition, a trajectory optimization strategy for AGVs is proposed to improve the overall effective computing rate of the system.

[0006] This application proposes a resource optimization method based on mobile edge computing, specifically including the following steps: simulating the radio map of an intelligent factory; in response to completing the simulation of the radio map of the intelligent factory, setting the initial trajectory of the automated guided vehicle and the positions of each user, and generating the channel state information corresponding to the initial trajectory; optimizing the movement trajectory of the automated guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automated guided vehicle, and generating the channel state information corresponding to the optimized stack; building a neural network model, and the neural network model uses the channel state information to generate offloading decisions; in response to generating offloading decisions, performing resource allocation to obtain the best offloading action; updating the offloading strategy according to the obtained best offloading action to obtain the optimal offloading strategy.

[0007] As described above, among them, the offloading decisions generated by the neural network model using the channel state information include the offloading decisions generated according to the channel state information corresponding to the initial trajectory and the offloading decisions generated according to the channel state information corresponding to the optimized trajectory; the best offloading actions include the best offloading actions corresponding to the initial trajectory obtained after executing the offloading decisions generated according to the channel state information corresponding to the initial trajectory, and the best offloading actions corresponding to the optimized trajectory obtained after executing the offloading decisions generated according to the channel state information corresponding to the optimized trajectory; the optimal offloading strategy includes the optimal offloading strategy corresponding to the initial trajectory obtained after executing the best offloading actions corresponding to the initial trajectory, and the optimal offloading strategy corresponding to the optimized trajectory obtained after executing the best offloading actions corresponding to the optimized trajectory.

[0008] As described above, among them, the system computing rate corresponding to the initial trajectory obtained according to the optimal offloading strategy corresponding to the initial trajectory and the system computing rate corresponding to the initial trajectory obtained according to the optimal offloading strategy corresponding to the optimized trajectory are used to compare the system computing rates obtained before and after trajectory optimization. If the difference in computing rates is less than the specified threshold, then output is performed.

[0009] As described above, among them, the simulation of the radio map of the intelligent factory includes setting D ∈ R 2 which is the region of the positions x of all possible AGVs in the intelligent factory, considering dividing D into K non - overlapping segments D = D1(X U ) ∪ D2(X U ) ∪....D K (X U ), where D k (X u ) represents the AGV position region where the AGV maintains a k - degree LOS obstacle for the user. The proposed piece - wise propagation model of G dB is specified as:

[0010]

[0011] d u (X) is the distance between the AGV and the user, the user position is (X U , H), the AGV position is (X, 0), X, X U ∈R 2 respectively represent the horizontal positions of the AGV and the user; thus (X U , H), (X, 0) are respectively the positions of the user and the AGV in the three - dimensional image; a k and b k are some parameters, Ⅱ{A} is an indicator function that takes the value 1 if the condition A is satisfied and 0 otherwise; the random variable ε kCapture the remaining shadow effects. Assume a k , b k and ε k whose statistical data are completely known, then the radio map of the entire intelligent factory is simulated.

[0012] As above, where the setting of the initial trajectory of the automated guided vehicle and the positions of each user includes, assuming there are N users, the horizontal positions G of the N users are respectively G = {G1, G2, G3... G N ), where G N represents the horizontal position of the Nth user, and the heights of the N users are uniformly set to H.

[0013] A resource optimization system based on mobile edge computing specifically includes: a simulation unit, an initial setting unit, an optimization unit, an offloading decision generation unit, a best offloading action acquisition unit, an optimal offloading strategy acquisition unit, and an output unit; the simulation unit is used to simulate the radio map of the intelligent factory; the initial setting unit is used to set the initial trajectory of the automated guided vehicle and the positions of each user; the optimization unit is used to optimize the movement trajectory of the automated guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automated guided vehicle; the offloading decision generation unit 240 is used to build a neural network model, and the neural network model generates offloading decisions using channel state information; the best offloading action acquisition unit is used to perform resource allocation and obtain the best offloading action; the optimal offloading strategy acquisition unit is used to update the offloading strategy according to the obtained best offloading action and obtain the optimal offloading strategy.

[0014] As above, where the offloading decisions generated by the neural network model in the offloading decision generation unit using channel state information include the offloading decisions generated according to the channel state information corresponding to the initial trajectory and the offloading decisions generated according to the channel state information corresponding to the optimized trajectory; the best offloading actions obtained by the best offloading action acquisition unit include the best offloading actions corresponding to the initial trajectory obtained after executing the offloading decisions generated according to the channel state information corresponding to the initial trajectory, and the best offloading actions corresponding to the optimized trajectory obtained after executing the offloading decisions generated according to the channel state information corresponding to the optimized trajectory; the optimal offloading strategies obtained by the optimal offloading strategy acquisition unit include the optimal offloading strategies corresponding to the initial trajectory obtained after executing the best offloading actions corresponding to the initial trajectory, and the optimal offloading strategies corresponding to the optimized trajectory obtained after executing the best offloading actions corresponding to the optimized trajectory.

[0015] As described above, in the optimal offloading strategy acquisition unit, the system computing rate corresponding to the initial trajectory obtained according to the optimal offloading strategy corresponding to the initial trajectory, and the system computing rate corresponding to the initial trajectory obtained according to the optimal offloading strategy corresponding to the optimized trajectory are compared. If the difference in the computing rate is less than the specified threshold, the result is output.

[0016] As described above, in which the simulation unit performs the simulation of the radio map of the intelligent factory, including setting D ∈ R 2 is the area of all possible positions x of the AGVs in the intelligent factory. Considering dividing D into K non-overlapping segments D = D1(X U ) ∪ D2(X U ) ∪.... D K (X U ), where D k (X u ) represents the AGV position area where the AGV maintains a k-degree LOS obstacle to the user. The proposed piecewise propagation model of G dB is specified as:

[0017]

[0018] d u (X) is the distance between the AGV and the user, the user position is (X U , H), the AGV position is (X, 0), X, X U ∈R 2 respectively represent the horizontal positions of the AGV and the user; thus (X U , H), (X, 0) are the positions of the user and the AGV in the three-dimensional image respectively; a k and b k are some parameters, Ⅱ{A} is an indicator function, which takes the value 1 if the condition A is satisfied, otherwise it takes the value 0; the random variable ε k captures the remaining shadow effect. Assuming that the statistical data of a k , b k and ε k are completely known, the radio map of the entire intelligent factory is simulated.

[0019] As described above, in which the initial setting unit sets the initial trajectory of the automatic guided vehicle and the positions of each user, including assuming that there are N users, and the horizontal positions G of the N users are respectively G = {G1, G2, G3... G N ), where G N represents the horizontal position of the Nth user, and the heights of the N users are uniformly set to H.

[0020] This application has the following beneficial effects:

[0021] (1) This application provides an idea of simulating the radio map of the entire intelligent factory, which facilitates the subsequent trajectory optimization of AGV and the construction of the MEC network.

[0022] (2) This application provides a solution for trajectory design of AGV by combining the radio map in the intelligent factory, which can effectively improve the overall computing rate of the system.

[0023] (3) This application provides an idea of decomposing the original optimization problem into two sub-problems: namely, the offloading decision sub-problem and the resource allocation sub-problem. Deep reinforcement learning is used to generate offloading decisions, and then convex optimization is used to optimize the charging time and offloading time resources of users in the intelligent factory. It is applicable to continuous state spaces and does not require the discretization of channel gains, making up for the problems of the previous curse of dimensionality and slow convergence speed, and maximizing the overall weighted sum computing rate of the MEC system in the intelligent factory. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0025] Figure 1 is a flowchart of a resource optimization method based on mobile edge computing according to an embodiment of the present application;

[0026] Figure 2 is an internal structure diagram of a resource optimization system based on mobile edge computing according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0028] This application designs an optimization study on trajectory design of AGV by combining radio maps in an intelligent factory and resource allocation using deep reinforcement learning. Specifically, the user needs to decide whether to perform local computing or offload tasks to the mobile edge computing server carried by the AGV based on the channel quality between their own location and the location of the AGV. During the movement of the AGV, the collected computing tasks need to be processed by the mobile edge computing server and then fed back to the user. In this paper, we consider a binary offloading strategy, that is, tasks are either locally computed at the user side or offloaded to the edge computing server carried by the AGV for computing. Our goal is to generate the optimal offloading decision from the state space to the action space through deep reinforcement learning, and maximize the system weighted sum computing rate by allocating resources for the user's own charging time and offloading time, and then improve the overall weighted sum computing rate of the system through AGV trajectory optimization. Considering that the AGV has a stable power supply, each user has a rechargeable battery that can store the energy obtained from charging to power the operation of the device. We assume that the computing speed and transmission power of the edge computing server carried by the AGV are much larger than those of the users limited by size and energy. Therefore, in this embodiment, the time spent by the edge computing server carried by the AGV in task computing and downloading is ignored, and the offloading computing rate is equal to the user data offloading capacity. In this way, each time frame is only occupied by the user's charging time and task offloading time. To maximize the computing rate, the user consumes all the energy it collects in both offloading and local computing modes. Among all system parameters, we assume that only the wireless channel gain is time-varying. Generally speaking, the trajectory optimization scheme proposed by the present invention improves the overall weighted sum computing rate of the system, and the proposed resource optimization research scheme for the edge computing system based on deep reinforcement learning achieves a performance close to the optimal similar to the existing benchmark test methods, making up for the deficiencies of the previous methods such as the curse of dimensionality, slow convergence speed, and inapplicability to handling high-dimensional action spaces.

[0029] Embodiment 1

[0030] As Figure 1 shown, it is a resource optimization method based on mobile edge computing provided by an embodiment of this application, mainly for trajectory design of AGV by combining radio maps in an intelligent factory and optimization of resource allocation using deep reinforcement learning, specifically including the following steps:

[0031] Step S110: Simulate the radio map of the intelligent factory.

[0032] Consider an intelligent factory in an urban environment, where users are evenly placed on the roof of the factory at a fixed height of H = 50 meters. Since various machines and buildings in the factory may be densely distributed, the signals transmitted from the users are likely to be severely blocked by obstacles. Assume that an AGV moves inside the factory. The channel model of the link between the AGV and the user is a segmented nested model: in the classical large-scale fading channel model, the channel gain is modeled as G dB = b - a log 10 d + ε, where ε is a random variable that captures the shadowing effect. G dB is divided into several components, each associated with a set of parameters a, b, and a random component ε with a specific degree of blockage k .

[0033] Specifically: Let D ∈ R 2 be the region of all possible positions x of the AGVs in the intelligent factory. Consider dividing D into K non-overlapping segments D = D1(X U ) ∪ D2(X U ) ∪....D K (X U ), where D k (X u ) represents the region of AGV positions where the AGV has a k-degree LOS obstacle to the user. The proposed segmented propagation model of G dB is specified as:

[0034]

[0035] d u (X) is the distance between the AGV and the user, where the user's position is (X U , H), the AGV's position is (X, 0), and X, X U ∈ R 2 represent the horizontal positions of the AGV and the user respectively. Thus, (X U , H), (X, 0) are the positions of the user and the AGV in the three-dimensional image respectively. a k and b k are some parameters, and Ⅱ{A} is an indicator function that takes the value 1 if the condition A is satisfied and 0 otherwise. The random variable ε k captures the remaining shadowing effect. Assume that the statistics of a k , b k and ε k are completely known, so that the radio map of the entire intelligent factory is simulated.

[0036] Step S120: In response to the completion of the simulation of the radio map of the intelligent factory, set the initial trajectory of the automatic guided vehicle (AGV) and the positions of each user, and generate the channel state information corresponding to the initial trajectory.

[0037] Among them, the set user positions are the horizontal position and height information of the set users. Suppose there are N users, and the horizontal positions G of the N users are respectively G = {G1, G2, G3... G N ), where G N represents the horizontal position of the Nth user. The heights of the N users are uniformly set to H.

[0038] Among them, the initial trajectory of the automatic guided vehicle can be set as a circle with a specified radius, or an initial trajectory can be randomly initialized.

[0039] After setting the initial trajectory of the AGV, it also includes collecting multiple groups of channel state information between the AGV positions and users on the initial trajectory using the radio map. Each group is respectively Among them, h1 represents the channel gain between the AGV and the first user, h2 represents the channel gain between the AGV and the second user, and so on, h N represents the channel gain between the AGV and the Nth user. Among them, i represents the position of the AGV, and for each AGV position, there is such a group of channel gains. Since the channel gain is time-varying, small-scale fading α t is mixed in, that is Among them, the small-scale fading α t is generated by the Rayleigh fading channel model and is an independent random channel fading factor following the unit mean exponential distribution.

[0040] Among them, after setting the initial trajectory of the automatic guided vehicle and the positions of each user, steps S140 - S160 are executed to obtain the computing rate obtained with the initial trajectory as the input.

[0041] Step S130: Optimize the movement trajectory of the automatic guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automatic guided vehicle, and generate the channel state information corresponding to the optimized stack.

[0042] Specifically, optimize the movement trajectory of the AGV according to the obtained radio map and the channel quality between the AGV and the users. Specifically, it includes the following sub-steps:

[0043] Step S1301: Print out multiple AGV positions collected on the initial trajectory of the AGV.

[0044] Step S1302: Set the possible movement directions according to the AGV positions.

[0045] Specifically, for the positions on the initial trajectory of each AGV, it is stipulated that there are eight possible positions after trajectory optimization. They are: moving one step length at 0 degrees, one step length at 45 degrees, one step length at 90 degrees... one step length at 315 degrees. That is, the AGV has eight possibilities of moving in different directions.

[0046] Step S1303: Determine the position of each AGV after movement according to the calculated possible movement directions.

[0047] Among them, for the position of each AGV, the position of the AGV after calculating eight possible movement directions, that is, the position of the AGV after moving according to these eight possible movement directions.

[0048] Among them, the position of the AGV after eight possible movement directions can be obtained by the python method. Specifically, set a group of data, which are [1, -1], [1, 0], [1, 1], [-1, 0], [0, 1], [0, -1], [-1, -1], [-1, 1]. Multiply each data by the step length and then add the initial trajectory point to obtain the position of the AGV after eight possible movement directions.

[0049] Step S1304: Calculate the combined channel gain between the position of each AGV after movement and each user.

[0050] Among them, according to the positions of the AGV after eight movement schemes, using the radio map, calculate the combined channel gain between each position and each user.

[0051] Among them, directly read the channel gains between the AGV and each user and accumulate them, that is, obtain the combined channel gain between each position and each user.

[0052] Step S1305: Select the scheme with the largest combined channel gain as the scheme for optimizing the trajectory of each AGV position according to the combined channel gain between the position of each AGV after movement and each user.

[0053] Among them, after obtaining the optimized trajectory, use the channel state information collected by the optimized trajectory as the input of the neural network model. That is, after executing step S130, execute steps S140 - S160 to obtain the computing rate obtained by using the optimized trajectory as the input.

[0054] Step S140: Build a neural network model, and the neural network model generates offloading decisions using the channel state information.

[0055] Among them, the channel state information for generating offloading decisions includes the channel state information obtained according to the initialization trajectory and the channel state information obtained according to the optimized trajectory, and offloading decisions are generated according to the two kinds of channel state information respectively.

[0056] Specifically, the constructed neural network model is a fully connected DNN composed of an input layer, two hidden layers, and an output layer. The first and second hidden layers have 120 and 80 hidden neurons respectively. Since a simple two-layer perceptron can already achieve satisfactory convergence performance, better convergence performance of the neural network model can be obtained by further optimizing the DNN parameters.

[0057] During the construction of the neural network model, an offloading policy function y is designed, which can quickly generate an optimal offloading action x. Therefore, by using DNN to generate offloading actions, its characteristic is the embedded parameter u, such as the weights connecting the hidden neurons.

[0058] Specifically, within the t-th time period, the DNN takes the channel gain (When the channel state information obtained by initializing the trajectory in step S120 is input into the neural network model, then this channel gain represents the channel gain corresponding to the initialized trajectory. When the channel state information of the optimized trajectory in step S130 is input into the neural network model, then this channel gain can also represent the channel gain corresponding to the optimized trajectory) as the input, and outputs a relaxed offloading action x according to the current offloading policy t , parameterized by u t , where where f ut represents a function parameterized by ut, which takes the channel gain as the input and obtains a relaxed offloading action after being parameterized by ut. Among them, the offloading action x t can also be expressed as x t ={x t,i ∈[0,1], i = 1...N}. x t,i represents the i-th entry of x t , and N represents the number of entries.

[0059] Furthermore, the ReLu function is applied to the neurons as the activation function. In the output layer, in this embodiment, a sigmoid activation function w = 1 / (1 + e -v ) is used, which can make the relaxed offloading action satisfy x t,i ∈[0,1]. Then the relaxed action is quantized into K binary offloading actions. Among them, K is a design parameter. The quantization function g k : x t →{x k ∈{0,1} N , k = 1...K}, specifically, in this embodiment, an order-preserving quantization method is used, and the quantization follows the following rules:

[0060] The first binary offloading decision X1 is:

[0061]

[0062] To generate the remaining K - 1 offloading decisions, first sort the entries of x t specifically by using the i-th order statistic to sort each of its entries, |x t,(1) - 0.5| ≤ |x t,(2) - 0.5| ≤... ≤ |x t,(i) - 0.5| ≤... ≤ |x t,(N) - 0.5|, where x t,i is the i-th order statistic of x t The result of the k-th offloading decision X k is based on the following rules: where k = 2...K.

[0063]

[0064] Since x t has N order statistics, and each order statistic can generate a quantization action from the above two rules, this order-preserving quantization method can generate at most (N + 1) quantization actions. That is, the value of K ranges between [1, N + 1].

[0065] The offloading decisions and offloading actions generated above can represent the offloading decisions and offloading actions obtained by using the channel gains corresponding to the initial trajectory of step S120 as input, or the offloading decisions and offloading actions obtained by using the channel gains corresponding to the optimized trajectory of step S130 as input.

[0066] Step S150: In response to generating the offloading decision, perform resource allocation to obtain the optimal offloading action.

[0067] Before performing resource allocation, it also includes dividing the system time into consecutive time frames of equal length T. Each time frame consists of the time when the user charges itself for energy and the time when the user offloads to the edge computing server carried by the AGV. Considering that the AGV has a stable power supply and each user has a rechargeable battery that can store the energy obtained from charging to power the operation of the device. To maximize the computing rate, the user consumes all the energy it collects whether in the offloading or local computing mode. Therefore, the resource allocation includes the following sub-steps:

[0068] Step S1501: Perform resource allocation in the local computing mode.

[0069] Step S1502: Perform resource allocation in the edge computing mode.

[0070] Among them, in step S1401, let fi represents the computing speed (cycles per second) of the user processor, t i represents the computing time, 0 ≤ t i ≤ T, where T represents the time of the entire time frame, and the number of bits processed by the user is where represents the number of cycles required to process one-bit task data. Among them, the energy consumption of user computing is limited, k i f i 3 t i ≤ E, E = F * aT. Among them, k i represents the computing energy efficiency coefficient. Within one time frame, aT is the user charging time, and F is the charging efficiency, where a ∈ [0, 1].

[0071] To process the maximum amount of data under energy constraints, the user should consume all the collected energy and calculate within the entire time frame. To maximize the calculation rate, calculate within the entire time frame, so the identification symbol asterisk is added. Therefore: t * = T, f i * = (E / k i T) 1 / 3 , Let μ1 = (F / k i ) 1 / 3 , and the local calculation rate (in bits per second) is

[0072] In step S1502, it is stipulated that the user can only unload its tasks to the AP after collecting energy in the offloading mode. Let τ i T represent the offloading time of the i-th user, τ i ∈ [0, 1]. Assuming that the computing speed and transmission power ratio of the edge computing server carried by the AGV are much larger than those of the user limited by size and energy, the time spent by the edge computing server carried by the AGV in task calculation and download is ignored. Each time frame is only occupied by the charging time and the task offloading time. The offloading calculation rate is equal to the user data offloading capacity: B represents the communication bandwidth, N0 represents the noise power, V u refers to the communication overhead in task offloading, and F represents the charging efficiency.

[0073] Therefore, the weighted sum calculation rate of the MEC network of the entire intelligent factory is expressed as:

[0074]

[0075] For the above problem of maximizing the computing rate, a double-section search algorithm can be used to calculate the optimal transmission time allocation to maximize the computing rate.

[0076] Each offloading decision X in step S140 k (At this time, if it is the execution process of step S120, it is the candidate action corresponding to the initial trajectory. If it is the execution process of step S130, it is the candidate action corresponding to the optimized trajectory) can achieve the maximization of the computing rate by solving the above convex problem. Therefore, the best offloading action in the t-th time frame is (At this time, if it is the execution process of step S120, it is the best offloading action corresponding to the initial trajectory. If it is the execution process of step S130, it is the best offloading action corresponding to the optimized trajectory).

[0077] Step S160: Update the offloading strategy according to the obtained best offloading action to obtain the optimal offloading strategy.

[0078] Among them, the best offloading action is used to update the offloading strategy of the DNN.

[0079] Specifically, this embodiment maintains an initially empty memory with a finite capacity. In the t-th time period, a new training data sample is added to the memory When the memory is full, the new data sample will replace the oldest data sample. Here, the experience replay technique is used to train the DNN using the stored sample data. In the t-th time frame, a batch of training data samples is randomly selected, with a set of time exponents as features. To reduce the average cross-entropy loss, the Adam algorithm is applied to update the parameters u of the DNN t . Among them, the average cross-entropy loss L(u t ) is expressed as:

[0080]

[0081] After collecting a sufficient number of new data samples, the DNN is trained every δ time frames. The DNN iteratively learns from the best state-action pairs and generates better offloading decision outputs over time, thereby completing the update of the offloading strategy every δ time frames.

[0082] Among them, the updated offloading strategy is repeatedly executed in steps S140 - S160 to continuously obtain a new average cross-entropy loss L(u t ). When the average cross-entropy loss L(u t ) is less than the specified threshold, the corresponding optimal offloading strategy is obtained. According to this offloading strategy, the weighted sum computing rate of the MEC network is the final system computing rate of the current time frame.

[0083] Among them, due to step S120, steps 140 - 160 can obtain the system calculation rate corresponding to the initial trajectory, and due to step S130, steps 140 - 160 can obtain the system calculation rate corresponding to the optimized trajectory. Therefore, compare the system calculation rate corresponding to the optimized trajectory with the system calculation rate corresponding to the initial trajectory. If the performance improvement (the difference between the calculation rates) is less than the specified threshold, it indicates that the performance converges, and then output the final system calculation rate. If there is a performance improvement (the difference between the calculation rates is greater than the specified threshold), then continue to repeat steps S120, 140 - 160, and steps S130 - 160.

[0084] Among them, the above-mentioned steps S120, 140 - 160, and steps S130 - 160 are executed in sequence, that is, first execute steps S120, 140 - 160, and after completion, execute steps S130 - 160, so as to complete the comparison of the system calculation rates before and after trajectory optimization.

[0085] Embodiment 2

[0086] As Figure 2 shown, it is a resource optimization system based on mobile edge computing provided by an embodiment of the present application, specifically including: a simulation unit 210, an initial setting unit 220, an optimization unit 230, an offloading decision generation unit 240, an optimal offloading action acquisition unit 250, and an optimal offloading strategy acquisition unit 260.

[0087] Among them, the simulation unit 210 is used to simulate the radio map of the intelligent factory.

[0088] The initial setting unit 220 is connected to the simulation unit 210 and is used to set the initial trajectory of the automated guided vehicle and the positions of each user.

[0089] The optimization unit 230 is respectively connected to the simulation unit 210 and the initial setting unit 220, and is used to optimize the movement trajectory of the automated guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automated guided vehicle.

[0090] The offloading decision generation unit 240 is respectively connected to the initial setting unit 220 and the optimization unit 230, and is used to build a neural network model, and the neural network model uses channel state information to generate offloading decisions corresponding to the initial trajectory and offloading decisions corresponding to the optimized trajectory.

[0091] The optimal offloading action acquisition unit 250 is connected to the offloading decision generation unit 240, and is used to perform resource allocation after generating offloading decisions corresponding to the initial trajectory and offloading decisions corresponding to the optimized trajectory, and obtain the optimal offloading action corresponding to the initial trajectory and the optimal offloading action corresponding to the optimized trajectory.

[0092] The optimal offloading strategy acquisition unit 260 is connected to the best offloading action acquisition unit 250, and is used to obtain the best offloading strategy corresponding to the initial trajectory and the best offloading strategy corresponding to the optimized trajectory according to the best offloading action corresponding to the obtained initial trajectory and the best offloading action corresponding to the optimized trajectory.

[0093] Wherein, the best offloading strategy corresponding to the initial trajectory and the best offloading strategy corresponding to the optimized trajectory are respectively input into the neural network model again, and the steps executed by the offloading decision generation unit 240, the best offloading action acquisition unit 250 and the optimal offloading strategy acquisition unit 260 are performed to obtain the system calculation rate corresponding to the initial trajectory and the system calculation rate corresponding to the optimized trajectory. The system calculation rates generated by the two are compared. If the performance improvement (the difference between the calculation rates) is less than the specified threshold, it means that the performance converges, and then output is performed. If there is performance improvement (the difference between the calculation rates is greater than the specified threshold), then the steps of the initial setting unit 220, the offloading decision generation unit 240, the best offloading action acquisition unit 250, and the optimal offloading strategy acquisition unit 260 are continued to be repeated, as well as the steps of the execution optimization unit 230, the offloading decision generation unit 240, the best offloading action acquisition unit 250, and the optimal offloading strategy acquisition unit 260.

[0094] The present application has the following beneficial effects:

[0095] (1) The present application provides an idea of simulating the radio map of the entire intelligent factory, which is convenient for subsequent trajectory optimization of AGV and the construction of the MEC network.

[0096] (2) The present application provides a solution for trajectory design of AGV in an intelligent factory in combination with a radio map, which can effectively improve the overall system calculation rate.

[0097] (3) The present application provides an idea of decomposing the original optimization problem into two sub-problems: namely, the offloading decision sub-problem and the resource allocation sub-problem. Deep reinforcement learning is used to generate offloading decisions, and then convex optimization is used to optimize the research on the charging time and offloading time resources of users in the intelligent factory. It is applicable to continuous state spaces and does not require discretization of channel gains, making up for the problems of the previous curse of dimensionality and slow convergence speed, and maximizing the overall weighted sum calculation rate of the MEC system in the intelligent factory.

[0098] Although the examples referred to in the current application are described, they are only for the purpose of explanation and not a limitation of the present application. Changes, additions, and / or deletions to the embodiments can be made without departing from the scope of the present application.

[0099] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.

Claims

1. A resource optimization method based on mobile edge computing, characterized in that Specifically, it includes the following steps: Simulate the radio map of the smart factory; simulating the radio map of the smart factory includes setting D ∈ R 2 is the area of the positions x of all possible AGVs in the smart factory. Consider dividing D into K non - overlapping segments D = D1(X U ) ∪ D2(X U ) ∪....D K (X U ), where k ≠ j, D k (X u ) represents the AGV position area where the AGV maintains a k - degree LOS obstacle for the user; the proposed G dB 's segmented propagation model is specified as: d u (X) is the distance between the AGV and the user, and the user's position is (X U ,H), the AGV's position is (X, 0), X, X U ∈R 2 represent the horizontal positions of the AGV and the user respectively; thus (X U ,H), (X, 0) are the positions of the user and the AGV in the three-dimensional image respectively; a k and b k are some parameters, Ⅱ{A} is an indicator function that takes the value 1 if the condition A is satisfied and 0 otherwise; the random variable ε k captures the remaining shadow effect; In response to the completion of the simulation of the radio map of the intelligent factory, set the initial trajectory of the automated guided vehicle and the positions of each user, and generate the channel state information corresponding to the initial trajectory; Optimize the movement trajectory of the automated guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automated guided vehicle, and generate the channel state information corresponding to the optimized stack; Build a neural network model, which uses channel state information to generate offloading decisions; among them, within the t-th time period, the DNN uses the channel gain as input and outputs a relaxed offloading action x according to the current offloading strategy t , parameterized by u t , where where f ut represents a function parameterized by ut, inputs the channel gain, and obtains a relaxed offloading action after being parameterized by ut; the offloading action x t can also be expressed as x t ={x t,i ∈[0,1], i = 1...N}, x t,i represents the i-th entry of x t , and N represents the number of entries; apply the ReLu function on the neurons as the activation function, and in the output layer, use a sigmoid activation function, the sigmoid activation function w = 1 / (1 + e -v ), and then quantize the relaxed action into K binary offloading actions, where K is a design parameter; the quantization function g k : x t →{x k ∈{0,1} N , k = 1...K}, and the quantization follows the following rules: the first binary offloading decision X1 is: To generate the remaining K - 1 offloading decisions, first sort the entries of x t , and sort its various entries by using the i-th order statistic, |x t,(1) -0.5| ≤ |x t,(2) -0.5| ≤... ≤ |x t,(i) -0.5| ≤... ≤ |x t,(N) -0.5|, x t,(i) is the i-th order statistic of x t ; the result of the k-th offloading decision X k is based on the following rules: where k = 2...K; Since x t has a total of N order statistics, the value of K is between [1, N + 1]; In response to the generation of the unloading decision, perform resource allocation to obtain the best unloading action; Update the unloading strategy according to the obtained best unloading action to obtain the optimal unloading strategy.

2. The resource optimization method based on mobile edge computing according to claim 1, wherein The unloading decisions generated by the neural network model using the channel state information include the unloading decisions generated according to the channel state information corresponding to the initial trajectory and the unloading decisions generated according to the channel state information corresponding to the optimized trajectory; The best unloading actions include the best unloading actions corresponding to the initial trajectory obtained after executing the unloading decisions generated according to the channel state information corresponding to the initial trajectory, and the best unloading actions corresponding to the optimized trajectory obtained after executing the unloading decisions generated according to the channel state information corresponding to the optimized trajectory; The optimal unloading strategies include the optimal unloading strategies corresponding to the initial trajectory obtained after executing the best unloading actions corresponding to the initial trajectory, and the optimal unloading strategies corresponding to the optimized trajectory obtained after executing the best unloading actions corresponding to the optimized trajectory.

3. The resource optimization method based on mobile edge computing according to claim 2, wherein Compare the system computing rates corresponding to the initial trajectory obtained according to the optimal unloading strategy corresponding to the initial trajectory and the system computing rates corresponding to the initial trajectory obtained according to the optimal unloading strategy corresponding to the optimized trajectory. If the difference in the computing rates is less than the specified threshold, output it.

4. The resource optimization method based on mobile edge computing according to claim 1, characterized in that The setting of the initial trajectory of the automated guided vehicle and the positions of each user includes that, assuming there are N users, the horizontal positions G of the N users are respectively G = {G1, G2, G3... G N ), where G N represents the horizontal position of the Nth user, and the heights of the N users are uniformly set to H.

5. A resource optimization system based on mobile edge computing, characterized in that Specifically, it includes: A simulation unit, an initial setting unit, an optimization unit, an unloading decision generation unit, a best unloading action acquisition unit, an optimal unloading strategy acquisition unit, and an output unit; The simulation unit is used to perform the simulation of the radio map of the intelligent factory; The simulation of the radio map of the smart factory includes setting \(D\in\mathbb{R}\) 2 is the area of the positions \(x\) of all possible AGVs in the smart factory. Consider dividing \(D\) into \(K\) non - overlapping segments \(D = D_1(X U )\cup D_2(X U )\cup....D K (X U ), where k\neq j, D k (X u ) represents the area of AGV positions where the AGV maintains a \(k\) - degree LOS obstacle for the user; the proposed piece - wise propagation model of \(G dB is specified as: d u (X) is the distance between the AGV and the user, and the user's position is (X U ,H), the AGV's position is (X,0), X, X U ∈R 2 represent the horizontal positions of the AGV and the user respectively; thus (X U ,H), (X,0) are the positions of the user and the AGV in the three-dimensional image respectively; a k and b k are some parameters, Ⅱ{A} is an indicator function that takes the value 1 if condition A is satisfied and 0 otherwise; the random variable ε k captures the remaining shadow effect; The initial setting unit is used to set the initial trajectory of the automated guided vehicle and the positions of each user; The optimization unit is used to optimize the movement trajectory of the automated guided vehicle according to the simulated radio map of the intelligent factory and the initial trajectory of the automated guided vehicle; An offloading decision generation unit is used to build a neural network model, and the neural network model generates offloading decisions using channel state information; where in the t-th time period, the DNN uses the channel gain as input and outputs a relaxed offloading action x according to the current offloading strategy t , which is parameterized by u t , where where f ut represents a function parameterized by ut, takes the channel gain as input, and after being parameterized by ut, obtains a relaxed offloading action; the offloading action x t can also be expressed as x t ={x t,i ∈[0,1], i = 1...N}, x t,i represents the i-th entry of x t , and N represents the number of entries; the ReLu function is applied on the neurons as the activation function, and in the output layer, a sigmoid activation function is used, the sigmoid activation function w = 1 / (1 + e -v ), and then the relaxed action is quantized into K binary offloading actions, where K is a design parameter; the quantization function g k : x t →{x k ∈{0,1} N , k = 1...K}, and the quantization follows the following rules: the first binary offloading decision X1 is: To generate the remaining K - 1 offloading decisions, first sort the entries of x t , and sort its respective entries by using the i-th order statistic, |x t,(1) - 0.5| ≤ |x t,(2) - 0.5| ≤... ≤ |x t,(i) - 0.5| ≤... ≤ |x t,(N) - 0.5|, x t,(i) is the i-th order statistic of x t ; the result of the k-th offloading decision X k is based on the following rules: where k = 2...K; Since x t has a total of N order statistics, the value of K ranges between [1, N + 1]; The best unloading action acquisition unit is used to perform resource allocation to obtain the best unloading action; The optimal unloading strategy acquisition unit is used to update the unloading strategy according to the obtained best unloading action to obtain the optimal unloading strategy.

6. The resource optimization system based on mobile edge computing according to claim 5, wherein, The unloading decisions generated by the neural network model in the unloading decision generation unit using the channel state information include the unloading decisions generated according to the channel state information corresponding to the initial trajectory and the unloading decisions generated according to the channel state information corresponding to the optimized trajectory; The best unloading actions obtained by the best unloading action acquisition unit include the best unloading actions corresponding to the initial trajectory obtained after executing the unloading decisions generated according to the channel state information corresponding to the initial trajectory, and the best unloading actions corresponding to the optimized trajectory obtained after executing the unloading decisions generated according to the channel state information corresponding to the optimized trajectory; The optimal offloading strategy obtained by the optimal offloading strategy acquisition unit includes the optimal offloading strategy corresponding to the initial trajectory obtained after performing the best offloading action corresponding to the initial trajectory, and the optimal offloading strategy corresponding to the optimized trajectory obtained after performing the best offloading action corresponding to the optimized trajectory.

7. The resource optimization system based on mobile edge computing according to claim 5, characterized in that, In the optimal offloading strategy acquisition unit, the system computing rate corresponding to the initial trajectory obtained according to the optimal offloading strategy corresponding to the initial trajectory, and the system computing rate corresponding to the initial trajectory obtained according to the optimal offloading strategy corresponding to the optimized trajectory are compared with the system computing rates obtained before and after trajectory optimization. If the difference in computing rates is less than the specified threshold, output is performed.

8. The resource optimization system based on mobile edge computing according to claim 5, characterized in that, The initial setting unit sets the initial trajectory of the automated guided vehicle and the positions of each user, including that, assuming there are N users, the horizontal positions G of the N users are respectively G = {G1, G2, G3... G N ), where G N represents the horizontal position of the Nth user, and the heights of the N users are uniformly set to H.

Citation Information

Patent Citations

  • Calculation unloading method and system in many-to-many edge computing scene

    CN113900739A

  • Computing unloading optimization method, device and system for mobile edge computing network

    CN114698125A