A Service Migration Method Based on Reinforcement Learning in Mobile Edge Computing
By using reinforcement learning methods in mobile edge computing to build reward function and state transfer matrix, combining value iteration algorithm and Sarsa algorithm, the problem that service migration strategies in the existing technology fail to fully consider environmental factors and user mobility characteristics is solved, efficient migration decisions and dynamic path selection are achieved, and service quality and system efficiency are improved.
Patent Information
- Application Number
- CN202111492744.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-08
AI Technical Summary
Service migration strategies in existing mobile edge computing often fail to fully consider environmental factors and user mobility characteristics, and migration decisions and path selection are not sufficient to cope with dynamic network environments.
A service transfer method based on reinforcement learning is proposed. By constructing a reward function and state transfer matrix that comprehensively considers delay and transfer consumption, a value iteration algorithm is used to make transfer decisions, and dynamic path selection is performed through Sarsa reinforcement learning algorithm.
It realizes migration decisions that comprehensively consider latency and migration consumption in a mobile edge environment, and can dynamically update migration paths to adapt to network changes, thereby improving service quality and system efficiency.
Smart Images

Figure CN114339879B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile edge computing, and in particular relates to a service migration method based on reinforcement learning in mobile edge computing. Background Art
[0002] Mobile edge computing provides better services to users by delegating computing resources to various nodes closer to users, thereby improving the system's quality of service (QoS) and user experience (QoE). The emergence of mobile edge computing has promoted the development of the Internet of Things, 5G, and personalized services of operators. It is also widely used in fields such as augmented reality (AR), video optimization and acceleration, video stream analysis, the Internet of Things (IoT), and Internet of Vehicles. The main problem in mobile edge computing is task offloading, which mainly includes three aspects: task offloading decision-making, resource allocation, and mobility management. Among them, the mobility management problem arises from the mobility of users. An effective way to solve the mobility management problem is service migration. Service migration greatly reduces latency and provides users with better services by migrating services running on servers far away from users to servers closer to users.
[0003] Existing service migration strategies are mainly implemented through Markov decision processes, time window technology, and prediction technology. As a problem caused by user mobility, service migration often considers long-term optimization. Therefore, the service migration problem can be modeled as a time series decision problem for solution. As a classic formal representation of time series decision problems, Markov decision processes can be used to study service migration problems. Time window technology and prediction technology can well predict future energy consumption, and then find the optimal service placement strategy, which can also be used to solve service migration problems.
[0004] Although service migration has been studied a lot using the above methods, studies based on Markov decision processes often do not adequately consider environmental factors, and rarely consider the actual mobility characteristics of users. The question of how to migrate after the migration decision is made is also rarely mentioned. Therefore, it is crucial to comprehensively consider multiple environmental factors to make migration decisions and choose the appropriate migration path.
[0005] After retrieval, the application publication number is CN110347495A, which is a task migration method for mobile edge computing using deep reinforcement learning. First, set the parameters of the system model, then describe the decision-making formula in reinforcement learning, and then give the task migration algorithm based on the formula. Through this method, an efficient task migration mechanism can be obtained, and the efficient task migration mechanism can improve the system real-time performance, make full use of computing resources, and reduce energy consumption. This method also uses the idea of deep reinforcement learning for task scheduling, that is, to decide whether to migrate the computing task. In particular, the Markov decision process is used, which can give a better solution in a very short time and has strong real-time performance. This method is applicable when the user is in a high-speed moving state to solve the problem of whether to change the used server base station. In this patent, the deep reinforcement learning algorithm is used to solve the task migration problem in mobile edge computing. However, due to the uncertainty of the moving state, it often cannot cover all the moving trajectories of the user. This patent predicts the user's future moving direction based on the user's previous moving direction, and then constructs a user movement model. Then, combined with the decision of whether to migrate, a state transition matrix is constructed, which can cover all possible user movement states, and then the migration decision problem more in line with the actual scenario can be solved. At the same time, this patent also uses the reinforcement learning algorithm to solve the problem of the selection of the migration path.
[0006] The application publication number is CN110830560A, which is a multi-user mobile edge computing migration method based on reinforcement learning, including the following steps: First, the mobile device determines its current workload arrival rate, renewable energy, battery power and other states. Then, by accessing the action-state value matrix, according to the ∈-greedy policy, it decides the amount of tasks to be processed locally and takes corresponding actions. Then, calculate the reward value that can reflect the quality of the current action and update the action-state value matrix with this. Finally, calculate the total cost of the mobile device (including delay cost and computing cost). The present invention applies reinforcement learning to the mobile edge computing technology, which is one of the key technologies of 5G, and combines the advantages of the model-free Q-learning to formulate a task allocation strategy for mobile devices, significantly reducing the cost of mobile devices. This patent solves the multi-user task migration problem through reinforcement learning, mainly used to solve the long-term cost optimization problem of the system during the task offloading process. Different from this patent's solution to the problem of formulating migration decisions under different positions and different moving directions, at the same time, this patent also solves the problem of the selection of the migration path. Summary of the Invention
[0007] The present invention aims to solve the service migration problem in existing mobile edge computing, proposes a migration decision-making model that comprehensively considers the influence of various environmental factors and mobile prediction, and uses the value iteration algorithm to solve the problem. At the same time, the reinforcement learning algorithm is used to realize the selection of the adaptive migration path in the dynamic network environment. The technical solution of the present invention is as follows:
[0008] A service migration method based on reinforcement learning in mobile edge computing, which includes the following steps:
[0009] S1. Construct a reward function based on the server location where the user task is located, the area location where the user is currently located, and the load of the server processing the task;
[0010] S2. Construct a state transition matrix according to the current location of the user, the previous moving direction, and the migration decision;
[0011] S3. Make a migration decision using the value iteration algorithm according to the reward function and the state transition matrix;
[0012] S4. Assign link consumption by normalizing the delay consumption and network consumption between routes;
[0013] S5. Select a path using the Sarsa reinforcement learning algorithm according to the normalized link consumption and adaptively update the link selection to adapt to the link changes of the dynamic network;
[0014] The construction of the reward function according to the server location where the user task is located, the area location where the user is currently located, and the load of the server processing the task specifically includes:
[0015] (S11) Use the distance d of the user from the server processing the task t and the load h of the server processing the task t to construct a user service satisfaction function;
[0016] (S12) Use the distance d of the user from the server processing the task t to construct a migration consumption function;
[0017] (S13) Use the weighted sum of the service satisfaction function and the migration consumption function as the reward function;
[0018] The construction of the user satisfaction c1(s t , a t ) using the distance of the user from the server processing the task and the load of the server processing the task, the specific formula is:
[0019] c1(s t , a t ) = D - μ1d t - μ2h t
[0020]
[0021] where D represents the maximum service satisfaction that the user can obtain, d t represents the distance of the user from the server processing the task at time t, ht Indicates the server load situation for processing tasks at time t. μ1 and μ2 are proportionality coefficients, indicating the influence degrees of distance and load on user service satisfaction; d t By calculating the user's current location l t =(x t , y t ) and the Euclidean distance from the location l s =(x s , y s ) of the server for processing tasks is obtained;
[0022] Using the distance d between the user and the server for processing tasks t Construct the migration consumption function c2(s t , a t ):
[0023] c2(s t , a t ) = μ3 + μ4d t
[0024] Among them, the linear function of the distance d t is used to represent the migration consumption, μ3 represents the constant consumption, and μ4 represents the influence coefficient of the distance;
[0025] Use the weighted sum of the user service satisfaction function and the migration consumption function as the reward function r(s, a):
[0026]
[0027] Among them, a represents the migration decision. a = 0 means no migration, and a = 1 means migration; d max represents the maximum distance allowed for the task to be processed. Beyond this distance, there will be a huge penalty M;
[0028] Said constructing the state transition matrix according to the user's current location, previous moving direction, and migration decision, including:
[0029] (S21) Record the user's current location and the moving direction of the user at the previous moment;
[0030] (S22) Different moving directions will affect the user's subsequent moving trajectory. The user's moving model is that the user has a high probability of not changing the direction and a low probability of changing the direction;
[0031] (S23) Based on the user's moving model and migration decision, determine the user's state at the next moment;
[0032] Said recording the moving direction z of the user at the previous moment t , using the user's current location l tWith the previous moving direction z t Indicates the current state s of the user t =(x t , y t , z t );
[0033] The different moving directions z t Will affect the user's subsequent movement trajectory. There is a relatively high probability p that the user will maintain the moving direction z at the next time sequence t Unchanged and reach the position At the same time, there is a relatively low probability that the user will Change the moving direction to Or And reach the position Or
[0034]
[0035] Based on the user's movement model and migration decision, determine the state transition probability P(s'|s,a):
[0036]
[0037] Among them, Indicates that after migration, the user and the server processing the task are in the same location; Indicates that after migration, the user's moving direction remains unchanged, and at the same time, there is a probability p that the user's moving direction remains unchanged when not migrating;
[0038] According to the reward function and the state transition matrix, use the value iteration algorithm to make migration decisions, including:
[0039] (S31) Randomly initialize the state value function v(s) of the user in different positions and different moving directions;
[0040] (S32) Based on the Bellman optimal equation, update the state value function value of the next iteration cycle using the state value function value of the previous iteration cycle. The specific formula is:
[0041]
[0042] Among them, v k+1 (s) represents the state value function corresponding to the state s in the (k + 1)-th iteration cycle, Represents the reward obtained by selecting the action a in the state s, Represents the probability of reaching the state s' by selecting the action a in the state s, and v k (s') represents the state value function corresponding to the state s' in the k-th iteration cycle;
[0043] (S33) Repeat step (S32) until the state value functions in different positions and directions converge;
[0044] The method of assigning link consumption c by normalizing the delay consumption t and network consumption p between routes includes the steps of:
[0045] Record the delay consumption t required for transmission in the link and the network consumption p;
[0046] After normalizing the two and performing weighted summation, assign the link consumption c:
[0047]
[0048] c i = ω t t i + ω p p i
[0049] where, t i and p i represent the delay consumption and network consumption corresponding to each link, represents the minimum value of the link delay consumption, represents the maximum value of the link delay consumption, represents the minimum value of the link network consumption, represents the maximum value of the link network consumption; ω t and ω p respectively represent the weighted coefficients of the delay consumption and the network consumption.
[0050] Furthermore, the method of using Sarsa reinforcement learning to select the migration path includes:
[0051] (1) Randomly initialize the link information connected by each route, including the delay consumption t and the network consumption p;
[0052] (2) Randomly select a path for data transmission from the source server to the target server;
[0053] (3) Record the delay consumption t generated during data transmission and the network consumption p generated, and after normalizing them, obtain the corresponding link consumption c by weighted summation;
[0054] (4) Each route selects the link for data transmission according to the ε-greedy policy, and at the same time records the link consumption for transmitting to the next route through this link. Each route updates its corresponding state-action Q value table according to the transmission of this data;
[0055] (5) Along with the data transmission, each route repeats step (4) to dynamically update the Q value table of this route and select a more optimized path.
[0056] Furthermore, the methods for selecting actions using the ε-greedy strategy and updating the state-action Q-value table are as follows:
[0057]
[0058] Q(S,A)←Q(S,A)+α(R+γQ(S',A')-Q(S,A))
[0059] where π(a|s) represents the probability of selecting action a in state s, a * represents the action that can maximize the Q-value in the current state s, m represents the number of available actions, Q(S,A) represents the state-action function values corresponding to different actions selected in each state, α is the learning rate parameter, γ is the decay factor, and Q(S',A') represents the state-action function value corresponding to the next state.
[0060] The advantages and beneficial effects of the present invention are as follows:
[0061] 1. The present invention comprehensively considers the delay factor concerned by users and the migration consumption factor concerned by merchants in the mobile edge environment, constructs a more realistic mobile model and state transition matrix based on mobile prediction, and finally uses the value iteration algorithm to obtain the migration decisions in different positions and different mobile directions. It guides when users should migrate to maximize the benefits. The migration decisions finally obtained by the present invention are different from other service migration strategies that have the same migration decisions at the same location. Instead, different migration decisions will be generated at the same location due to different previous mobile directions of users, which is more in line with the actual scenario.
[0062] 2. After it is determined that the service needs to migrate in the mobile edge environment, the present invention comprehensively considers the interests of users and merchants, assigns link consumption based on delay and network consumption, and uses the reinforcement learning algorithm to solve the problem of adaptive migration path selection in the dynamic network environment. The migration path finally obtained by the present invention will be updated in real time dynamically as the network link state changes, and can provide a better migration path for the service. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is the migration decision-making algorithm based on value iteration provided by the preferred embodiment of the present invention;
[0064] Figure 2 is the dynamic path selection algorithm based on Sarsa;
[0065] Figure 3 is the flow chart of the service migration method based on reinforcement learning in mobile edge computing. DETAILED DESCRIPTION OF THE INVENTION
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0067] The technical solution for the present invention to solve the above technical problems is:
[0068] As Figure 3 shown, the present invention discloses a service migration method based on reinforcement learning in mobile edge computing, including the following steps:
[0069] S1. Based on the server location l where the user service is located s , the regional location l where the user is currently located t and the server load h for processing tasks t construct a reward function r(s, a). Use the reward function to represent the benefits that the user can obtain each time a migration decision is made, and regard it as an equilibrium manifestation of the user service experience and migration consumption;
[0070] S2. Based on the current location l of the user t , the previous moving direction z t and whether to migrate, construct a state transition matrix P. Use the state transition matrix to represent the state changes generated each time the user makes a migration decision, including changes in location and moving direction;
[0071] S3. Based on the reward function r(s, a) and the state transition matrix P, use the value iteration algorithm to solve the problem of formulating migration decisions. Furthermore, determine whether the user service needs to migrate in different positions and different moving directions.
[0072] S4. Based on the delay consumption t and network consumption p between routes, normalize the two and assign the link consumption c;
[0073] S5. Based on the normalized link consumption c, use the Sarsa reinforcement learning algorithm to select a migration path and adaptively update the link selection to adapt to the link changes of the dynamic network.
[0074] In this embodiment, the method for constructing the reward function r(s, a) according to the server location l where the user service is located s , the regional location l where the user is currently located t and the server load h for processing tasks t in step S1 includes the steps:
[0075] (1) Use the distance d between the user and the server for processing tasks t and the server load h for processing tasks t to construct a user satisfaction function c1(s t , a t):
[0076] c1(s t ,a t ) = D - μ1d t - μ2h t
[0077]
[0078] Among them, D represents the maximum service satisfaction that the user can obtain, d t represents the distance of the user from the task processing server at time t, h t represents the server load situation of the task being processed at time t, and μ1 and μ2 are proportionality coefficients, indicating the influence degrees of distance and load on the user service satisfaction. d t By calculating the Euclidean distance between the user's current position l t =(x t ,y t ) and the position l of the task processing server s =(x s ,y s ).
[0079] (2) Use the distance d of the user from the task processing server t to construct the migration consumption function c2(s t ,a t ):
[0080] c2(s t ,a t ) = μ3 + μ4d t
[0081] Among them, a linear function of the distance d t is used to represent the migration consumption, μ3 represents the constant consumption, and μ4 represents the influence coefficient of the distance.
[0082] (3) Use the weighted sum of the user service satisfaction function and the migration consumption function as the reward function r(s,a):
[0083]
[0084] Among them, a represents the migration decision, a = 0 means no migration, and a = 1 means migration; d max represents the maximum distance allowed for processing the service, and exceeding this distance will result in a huge penalty M.
[0085] In this embodiment, the method for constructing the state transition matrix according to the user's current position l t , the previous moving direction z t and whether to migrate in step S2 includes the steps:[[]]
[0086] (1) Record the user's moving direction z at the previous moment t , and use the user's current location l t and the previous moving direction z t to represent the user's current state s t = (x t , y t , z t ).
[0087] (2) Different moving directions z t will affect the user's subsequent moving trajectory. There is a relatively high probability p that the user will maintain the moving direction z t unchanged and reach the location At the same time, there is a relatively low probability that the user will change the moving direction to or and reach the location or
[0088]
[0089] (3) Based on the user's movement model and migration decision, determine the user's state transition probability P(s'|s,a) as follows:
[0090]
[0091] Among them, indicates that after migration, the user and the server processing the task are in the same location; indicates that after migration, the user's moving direction remains unchanged, and at the same time, there is a probability p that the user's moving direction remains unchanged when not migrating.
[0092] In this embodiment, the method for making a migration decision using the value iteration algorithm according to the reward function r(s,a) and the state transition matrix P in step S3 includes the steps:
[0093] (1) Randomly initialize the state value function v(s) of the user in different positions and different moving directions;
[0094] (2) Update the state value function value of the next iteration cycle based on the Bellman optimal equation using the state value function value of the previous iteration cycle. The specific formula is:
[0095]
[0096] Among them, v k+1 (s) represents the state value function corresponding to the state s in the (k + 1)-th iteration cycle, represents the reward obtained by selecting the action a for the state s, Denote the probability that state s selects action a and reaches state s', and v k (s') represents the state value function corresponding to state s' in the k-th iteration cycle;
[0097] (3) Repeat step (2) until the state value functions in different positions and different directions converge.
[0098] In this embodiment, the method for assigning the link consumption c by normalizing the delay consumption t and network consumption p between routes in step S4 includes the steps:
[0099] (1) Record the delay consumption t required for transmission in the link and the network consumption p;
[0100] (2) After normalizing the two, perform weighted summation to assign the link consumption c:
[0101]
[0102] c i = ω t t i + ω p p i
[0103] where t i and p i represent the delay consumption and network consumption corresponding to each link, represents the minimum value of the link delay consumption, represents the maximum value of the link delay consumption, represents the minimum value of the link network consumption,
[0104] represents the maximum value of the link network consumption; ω t and ω p respectively represent the weighting coefficients of the delay consumption and network consumption.
[0105] In this embodiment, the method for path selection using the Sarsa reinforcement learning algorithm and adaptively updating the link selection to adapt to the link changes in the dynamic network according to the normalized link consumption includes the steps:
[0106] (1) Randomly initialize the link information connected by each route, including the delay consumption t and network consumption p;
[0107] (2) Randomly select a path for data transmission from the original server to the target server;
[0108] (3) Record the delay consumption t and network consumption p generated during data transmission, and normalize and weight them to obtain the corresponding link consumption c;
[0109] (4) Each route selects a data transmission link according to the ε-greedy policy, and at the same time records the link consumption for transmitting to the next route through this link. Each route updates its corresponding state-action Q-value table according to the current data transmission. The methods for selecting actions using the ε-greedy policy and updating the state-action Q-value table are as follows:
[0110]
[0111] Q(S,A)←Q(S,A)+α(R+γQ(S',A')-Q(S,A))
[0112] Among them, π(a|s) represents the probability of selecting action a in state s, a * represents the action that can maximize the Q-value in the current state s, and m represents the number of available actions. Q(S,A) represents the state-action function values corresponding to different actions in each state. α is the learning rate parameter, γ is the decay factor, and Q(S',A') represents the state-action function value corresponding to the next state.
[0113] The present invention comprehensively considers various environmental factors to formulate migration decisions and select migration paths. Compared with existing service migration methods, the present invention has the following main advantages: (1) Comprehensively considering various environmental factors, introducing server load as a factor affecting user service experience, and at the same time introducing the previous moving direction of the user as a prediction index, so that it affects the user's subsequent movement, which is more in line with the actual scenario; (2) Comprehensively considering the delay factor concerned by users and the network consumption factor concerned by service providers, and using reinforcement learning to solve the adaptive service migration path in a dynamic network environment.
[0114] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to this process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.
[0115] The above embodiments should be understood as being only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the content recorded in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A service migration method based on reinforcement learning in mobile edge computing, characterized in that It includes the following steps: S1. Construct a reward function based on the server location where the user task is located, the area location where the user is currently located, and the server load of the currently processed task; S2. Construct a state transition matrix based on the user's current location, the previous moving direction, and the migration decision; S3. Make a migration decision using the value iteration algorithm according to the reward function and the state transition matrix; S4. Assign link consumption by normalizing the delay consumption and network consumption between routes; S5. According to the normalized link consumption, use the Sarsa reinforcement learning algorithm to select a path and adaptively update the link selection to adapt to the link changes of the dynamic network; The construction of the reward function based on the server location where the user task is located, the area location where the user is currently located, and the server load of the processed task server specifically includes: (S11) Use the distance d between the user and the processing task server t and the load h of the processing task server t to construct a user service satisfaction function; (S12) Use the distance d between the user and the task server for processing t Construct a migration consumption function; (S13) Use the weighted sum of the service satisfaction function and the migration consumption function as the reward function; The user satisfaction c1(s t , a t ) is constructed using the distance between the user and the processing task server and the load of the processing task server. The specific formula is as follows: c1(s t ,a t ) = D - μ1d t - μ2h t Among them, D represents the maximum service satisfaction that the user can obtain, and d t represents the distance between the user and the task processing server at time t, and h t represents the server load situation for task processing at time t. μ1 and μ2 are proportionality coefficients, indicating the influence degrees of distance and load on the user's service satisfaction; d t By calculating the Euclidean distance between the user's current position l t =(x t , y t ) and the position of the task processing server l s =(x s , y s ), it is obtained; Use the distance d between the user and the task server to process the task t Construct the migration consumption function c2(s t ,a t ): c2(s t ,a t ) = μ3 + μ4d t where the migration consumption is represented by a linear function of the distance d, μ3 represents the constant consumption, and μ4 represents the influence coefficient of the distance; t Use the weighted sum of the user service satisfaction function and the migration consumption function as the reward function r(s,a): Where, a represents the migration decision, a = 0 means not to migrate, and a = 1 means to migrate; d max Indicates the maximum distance allowed for task processing, and there will be a huge penalty M if the distance is exceeded; The construction of the state transition matrix based on the user's current location, the previous moving direction, and the migration decision includes: (S21) Record the user's current location and the moving direction of the user at the previous moment; (S22) Different moving directions will affect the user's next moving trajectory. The user's moving model is that the user has a greater probability of not changing the direction and a smaller probability of changing the direction; (S23) Based on the user's moving model and the migration decision, determine the user's state at the next moment; The recorded moving direction z of the user at the previous moment t , using the current location l of the user t and the previous moving direction z t to represent the current state s of the user t =(x t , y t , z t ); The different moving directions z t will affect the user's subsequent movement trajectory, and there is a relatively high probability p that the user will maintain the moving direction z t unchanged and reach the position Meanwhile, there is a relatively low probability that the user will change the moving direction to or and reach the position or Based on the user's moving model and the migration decision, determine the state transition probability P(s'|s,a): Among them, indicates that after migration, the user is in the same location as the server handling the task; indicates that after migration, the user's moving direction remains unchanged, and there is a probability p that the user's moving direction remains unchanged when there is no migration; Making a migration decision using the value iteration algorithm according to the reward function and the state transition matrix includes: (S31) Randomly initialize the state value function v(s) of the user in different positions and different moving directions; (S32) Update the state value function value of the next iteration cycle based on the Bellman optimal equation using the state value function value of the previous iteration cycle. The specific formula is: where v k+1 (s) represents the state value function corresponding to the state s in the (k + 1)-th iteration cycle, represents the reward obtained by selecting action a in state s, represents the probability of reaching state s' by selecting action a in state s, and v k (s') represents the state value function corresponding to the state s' in the k-th iteration cycle; (S33) Repeat step (S32) until the state value functions in different positions and different directions converge; The method for assigning link consumption c by normalizing the delay consumption t and network consumption p between routes includes the steps: Record the delay consumption t and network consumption p required for transmission in the link; After normalizing the two and performing weighted summation, assign the link consumption c: c i = ω t t i + ω p p i Among them, t i and p i represent the delay consumption and network consumption corresponding to each link, represents the minimum value of the link delay consumption, represents the maximum value of the link delay consumption, represents the minimum value of the link network consumption, represents the maximum value of the link network consumption; ω t and ω p represent the weighted coefficients of the delay consumption and the network consumption respectively.
2. The service migration method based on reinforcement learning in mobile edge computing according to claim 1, characterized in that, The method for using Sarsa reinforcement learning to select a migration path includes: (1) Randomly initialize the link information connected by each route, including the delay consumption t and network consumption p; (2) Randomly select a path to transmit the data information from the original server to the target server; (3) Record the delay consumption t and network consumption p generated during the data transmission process, and standardize them and then perform weighted summation to obtain the corresponding link consumption c; (4) Each route selects the link for data transmission according to the ε-greedy strategy, and at the same time records the link consumption for transmitting to the next route through this link. Each route updates its corresponding state-action Q-value table according to the current data transmission. (5) Along with the data transmission, each route repeats step (4) to dynamically update the Q-value table of this route and select a more optimized path.
3. The service migration method based on reinforcement learning in mobile edge computing according to claim 2, characterized in that, The methods for selecting actions using the ε-greedy strategy and updating the state-action Q-value table are respectively: Q(S,A)←Q(S,A)+α(R+γQ(S',A')-Q(S,A)) Among them, π(a|s) represents the probability of selecting action a in state s, and a * represents the action that can maximize the Q value in the current state s. m represents the number of available actions. Q(S,A) represents the state-action function values corresponding to different actions selected in each state. α is the learning rate parameter, γ is the decay factor, and Q(S',A') represents the state-action function value corresponding to the next state.
Citation Information
Patent Citations
Task migration method for carrying out mobile edge calculation by using deep reinforcement learning
CN110347495A
Multi-user mobile edge calculation migration method based on reinforcement learning
CN110830560A
Software defined networking load balancingdevice and method
CN105516312A
Super-dense edge computing network mobility management method based on deep reinforcement learning
CN111666149A