An Industrial Internet Edge Service Caching Decision Method and System

Through distributed deep reinforcement learning algorithms, a method to optimize edge service caching strategy is built, which solves the problem of limited edge storage capacity, achieves the effect of minimizing service access delay and energy consumption, and improves the performance of industrial Internet systems.

CN114328291BActive Publication Date: 2025-05-27SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111556974.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-18
Publication Date
2025-05-27
Estimated Expiration
2041-12-18

AI Technical Summary

Technical Problem

The prior art is difficult to flexibly configure edge service caches within limited edge storage capacity, effectively improving the performance of industrial Internet systems, especially in terms of latency and economics.

Method used

Using a distributed deep reinforcement learning algorithm, through mathematical modeling and optimization goals, an algorithm that can achieve the optimal edge service caching strategy is built, aiming to minimize service access delay and energy consumption.

Benefits of technology

Through deep reinforcement learning algorithms, service caching strategies can be effectively optimized under the limited edge storage capacity, significantly reduce service access latency and energy consumption, and improve system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328291B_ABST
    Figure CN114328291B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of industrial Internet, and particularly relates to an industrial Internet edge service caching decision method and system. The industrial Internet edge service caching decision method in the embodiments of the present invention calculates the optimal solution for the edge caching policy mathematical model through an algorithm constructed based on the distributed deep reinforcement learning method, and can solve the optimization problem of the mathematical model of the system. This method is based on the establishment of a network mathematical model and the determination of an optimal objective, and based on the combination of reinforcement learning and deep learning technologies. According to a large amount of user historical data, the machine is allowed to learn and predict the user preference degree and the change trend of the content popularity in the network, and the service caching policy is adjusted according to the learning results. It can effectively give the optimal solution of the service caching decision. Its corresponding system also has the same technical effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial Internet, and more specifically, to an industrial Internet edge service caching decision method and system. Background Art

[0002] As more and more industrial devices are connected to the Internet, it is difficult to meet the requirements of industrial applications in terms of latency and economy only by relying on the traditional cloud computing mode. Edge computing, as an emerging computing paradigm, can alleviate the physical resource bottleneck of intelligent devices. In an edge computing system, service caching can be used to improve traffic load and service quality. However, how to flexibly configure edge service caching within limited edge storage capacity to improve system performance is extremely challenging.

[0003] The existing technologies mainly focus on the research of mobile edge computing caching problems. However, most of the work mainly focuses on improving some caching strategies on traditional networks according to the new characteristics of mobile edge computing networks. There is also a part of the work exploring new caching schemes, such as caching strategies based on user preferences, learning, or multi-edge node cooperation. However, because the content popularity and user preference degrees change with time and are unpredictable. At the same time, the service caching problem is an integer linear programming problem, which cannot be solved within polynomial time, and traditional optimization methods are difficult to effectively achieve the result of optimizing service caching. There are deficiencies in the existing technologies. Summary of the Invention

[0004] Embodiments of the present invention provide an industrial Internet edge service caching decision method and system, which solve at least one of the above technical problems by using a distributed deep reinforcement learning algorithm to solve the optimal edge service caching strategy, so as to achieve the purpose of minimizing service access latency and energy consumption.

[0005] According to an embodiment of the present invention, an industrial Internet edge service caching decision method is provided, including the following steps:

[0006] S1. Based on the fact that only when the corresponding service data is cached in the server can the task corresponding to the service be executed, a mathematical model of the industrial Internet system is established; all the data required for all services is cached in the cloud server of the system model;

[0007] S2. Establish a mathematical model for the service access latency in the edge-cloud collaboration system;

[0008] S3. According to the power of data transmission between edge servers and between edge servers and the cloud server, and the computing power of edge servers and the cloud server, a mathematical model of the energy consumption of the industrial Internet system is established;

[0009] S4. Based on the system model, delay model, and energy consumption model, establish an optimization objective that minimizes service access latency and energy consumption.

[0010] S5. Based on the distributed deep reinforcement learning method, construct an algorithm that can achieve the above optimization objective.

[0011] The present invention also provides an industrial Internet edge service caching decision-making system using the method described in any one of the above, including: a mathematical modeling module and a service caching decision-making module;

[0012] The mathematical modeling module performs mathematical modeling on the industrial Internet system based on the fact that only when the corresponding service data is cached in the server can the tasks corresponding to the service be executed; all the data required for all services is cached in the cloud server of the system model;

[0013] Establish a mathematical model for service access latency in the edge-cloud collaboration system;

[0014] According to the power of data transmission between edge servers and between edge servers and cloud servers, and the computing power of edge servers and cloud servers, perform mathematical modeling on the energy consumption of the industrial Internet system;

[0015] Based on the system model, delay model, and energy consumption model, establish an optimization objective that minimizes service access latency and energy consumption.

[0016] The service caching decision-making module constructs an algorithm that can achieve the above optimization objective based on the distributed deep reinforcement learning method.

[0017] In the industrial Internet edge service caching decision-making method and system of the embodiments of the present invention, the optimal solution of the edge caching policy mathematical model can be calculated through an algorithm constructed based on the distributed deep reinforcement learning method, which can solve some of the above problems. This method determines the digital modeling and optimal solution objective of the industrial Internet system, and based on the combination of reinforcement learning and deep learning technologies, according to a large amount of user historical data, enables the machine to learn and predict the user's preference degree and the changing trend of content popularity in the network, and adjusts the service caching policy according to the learning results. It can effectively give the optimal solution of the service caching decision. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0019] Figure 1 It is a flowchart of the industrial Internet edge service caching decision-making method of the present invention;

[0020] Figure 2 Schematic diagram of the industrial Internet of Things edge service caching decision method of the present invention;

[0021] Figure 3 Schematic diagram of the edge-cloud collaborative service structure of the industrial Internet of Things edge service caching decision system of the present invention. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] As shown in the attached Figure 3 figure, in order to facilitate the modeling (construction of a mathematical model) of the industrial Internet of Things system, the present invention discretizes time into uniformly distributed time slices The interval of each time slice is Δt. Considering that the industrial Internet of Things system consists of N edge servers, denoted as Each edge server can provide data analysis and processing services for sensor devices and industrial devices. Compared with the cloud server, the computing and storage resources on the edge server are limited. The computing power and storage capacity of edge server n are respectively denoted as and Let F cloud represent the computing power of the cloud server.

[0025] Referring to Figures 1-3 , according to an embodiment of the present invention, an industrial Internet of Things edge service caching decision method is provided, including the following steps:

[0026] S1. Mathematically model the industrial Internet system based on the fact that only when the corresponding service data is cached in the server can the tasks corresponding to the service be executed; all the data required for the services is cached in the cloud server of the system model.

[0027] When specifically implementing the mathematical modeling, to execute a certain type of task on the edge server, the corresponding service should be placed first. A service is an abstraction of an application. To run a specific service, the edge server should cache a related data, including the software and database required by the application. Only when the corresponding service data is cached in the server can the tasks corresponding to the service be executed. In the modeling process of the present invention, it is assumed that all the data required for the services is cached in the cloud server.

[0028] In the industrial Internet system, the generated service set is represented as The present invention assumes that different services have different data volumes and require different computing resources and storage resources to process, which are represented by f l and m l respectively, where Each edge server can cache one or more services. The present invention defines a binary variable x l,n (t) ∈ {0, 1} to represent whether the service is cached on the edge server. Then the service caching policy is If the service l is cached on the edge server n at time t, then x l,n (t) = 1; otherwise, x l,n (t) = 0. Let p l,n represent the computing power allocated by the edge server n to the service l. Since service caching is limited by the storage space and computing power of the edge server, therefore:

[0029]

[0030]

[0031] To analyze the access delay of the service, assume that at time t, the number of requests from the industrial device to the edge server n for the service l is λ l,n (t). Due to the diversity of the data collected by the device and the dynamic characteristics of the service requests, λ l,n (t) is dynamically changing. It should be noted that when the requested service is not cached on this edge server, workload scheduling is required to schedule the arriving service requests to the edge server that has cached the service or to execute in the cloud server. Let μ l,n (t) ∈ [0, 1] represent the load ratio of the service l executed on the edge server n. μ l,c (t) represents the ratio of the service l executed on the cloud server, μ l,n , μl,c The value can be set using a variety of different strategies, but generally needs to meet the following conditions:

[0032]

[0033] At time t, the total number of requests for service l in the industrial Internet system is expressed as:

[0034]

[0035] Therefore, the total workload and data size of the edge server n in the system running service l are respectively:

[0036] F l,n (t) = μ l,n f l λ l (t)

[0037] M l,n (t) = μ l,n m l λ l (t);

[0038] Finally, a heterogeneous edge-cloud collaborative offloading framework is realized using a mathematical model, as shown in the appendix Figure 3 As shown, the framework includes a large number of industrial devices and sensor devices, multiple edge servers and a cloud server. The industrial devices and sensor devices communicate with the edge servers through wireless channels, while the edge servers are connected to the remote cloud through wired links. Device tasks can be offloaded to the edge servers or the cloud server for execution. For edge servers without cache services or without sufficient computing power, the corresponding tasks can be offloaded to nearby edge servers with existing cache services or the cloud server for execution. The cooperation between edge nodes can make full use of the heterogeneous edge server resource capacity and alleviate the problem of resource capacity mismatch of a single edge node.

[0039] S2. Establish a mathematical model for the service access delay in the edge-cloud collaborative system.

[0040] In specific implementation, the computing delay of executing service l on edge server n is expressed as:

[0041]

[0042] The computing delay of executing service l in the cloud server is expressed as:

[0043]

[0044] Since the industrial devices are close to the edge servers, the transmission delay between the devices and the edge servers is ignored, and only the workflow transmission delay between adjacent edge servers is considered. Let the data transmission rate between edge servers be r e , then the data transmission delay between edge servers is expressed as:

[0045]

[0046] When the task is offloaded to the remote cloud server for execution, let the data transmission rate of the core network be r c , and the transmission delay of the task offloaded to the cloud for execution is expressed as:

[0047]

[0048] The access delay of running service l includes the transmission delay and the computing delay of the service:

[0049]

[0050] S3. According to the power of data transmission between edge servers and between edge servers and cloud servers, as well as the computing power of edge servers and cloud servers, a mathematical model of the energy consumption of the industrial Internet system is established.

[0051] Specifically, if represents the power of data transmission between edge servers, represents the computing power of edge servers, then the energy consumed by running service l on edge server n at time t is:

[0052]

[0053] If represents the power of core network data transmission between edge servers and cloud servers, represents the computing power of cloud servers, then the energy consumed by running service l on the cloud server at time t is:

[0054]

[0055] Therefore, at time t, the total energy consumption of running service l is expressed as:

[0056]

[0057] S4. Based on the system model, delay model, and energy consumption model, an optimization goal of minimizing the service access delay and minimizing the energy consumption is established;

[0058] In specific implementation, the optimization objective of the present invention solves for the optimal service caching strategy to achieve the goal of minimizing service access latency and energy consumption, where β represents the weight of energy consumption. The optimization objective is as follows:

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065] S5. Construct an algorithm that can achieve the optimization objective based on the distributed deep reinforcement learning method.

[0066] Furthermore, step S5 further includes the following steps:

[0067] S51. Combine multiple parallel deep neural networks DNN with the reinforcement learning algorithm Q-learning to construct a parallel deep reinforcement learning algorithm for service caching decision-making.

[0068] Specifically, in order to minimize service access latency and energy consumption, a parallel deep reinforcement learning algorithm is designed by combining multiple parallel deep neural networks DNN with the reinforcement learning algorithm Q-learning for service caching decision-making.

[0069] Deep reinforcement learning usually defines the problem as a Markov decision process. It mainly focuses on how the agent interacts with the environment, takes different actions, and maximizes the cumulative reward. Its main components include the agent, the environment, the state, the action, and the reward. Therefore, the present invention describes the service optimization caching problem as a Markov decision process. It consists of three parts: the state space S, the action space A, and the reward function R, and the definitions are as follows:

[0070] State space: s n,t ∈S represents the state of the edge server n at time slot t. respectively represent the storage and computing capabilities of the edge server n, the arriving service requests, and the edge service caching strategy.

[0071] Action space: In each time slot t, the edge server needs to make a service caching decision based on the current state, a t =Υ t .

[0072] Reward function: The optimization objective of edge service caching is to minimize service access latency and energy consumption. Therefore, the designed reward function is as follows:

[0073]

[0074] Define the action-value function as Q(s t ,a t ), and the Q-value update:

[0075]

[0076] α ∈ (0, 1] represents the learning rate, and the reward decay factor γ ∈ [0, 1].

[0077] Set m parallel neural network units combined with Q-learning to generate actions. Each neural network action execution is parallel, which includes two DNNs with the same structure but different parameters. One is the main neural network for predicting the Q-estimation value, which has the latest network parameters θ, and the other is the target neural network for predicting the actual Q-value, which uses the parameters θ from some time ago * , and remains unchanged for a period of time. When the main neural network has learned a certain number of times, update the parameters of the target neural network. Each neural network unit will give the selected action value according to the Q-value calculated by the greedy algorithm. The loss function is defined as:

[0078]

[0079] S52. In the training phase, select an action to execute according to the greedy policy based on the current state, obtain the reward and the next state, and store the obtained state transition in the experience pool; when the capacity of the experience pool D is large enough, extract a certain number of state transitions from the experience pool to train the network parameters.

[0080] In the training phase, based on the current state s t select an action a according to the ε-greedy policy t , execute this action to obtain the reward R t and the next state s t+1 , and store the obtained state transition [s t ,a t ,R t ,s t+1 in the experience pool D. When the capacity of the experience pool D is large enough, randomly extract a certain number of state transitions from the experience pool to train the network parameters θ.

[0081] The algorithm flow of the training phase of the edge service caching algorithm based on distributed deep reinforcement learning is as follows:

[0082] Algorithm Input: Environmental information Ψ, reward decay factor γ, learning rate α, exploration-exploitation trade-off parameter ε, experience pool D, number of iterations M, number of update steps C.

[0083] Algorithm Output: Parameters of m DNN networks. The specific training steps are as follows:

[0084] Step 1: Initialize parallel DNNs and initialize the experience pool D.

[0085] Step 2: Iterate for epoch from 1 to M.

[0086] Step 3: Initialize the state s for m parallel DNNs t 。

[0087] Step 4: Iterate for time step t from 1 to T.

[0088] Step 5: Select m actions through the ε-greedy policy i = 1, 2, 3..., m.

[0089] Step 6: In the m DNNs, respectively execute Observe the obtained reward and get the new state s t+1 。

[0090] Step 7: Store in the experience pool D.

[0091] Step 8: Randomly sample n samples [s j , a j , R j (s j , a j ), s j+1 from the experience pool D, where j = 1, 2,..., n.

[0092] Step 9: Calculate the loss function Loss and update the parameters θ of the m main neural networks through neural network gradient backpropagation.

[0093] Step 10: If t % C == 0, assign the parameters θ of the m main neural networks to the target neural network θ * 。

[0094] Step 11: If t <= T, enter the next time step and return to Step 5.

[0095] Step 12: If epoch == M, end the iteration and output the parameters of the m DNN networks.

[0096] S53. In the decision-making stage, multiple caching decisions are generated by multiple parallel deep neural networks and stored in the action set, and the corresponding rewards obtained for each caching decision are calculated, and the caching decision with the maximum reward is used as the output action.

[0097] In the decision-making stage, an optimal service caching policy needs to be obtained. First, the present invention generates m caching decisions a through m parallel DNNs i , and stores them in the action set A, calculates the reward R(s, a i ) and Q(s, a i ) obtained by executing the action a i , and finds the action a = argmax A Q(S t , A) from the action set A, and outputs the action a as the edge service caching policy.

[0098] The algorithm flow of the decision-making stage of the edge service caching algorithm based on distributed deep reinforcement learning is as follows:

[0099] Algorithm input: m DNN network parameters θ.

[0100] Algorithm output: Service caching policy a. The specific steps are as follows:

[0101] Step 1: Iterate from i = 1 to m.

[0102] Step 2: Generate the action a through the i-th DNN i and input it into the action set A (A = {a 1 , a 2 , a 3 ,..., a i}).

[0103] Step 3: Execute a i , obtain the reward R(s, a i ), and calculate Q(s, a i ).

[0104] Step 4: If i == m, select the action a = argmax A Q(S t , A) from the action set A.

[0105] Step 5: Output the action a as the service caching policy.

[0106] Furthermore, in step S1, the edge server provides data analysis and processing services for sensor devices and industrial devices; compared with the cloud server, the edge server has limited computing resources and storage resources.

[0107] Further, in step S1, when the requested service is not cached on the edge server closest to it, the service is executed on the cloud server or another edge server that has cached the service.

[0108] Further, in step S5, the service optimal caching problem is described as a Markov decision process; the Markov decision process consists of three parts: a state space S, an action space A, and a reward function R.

[0109] Further, in step S51, the actions of the deep neural network DNN are executed in parallel. The deep neural network DNN includes two neural network structures with the same structure but different parameters; one is the main neural network for predicting the Q-estimation value of the reinforcement learning algorithm Q-learning, which has the latest network parameters; the other is the target neural network for predicting the actual Q-value of the reinforcement learning algorithm Q-learning, which uses parameters from some time ago and remains unchanged for some time.

[0110] The present invention also provides an industrial Internet edge service caching decision system adopting the method as described in any one of the above, including: a mathematical modeling module and a service caching decision module;

[0111] The mathematical modeling module performs mathematical modeling on the industrial Internet system based on the fact that only when the corresponding service data is cached in the server can the task corresponding to the service be executed; all the data required for all services is cached in the cloud server of the system model;

[0112] Establish a mathematical model for the service access delay in the edge-cloud collaboration system;

[0113] According to the power of data transmission between edge servers and between edge servers and the cloud server, and the computing power of edge servers and the cloud server, perform mathematical modeling on the energy consumption of the industrial Internet system;

[0114] Based on the system model, the delay model, and the energy consumption model, establish an optimization goal of minimizing the service access delay and minimizing the energy consumption;

[0115] The service caching decision module constructs an algorithm that can achieve the optimization goal based on the distributed deep reinforcement learning method.

[0116] Further, the service caching decision module further includes: an algorithm construction unit, a training unit, and a decision unit;

[0117] The algorithm construction unit combines multiple parallel deep neural networks DNN with the reinforcement learning algorithm Q-learning to construct a parallel deep reinforcement learning algorithm for service caching decision;

[0118] During the training phase, the training unit selects an action to execute according to the current state using a greedy strategy, obtains a reward and the next state, and stores the obtained state transition in the experience pool; when the capacity of the experience pool D is large enough, a certain number of state transitions are extracted from the experience pool to train the network parameters;

[0119] During the decision-making phase, the decision-making unit generates multiple caching decisions through multiple parallel deep neural networks and stores them in the action set, calculates the corresponding rewards obtained for each caching decision, and uses the caching decision with the maximum reward as the output action.

[0120] Furthermore, the algorithm construction unit describes the service optimization caching problem as a Markov decision process; the Markov decision process consists of three parts: a state space S, an action space A, and a reward function R.

[0121] Furthermore, in the parallel deep reinforcement learning algorithm constructed by the algorithm construction unit, the action execution of the deep neural network DNN is executed in parallel. The deep neural network DNN includes two neural network structures with the same structure but different parameters; one is the main neural network for predicting the Q estimate value of the Q-learning reinforcement learning algorithm, which has the latest network parameters; the other is the target neural network for predicting the actual Q value of the Q-learning reinforcement learning algorithm, and the parameters used are the parameters from some time ago and remain unchanged for some time.

[0122] The present invention calculates the optimal solution for the mathematical model of the edge caching policy based on the combination of multiple parallel deep learning networks and reinforcement learning algorithms, and can solve the technical problem of inaccurate prediction of caching policies in the prior art. These machine learning combined with deep learning algorithms, based on a large amount of user historical data, enable the machine to learn and predict the user's preference degree and the changing trend of content popularity in the network, and adjust the service caching policy according to the learning results.

[0123] Although quite a lot of research has been carried out on the edge computing service caching problem at present, there is less research on using the deep reinforcement learning method for edge service caching, and even less application to the industrial Internet field. In addition, compared with the traditional deep reinforcement learning method, the present invention designs a distributed method, uses multiple parallel DNNs for service caching decision-making, and is superior in terms of minimizing service access latency and energy consumption performance.

[0124] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An industrial Internet of Things (IIoT) edge service caching decision-making method, characterized in that, it includes the following steps: S1. Based on the fact that only when the corresponding service data is cached in the server can the tasks corresponding to the service be executed, a mathematical model of the industrial Internet system is established; all the data required for all services is cached in the cloud server of the system model; S2. A mathematical model of the service access delay in the edge-cloud collaboration system is established; S3. According to the power of data transmission between edge servers and between edge servers and the cloud server, and the computing power of edge servers and the cloud server, a mathematical model of the energy consumption of the industrial Internet system is established; S4. Based on the system model, delay model, and energy consumption model, an optimization goal of minimizing service access delay and minimizing energy consumption is established; wherein, the optimal service caching policy is solved based on the optimization goal to achieve the purpose of minimizing service access delay and energy consumption, where β represents the weight of energy consumption; the optimization goal is as follows: Discretize time into evenly distributed time slices The interval of each time slice is t; consider that the industrial Internet system consists of N edge servers, denoted as The computing power and storage capacity of edge server n are respectively denoted as and In the industrial Internet system, the generated service set is denoted as Different services have different data volumes and require different computing resources and storage resources to process, denoted by f l and m l respectively, where The binary variable x l,n (t) ∈ {0, 1} indicates whether the service is cached on the edge server. The service caching policy is If service l is cached on edge server n at time t, then x l,n (t) = 1; otherwise, x l,n (t) = 0; p l,n represents the computing power allocated by edge server n to service l; μ l,n (t) ∈ [0, 1] represents the load ratio of service l executed on edge server n; S5. Based on the distributed deep reinforcement learning method, an algorithm capable of achieving the optimization goal is constructed.

2. The method according to claim 1, characterized in that, step S5 further includes the following steps: S51. Combine multiple parallel deep neural networks (DNNs) with the Q-learning reinforcement learning algorithm to construct a parallel deep reinforcement learning algorithm for service caching decision-making; S52. In the training stage, select an action to execute according to the current state with a greedy policy to obtain a reward and the next state, and store the obtained state transition in the experience pool; when the capacity of the experience pool D stored is large enough, extract a certain number of state transitions from the experience pool to train the network parameters; S53. In the decision-making stage, generate multiple caching decisions through multiple parallel deep neural networks and store them in the action set, calculate the corresponding rewards obtained by each caching decision, and use the caching decision with the maximum reward as the output action.

3. The method according to claim 2, characterized in that, in step S1, the edge server provides data analysis and processing services for sensor devices and industrial devices; compared with the cloud server, the computing resources and storage resources of the edge server are limited.

4. The method according to claim 3, characterized in that, in step S1, when the requested service is not cached on the edge server closest to it, the service is executed on the cloud server or another edge server that has cached the service.

5. The method according to claim 4, characterized in that, in step S5, the service optimal caching problem is described as a Markov decision process; the Markov decision process consists of three parts: a state space S, an action space A, and a reward function R.

6. The method according to claim 5, characterized in that, In the step S51, the action execution of the deep neural network DNN is performed in parallel. The deep neural network DNN includes two neural network structures with the same structure but different parameters. One is the main neural network for predicting the Q-estimation value of the Q-learning algorithm in the reinforcement learning algorithm, which has the latest network parameters. The other is the target neural network for predicting the actual Q-value of the Q-learning algorithm in the reinforcement learning algorithm, and the parameters used are the parameters from some time ago and remain unchanged for a certain period of time.

7. An industrial Internet edge service caching decision system adopting the method according to any one of claims 1-6, comprising: a mathematical modeling module and a service caching decision module; characterized in that, the mathematical modeling module performs mathematical modeling on the industrial Internet system based on the fact that only when the corresponding service data is cached in the server can the tasks corresponding to the service be executed; all the data required for the services are cached in the cloud server of the system model; establish a mathematical model for the service access delay in the edge-cloud collaborative system; perform mathematical modeling on the energy consumption of the industrial Internet system according to the power of data transmission between edge servers and between edge servers and the cloud server, and the computing power of edge servers and the cloud server; based on the system model, the delay model and the energy consumption model, establish an optimization goal of minimizing the service access delay and minimizing the energy consumption; the service caching decision module constructs an algorithm capable of achieving the optimization goal based on the distributed deep reinforcement learning method.

8. The system according to claim 7, characterized in that, the service caching decision module further includes: an algorithm construction unit, a training unit and a decision unit; the algorithm construction unit combines multiple parallel deep neural networks DNN with the Q-learning algorithm in the reinforcement learning algorithm to construct a parallel deep reinforcement learning algorithm for service caching decision-making; the training unit selects an action to execute according to the greedy strategy based on the current state in the training stage, obtains the reward and the next state, and stores the obtained state transition in the experience pool; when the capacity of the experience pool D stored is large enough, a certain number of state transitions are extracted from the experience pool to train the network parameters; the decision unit generates multiple caching decisions through multiple parallel deep neural networks and stores them in the action set in the decision-making stage, calculates the corresponding rewards obtained by each caching decision, and uses the caching decision with the maximum reward as the output action.

9. The system according to claim 8, characterized in that, the algorithm construction unit describes the service optimal caching problem as a Markov decision process; the Markov decision process consists of three parts: a state space S, an action space A and a reward function R.

10. The system according to claim 9, characterized in that, In the parallel deep reinforcement learning algorithm constructed by the algorithm construction unit, the action execution of the deep neural network DNN is executed in parallel. The deep neural network DNN includes two neural network structures with the same structure but different parameters. One is the main neural network for predicting the Q estimate value of the Q-learning reinforcement learning algorithm, which has the latest network parameters. The other is the target neural network for predicting the actual Q value of the Q-learning reinforcement learning algorithm, which uses parameters from some time ago and remains unchanged for some time.

Citation Information

Patent Citations

  • Internet of vehicles edge caching method based on multi-agent deep reinforcement learning

    CN113094982A

  • Video cache updating method for adaptive code rate selection in mobile edge computing

    CN113114756A