Dependent task unloading method based on reliability perception of topology reconstruction in industrial internet edge computing

By building a system model in industrial Internet edge computing and combining Ford-Fulkerson approximation algorithm and deep Q networks to optimize microservice deployment and computing offloading strategies, the reliability problems under dynamic network topology and complex tasks are solved, and the efficient reliability and resource utilization of the system are optimized.

CN120353590APending Publication Date: 2025-07-22HUBEI UNIV OF ARTS & SCI
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510439184.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When facing dynamic network topology and complex task dependence, it is difficult to achieve high reliability and optimization of system utilization. Especially under dynamic network topology and complex task dependence, the existing industrial Internet edge computing methods fail to effectively integrate the efficient path planning capabilities of network flow algorithms, resulting in a refined trade-off between communication costs and computing delays.

Method used

Build a system model covering edge cloud networks, industrial cloud platforms and Internet of Things devices, combine the Ford-Fulkerson approximation algorithm and deep Q network, and reduce communication costs and improve system reliability through microservice topology reconstruction and dynamic computing offload optimization.

Benefits of technology

It has achieved high reliability and optimization of resource utilization for IoT applications in dynamic industrial scenarios, significantly reducing communication costs and improving system reliability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353590A_ABST
    Figure CN120353590A_ABST
Patent Text Reader

Abstract

The invention provides a reliability-aware dependent task unloading method based on topology reconstruction in industrial internet edge computing, which comprises the following steps of: constructing a system model covering an edge cloud network platform, an industrial cloud platform and internet of things equipment, and establishing a transmission delay model and a reliability model; constructing a task unloading mathematical model with the maximum reliability level; the method comprises the following steps: modeling a micro-service dependency relationship into a weighted directed acyclic graph based on a network flow theory, and carrying out topology reconstruction on micro-services applied to the Internet of Things through a Ford-Fulkerson approximation algorithm to obtain a micro-service grouping structure formed by minimum cut division; the micro-service grouping structure serves as priori knowledge to be input into the deep Q network, the deep Q network is used for solving a task unloading mathematical model, a task unloading strategy is dynamically adjusted, resource allocation is optimized, and an optimal calculation unloading scheme in the industrial internet edge calculation environment is obtained. The micro-service deployment is optimized, the communication overhead is reduced, and the system reliability and the resource utilization rate are improved through real-time network state dynamic decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial Internet edge computing, and particularly to a reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing. Background Technique

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, the industrial Internet has become the core technology to promote the transformation of industrial production towards intelligence, automation, and high efficiency. The industrial Internet connects a large number of sensors, devices, and systems to the network, enabling real-time data collection, transmission, and analysis, thereby providing unprecedented flexibility and efficiency for industrial production. According to relevant market research reports, the global industrial Internet market size is expected to continue to grow in the next few years, showing great potential in improving production efficiency, optimizing resource allocation, and enhancing product quality.

[0003] However, with the continuous expansion of industrial production scale and the increasing complexity of production processes, the industrial Internet faces huge data processing and real-time challenges. The high communication costs and task latency problems caused by dynamic network topologies and complex task dependencies in industrial production seriously threaten the reliability of the system. Traditional cloud computing models are difficult to meet the strict requirements for real-time and reliability in industrial scenarios due to data transmission delays and the load limitations of central cloud servers. For example, in an intelligent manufacturing environment, the real-time monitoring and rapid response of production tasks are crucial for ensuring production quality and efficiency, while the data transmission delay in the cloud computing model may lead to decision-making delays in the production process, thus affecting production efficiency and product quality.

[0004] In this context, edge computing, as an emerging distributed computing paradigm, has gradually become an effective means to solve the real-time problems in the industrial Internet. Edge computing deploys computing resources on the edge side of the network, transfers data processing and analysis tasks from the cloud to edge nodes close to the data source, thereby significantly reducing data transmission delays and improving the system's response speed. In addition, edge computing can also reduce the dependence on central cloud servers, reduce network bandwidth consumption, and improve the reliability and stability of the system.

[0005] Although edge computing has shown great advantages in the industrial Internet, it still faces many challenges in practical applications. First, the complexity of Internet of Things (IoT) applications and the diversity of tasks make task offloading and resource allocation more complex. IoT applications usually consist of multiple interdependent microservices, and the processing time and communication cost of these microservices directly affect the completion time and reliability of the entire application. Second, factors such as dynamic changes in network topology, node failures, and uneven loads in the industrial Internet environment also pose severe challenges to the reliability of the system. The existence of these problems limits the widespread application of the industrial Internet in complex industrial scenarios and highlights the timeliness and importance of research.

[0006] Previous studies have mostly focused on how to reduce task offloading time or energy consumption, but have paid insufficient attention to the overall reliability of the system. In addition, existing methods often struggle to achieve global optimization when faced with dynamic network topologies and complex task dependencies.

[0007] With the rapid development of edge computing technology in the industrial Internet, the importance of computing offloading and resource management in industrial scenarios has become increasingly prominent. The complexity of the industrial Internet and its high requirements for real-time performance and reliability have made the optimization of computing offloading strategies gradually become a hot issue of common concern in academia and industry. Existing research mainly focuses on aspects such as task offloading efficiency, resource consumption, and reliability level, and has proposed diverse solutions in different application scenarios to meet the high requirements of the industrial Internet for real-time performance, reliability, and resource utilization.

[0008] For example, Deng et al. proposed two reinforcement learning-based offloading strategies and an autonomous partial offloading system for latency-sensitive computing tasks in a multi-user industrial Internet edge computing system, achieving smaller latency and quality of service (QoS) and optimizing task offloading time. Dai et al. proposed a co-offloading framework and studied a learning-based task co-offloading algorithm, enabling IoT devices to observe and understand system costs from candidate edge nodes and thus select the best edge node. You et al. proposed a task offloading strategy based on particle swarm optimization, reducing the latency and energy consumption of edge servers. Wang et al. designed an offloading algorithm based on an improved simulated annealing particle swarm algorithm to reduce costs and data transmission latency. However, the above research work mainly focuses on how to achieve a dynamic and effective task offloading mechanism in the industrial Internet edge computing environment, without considering the key factor of the reliability level of IoT applications.

[0009] For this reason, Zhao et al. proposed a distributed redundant scheduling algorithm to solve the availability problem of microservice applications caused by container failures. Long et al. proposed a semi-online learning service offloading strategy to optimize service offloading efficiency, energy consumption, and system reliability, and used an adaptive checkpoint mechanism to improve task reliability. Liang et al. studied reliability-aware data compression and task offloading for data-intensive applications in mobile edge computing, transformed the original problem into a convex optimization problem, and proposed a distributed reliability-aware task processing and offloading algorithm to minimize latency while satisfying reliability and energy consumption constraints. Although the above literature has improved the reliability level of IoT applications in different ways, the current research has not yet studied how to improve the reliability of IoT applications from the perspective of the completion time of IoT applications.

[0010] For this reason, Liu et al. proposed two reliability-enhanced task offloading methods, considering both the bandwidth consumption of IoT applications and the completion time of IoT applications to maximize the reliability level of these IoT applications. Liu et al. proposed an efficient task scheduling algorithm to minimize the average completion time of multiple applications, which guarantees the completion time constraints of applications and the processing dependency requirements of tasks by prioritizing multiple applications and multiple tasks. Goudarzi et al. proposed a parallel IoT batch application layout decision method based on a memetic algorithm to minimize the execution time and energy consumption of IoT applications. Ming et al. proposed a new hybrid task offloading scheme to jointly offload a series of dependent tasks to a mobile edge network supporting device-to-device communication, thereby minimizing the maximum latency for processing a series of dependent tasks in the mobile edge network. Taka et al. proposed a service placement and user allocation model using preventive startup time optimization in a mobile edge computing network to cope with single base station failures, in order to minimize the maximum penalty in all failure modes.

[0011] In addition, Hammadi et al. proposed four scheduling algorithms to reduce the total execution time, thereby enhancing the performance in meeting the deadlines of emergency tasks. Yan et al. designed a joint trajectory control and offloading allocation algorithm based on deep reinforcement learning to solve the computational offloading problem and reduce the system cost. Li et al. proposed a computational offloading scheme based on planned multi-agent deep reinforcement learning to make the most effective decisions under resource constraints. Yu et al. proposed a multi-agent deep reinforcement learning algorithm based on QMIX, effectively achieving cooperative offloading for aerial edge computing. Wang et al. developed a two-stage computational offloading mechanism, effectively improving the computational offloading performance. Naouri et al. proposed a three-layer task offloading framework and introduced a greedy task graph partitioning offloading algorithm, thereby minimizing the task communication cost. Ning et al. proposed two learning algorithms to achieve a Nash equilibrium with polynomial-time computational complexity, realizing efficient computational offloading and appropriate server deployment for multiple users and edge servers in a dynamic environment.

[0012] Current research has achieved remarkable results in the field of task offloading and reliability optimization for industrial Internet edge computing, but there are still multi-dimensional optimization bottlenecks. Although existing dynamic offloading strategies (such as reinforcement learning and swarm intelligence algorithms) improve the task execution efficiency through environmental perception and adaptive decision-making, they do not incorporate the dynamic network topology evolution, complex task dependency chains, and core reliability indicators into a unified optimization framework, resulting in the difficulty of stably guaranteeing system reliability in high-dynamic industrial scenarios. Research on reliability enhances the system robustness through mechanisms such as redundant scheduling and checkpoint fault tolerance, but its static resource allocation assumption and single-point failure model are difficult to adapt to the continuously fluctuating node loads and network topology reconstruction requirements in the industrial Internet. At the task topology optimization level, although existing microservice splitting methods (such as graph theory decomposition and priority sorting) reduce the basic communication cost, their fixed splitting rules are decoupled from the real-time resource status of edge nodes and cannot achieve a globally optimal microservice deployment under dynamic loads. At the same time, although the offloading scheme based on deep reinforcement learning shows the advantage of environmental adaptability, its action space design often ignores the dynamic constraints of the dependencies between microservices on topology reconstruction and does not effectively integrate the efficient path planning ability of network flow algorithms, resulting in the difficulty of achieving a refined trade-off between communication cost and computational latency. Most of these studies do not consider the impact of dynamic network topology and complex task dependencies on system reliability. Summary of the Invention

[0013] Aiming at the computing offloading challenges brought about by the dynamic network topology, complex task dependencies, and high reliability requirements in industrial Internet edge computing, the present invention proposes a reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing, focusing on the reliability computing offloading problem in the industrial Internet edge computing environment, and aiming to improve the reliability and performance of the system by optimizing the microservice deployment and computing offloading strategy of Internet of Things applications. Compared with previous studies, the present invention integrates the reliability optimization method of Internet of Things application topology reconstruction and deep reinforcement learning, not only paying attention to the efficiency of task offloading, but also particularly emphasizing the reliability and resource utilization rate of the system.

[0014] To achieve the above objectives, the technical solution of the present invention is realized as follows: A reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing, the steps of which are as follows:

[0015] Step 1: Construct a system model covering the edge cloud network platform, industrial cloud platform, and Internet of Things devices, establish a transmission delay model and a reliability model, and construct a task offloading mathematical model that maximizes the reliability level;

[0016] Step 2: Based on the network flow theory, model the microservice dependency relationship as a weighted directed acyclic graph, perform topology reconstruction on the microservices of the Internet of Things application through the Ford-Fulkerson approximation algorithm, obtain the minimum cut partition to form a tightly connected microservice grouping structure, so as to reduce unnecessary communication duration; The microservice grouping structure after topology reconstruction is used as prior knowledge to input into the deep Q network, and the deep Q network is used to solve the task offloading mathematical model in Step 1, dynamically adjust the task offloading strategy, optimize resource allocation, and obtain the optimal computing offloading scheme in the industrial Internet edge computing environment.

[0017] Preferably, the edge cloud network platform is respectively connected to the industrial cloud platform and the Internet of Things devices; the Internet of Things devices serve as the terminal layer and establish a connection with the edge access platform through wireless communication technology, and use lightweight protocols to transmit the collected industrial field data; the edge access platform serves as the gateway between the Internet of Things devices and the edge cloud network platform. After authenticating the identity, protocol conversion, and data preprocessing of the connected Internet of Things devices, it distributes the data to the edge cloud network platform through the edge network; the edge cloud network platform is composed of a distributed edge server cluster and is interconnected through an intranet; each edge server runs multiple containers, and the containers communicate through virtual network interfaces; the Internet of Things applications are split into multiple microservices. After receiving the data stream from the access platform, the edge cloud network platform dynamically schedules the computing tasks to the optimal container according to the microservice topology relationship, and the processing results are uploaded to the industrial cloud platform through the edge-cloud collaboration channel; the industrial cloud platform is interconnected with multiple edge cloud network platforms through dedicated lines or VPNs to achieve global resource orchestration and big data analysis; after receiving the edge cloud aggregated data, the edge cloud network platform performs complex calculations, and the generated optimization model / strategy is sent to the edge server through the reverse channel to update the local processing logic, forming a closed-loop data stream of "device-edge-cloud".

[0018] The system model randomly pre-deploys P heterogeneous edge servers into H edge cloud network platforms of the industrial Internet edge computing system; M containers indexed by the container set Z = {1, 2,..., M} are randomly assigned to H edge servers. Each Internet of Things device generates an Internet of Things application at certain times, with a total of I1 Internet of Things applications indexed by the set I = {1, 2,..., i,..., I1}. Each Internet of Things application contains multiple mutually dependent microservices. In the Internet of Things application, a series of microservices with dependencies are established as a directed acyclic graph G = (V, E), where V and E represent the node set and the edge set respectively; each microservice is indivisible and can only be processed by container j, and container j has a certain processing capacity C j ; in the established directed acyclic graph, each node represents a microservice, the weight of each node represents the average processing time of the microservice, and each edge represents the communication time between microservices;

[0019] The reliability of an Internet of Things application is determined by the path with the maximum sum of the average processing time, processing delay, and communication time between microservices from the start node to the end node.

[0020] Preferably, the task offloading mathematical model is: the modeling function is to maximize the reliability level of multiple Internet of Things applications:

[0021]

[0022] The constraints are as follows: 1) Each task can only be deployed to one container:

[0023]

[0024] 2) The memory capacity of each container is greater than or equal to the total memory required by the deployed tasks:

[0025]

[0026] Among them, x ixr represents whether the x-th microservice of the Internet of Things application i is deployed to the r-th container; l ix represents the memory capacity of the x-th microservice of the Internet of Things application i, and L r represents the memory capacity of the r-th container, and X i represents the microservice set of the Internet of Things application i, and |X i | represents the number of all microservices; I represents the Internet of Things application set containing I1 Internet of Things applications, and F(T i -D i ) is the difference function of the probability of violating QoS, T i is the completion time of the Internet of Things application i, D i is the deadline, and Z is the container set containing M containers.

[0027] Preferably, the difference function of the probability of violating QoS is

[0028]

[0029] i -D i ) represents whether the QoS requirement of the Internet of Things application i is violated;

[0030] The completion time T i of the Internet of Things application i is:

[0031]

[0032] Among them, represents the transmission delay for the Internet of Things device that generates the Internet of Things application i to transmit data d i to the edge cloud h, represents the link data transmission rate between the Internet of Things device that generates the Internet of Things application i and the edge cloud h;

[0033] T ixj is the total processing time of the x-th microservice of the Internet of Things application i on the j-th container, which consists of the microservice processing time, and:

[0034]

[0035] Among them, t ixj is the processing time of microservice x of the Internet of Things application i, and Among them, L ix is to divide the Internet of Things application i into |X i | microservices, and the instruction length of microservice x ∈ X i , and C j represents the processing capacity of container j;

[0036] The processing delay d of the microservice ixj ~Uniform(2, 6), where Uniform represents the uniform distribution function;

[0037] T ixj is the communication time between microservice x and microservice y, and

[0038]

[0039] Among them, α ixy represents whether microservice x and microservice y are on the same edge server. If so, the value of α ixy is 0, otherwise it is 1; β ixy represents whether microservice x sends data to microservice y. If so, the value of β ixy is 1, otherwise it is 0; in xy and out yx respectively represent the output data size from microservice x to microservice y and the output data size from microservice y to microservice x; B upload and B download respectively represent the network upload bandwidth and download bandwidth.

[0040] Preferably, the Ford-Fulkerson approximation algorithm performs minimum cut partitioning on the weighted directed acyclic graph, aggregates communication-intensive microservices into homogeneous sub-task groups, and by introducing virtual source nodes and sink nodes, transforms the topology reconstruction problem into a maximum flow optimization problem, and realizes the minimum cut solution by iteratively finding the augmenting path; the deep Q network drives the agent to learn the optimal offloading strategy in a dynamic environment by designing a composite reward function that includes instant communication cost, long-term reliability reward, and resource utilization penalty, and realizes the closed-loop feedback of static topology optimization and dynamic resource scheduling by using the experience replay mechanism. It can not only maintain the efficient communication mode between microservices, but also adapt to the real-time load fluctuations of edge nodes. Finally, on the premise of meeting the maximized reliability level, it realizes the multi-objective balance of computing offloading efficiency and system stability, and outputs the optimal offloading strategy π*, the converged Q value function Q*, and the results of the performance index set O, which together constitute the optimal computing offloading solution in the industrial Internet edge computing environment.

[0041] Preferably, each node in the weighted directed acyclic graph G=(V,E) represents a microservice, and the value on the connection line between adjacent microservices represents the communication time between microservices, denoted by the edge weight ω(e(v,u)), where e(v,u) represents the edge connecting nodes v and u; the Ford-Fulkerson approximation algorithm optimizes the splitting of microservices by finding an augmenting path. Each time, it finds an augmenting path with the maximum residual capacity from the virtual source node to the sink node, and increases the flow through the augmenting path, reducing the residual capacity of each edge on the augmenting path; this process is repeated until no more augmenting paths can be found, obtaining the maximum flow, which represents the optimal microservice splitting scheme;

[0042] The implementation method of the Ford-Fulkerson approximation algorithm is as follows: Based on the directed acyclic graph G, retain the dependency edges between the original microservices and directly use their edge weights ω(e(v,u)) as capacities to construct an auxiliary flow network G' that includes a virtual source node s and a sink node t. The source node s is the entry of the microservice task flow, and the sink node t is the exit of the microservice task flow. The source node s is connected to all nodes, and the edge capacity is the sum of the communication costs of the outgoing edges of each node. The sink node t is connected to all microservice nodes, and the edge capacity is the sum of the communication costs of the incoming edges of each node;

[0043] Calculate the edge capacity: The edge capacity from the source node s to each node v is initially set to the sum of the communication costs of all outgoing edges, and the edge capacity from the node to the sink node t is initially set to the sum of the communication costs of all incoming edges;

[0044] By iteratively searching for the s-t augmenting path in the auxiliary flow network G', calculate the minimum flow of the augmenting path and update the residual capacity to dynamically divide the communication-intensive microservice groups.

[0045] Preferably, the method for dynamically dividing the communication-intensive microservice groups is as follows: Based on breadth-first search, find a feasible path path from the virtual source node s to the sink node t in the auxiliary flow network G'. The generation condition of the feasible path path requires that the residual capacity cap(u,v) of all constituent edges (u,v)>0; for each discovered feasible path path, calculate the bottleneck flow min_flow = min{cap(u,v)|(u,v)∈path}; in the flow pushing stage, perform a capacity reduction operation of cap(u,v)-min_flow on all forward edges along the feasible path path, and at the same time perform a capacity compensation of cap(v,u)+min_flow on the reverse edges; in the implementation of microservice grouping, after each flow push, recalculate the reachable node set reachable starting from the source node s through breadth-first search, and the intersection of the reachable nodes and the remaining nodes remaining_nodes gives the group S kForm the currently communication-intensive microservice groups until the maximum number of groups K is reached or all nodes are allocated; the unallocated nodes are incorporated into the result set S as independent groups, and mutual exclusivity is ensured by merging overlapping groups;

[0046] When the nodes remain connected in the auxiliary flow network G', there is still remaining capacity in the communication links between the nodes, so they are divided into the same group S k ; conversely, if they are disconnected due to traffic saturation, they will be assigned to different groups; among them,

[0047] The method for merging and optimizing the overlapping groups to form the final group structure is as follows: detect all groups in the result set S that have common nodes, take the union of the groups with common nodes to form new groups, and remove the atomic groups and retain the merged new groups;

[0048] The finally output result set S satisfies: 1) The communication intensity between nodes within the group is higher than that between groups; 2) All groups do not overlap; 3) The total number of groups ≤ K1, where K1 ≤ P / 2 and P is the number of edge servers.

[0049] Preferably, the agent in the deep Q-network observes the system state s at each decision-making moment t t , and the state s t includes the container load L r , the channel gain the edges between microservices, and generates a q-dimensional offloading action vector At = [a1, a2,..., aq] through the deep Q-network, where each action component corresponds to a deployment decision of a microservice group between edge nodes, that is, an offloading strategy;

[0050] Based on the ε-greedy strategy, independently select the target node with the largest Q value for each microservice group, and keep the output dimension constantly q through the virtual filling or group merging mechanism; drive the update of network parameters through the designed composite reward function;

[0051] The obtained interaction data <s, a, r, s'> is stored in the experience replay pool. When there is enough data in the experience replay pool, randomly take out a batch_size-sized data from the experience pool and use the prediction network to calculate the predicted value of Q: Encode the state s t , and the prediction network maps the encoded state to a q×|Z|-dimensional Q-value matrix through a three-layer neural network to predict the long-term benefits of deploying each microservice group to different edge nodes; among them, s and a respectively represent the current state and action, the action a is the microservice deployment decision generated based on the current state, and r and s' respectively represent the reward value of the current action and the new state at the next moment;

[0052] Use the Q-Target network to calculate the Q target value, based on the next state s t+1 and the reward rt , the target network parameters θ updated by delay are used for forward propagation to obtain maxQ(s t+1 , a|θ-), and then the final target value is calculated in combination with the Bellman equation; the loss function between the target value and the predicted value is calculated, and the gradient descent is used to update the current network parameters, the gradient of the loss function with respect to the current network parameters is calculated, and the network weights are adjusted by backpropagation according to the set learning rate, so that the predicted value gradually approaches the target value, thereby optimizing the strategy; after repeating several times, the parameters of the prediction network are copied to the Q-Target network.

[0053] Preferably, the Q-Target network is denoted as represents the neural network for generating the target value; the prediction network is denoted as Q(s, a|θ), which represents the neural network for predicting the state-action value function and:

[0054] Q q+1 (s t , a t |θ) = Q q (s t , a t |θ) + a q E q ;

[0055]

[0056] Among them, a q and γ respectively represent the learning rate and the discount factor, s' is the predicted value of the state after performing the action a t in the q-th iteration, a' is the action with the largest reward value under the state s', E q is the cumulative reward value in the q-th iteration during the iteration process, a t is the action selected by the agent at time t, r t is the immediate reward obtained after performing the action a t at time t, θ, respectively represent the weight parameters of the prediction network and the weight parameters of the Q-Target network;

[0057] The agent randomly extracts a small part of the experience e t = {e1, e2,..., e t} from the experience replay pool D t = (s t , a t , r t , s') for calculating the target value When reaching the q-th time, the network parameters θ are updated by minimizing the composite reward function to realize the update of the offloading decision; the composite reward function is:

[0058]

[0059] Preferably, the implementation method of the deep Q-network is as follows: create an experience replay pool for storing the experience data generated during the interaction between the agent and the environment, including the current state, the actions taken, the rewards obtained, and the next state, and at the same time set the maximum capacity of the experience replay pool to K; construct two neural networks - a prediction network and a Q-Target network. Initially, copy the weights of the prediction network to the Q-Target network so that they have the same initial parameters; use M rounds of iterative training to optimize the dynamic offloading strategy, combine the ε-greedy strategy with the microservice topology reconstruction result, and generate the container deployment action a t : According to the ε-greedy strategy, the agent randomly selects an action with probability ∈ and selects the optimal action in the current state with probability 1 - ∈. Combine the microservice grouping with the resource status of the edge node or the cloud using the result of the microservice topology reconstruction to generate the specific container deployment action a t When the agent executes the action a t in the environment, observe the reward r t given by the environment and the new state s t+1 , and store the quadruple (φ t , a t , r t , φ t+1 ) into the experience replay pool, where φ t is the preprocessed representation of the current state s t , and φ t+1 is the preprocessed representation of the new state s t+1 ; update the Q-network parameter θ based on the temporal difference error, and minimize the weighted loss of the communication cost T ixy and the processing delay T ixj ; by randomly sampling a batch of experience samples from the experience replay pool, calculate the target value of each sample, calculate the difference between the predicted value of the prediction network and the target value as the temporal difference error, and use the temporal difference error to update the parameters of the Q-network through the gradient descent method; periodically synchronize the Q-Target network parameter θ- by copying the parameters of the prediction network to the parameters of the Q-Target network; output the optimal offloading strategy, the converged action value function, and the specific numerical values of the obtained performance metrics; determine the optimal offloading strategy by evaluating the state-action value function of the prediction network after the training is completed;

[0060] Traverse all possible states s and actions a, and select the action that maximizes the value of the state-action value function as the optimal offloading strategy π; output the converged state-action value function as the action value function Q

[0061] Compared with the prior art, the beneficial effects of the present invention are as follows: First, a system model covering the edge cloud network, industrial cloud platform, and Internet of Things devices is constructed, and a mathematical framework including transmission delay and reliability models is established, with the Internet of Things application completion time as the core reliability index. Second, the Ford-Fulkerson approximation algorithm is innovatively combined with the deep Q-network. Through the collaborative optimization of microservice topology reconstruction and dynamic computing offloading, the communication cost is reduced and the system reliability is improved, solving the limitations of existing methods in dynamic industrial scenarios. The Ford-Fulkerson approximation algorithm realizes the minimum-cost reconstruction of the microservice topology through the maximum flow-minimum cut principle, and the deep Q-network dynamically generates offloading strategies based on real-time network status, resource load, and task dependency graphs, thus achieving a double breakthrough in reliability improvement and resource efficiency optimization with the application completion time as the core in a complex industrial environment, providing a system-level solution for the computing offloading problem in dynamic topology and multi-constraint scenarios. Experimental results show that the present invention is significantly superior to existing methods in terms of Internet of Things application completion time and reliability level, providing a new technical path for the efficient and reliable operation of industrial Internet edge computing. The main contributions of the present invention are as follows:

[0062] (1) An industrial Internet edge computing system model covering key components such as edge cloud networks, industrial cloud platforms, Internet of Things applications, and Internet of Things devices is constructed, providing a comprehensive framework foundation for research.

[0063] (2) A mathematical model including a transmission delay model and a reliability model is established, with the Internet of Things application completion time as the core reliability index.

[0064] (3) A joint offloading method (FFAA-DQN) based on FFAA and DQN is proposed to optimize microservice deployment to reduce communication overhead, and improve system reliability and resource utilization through dynamic decision-making based on real-time network status.

[0065] (4) Through the construction of a simulation environment, a comprehensive performance evaluation of the method proposed by the present invention is carried out, verifying its superiority in terms of reliability level, Internet of Things application completion time, etc., and proving that the present invention can effectively adapt to complex industrial Internet environments. Description of the Drawings

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0067] Figure 1It is a schematic block diagram of the system model of industrial Internet edge computing of the present invention.

[0068] Figure 2 It is a schematic diagram of a directed acyclic graph of an Internet of Things application of the present invention.

[0069] Figure 3 It is an example diagram of topological reconstruction of an Internet of Things application of the present invention.

[0070] Figure 4 It is a framework diagram of dynamic offloading of Internet of Things applications based on DQN of the present invention.

[0071] Figure 5 It is a curve comparison diagram of the reliability levels obtained by the present invention and three other methods under different numbers of Internet of Things applications.

[0072] Figure 6 It is a curve comparison diagram of the Internet of Things application completion times obtained by the present invention and three other methods under different numbers of Internet of Things applications.

[0073] Figure 7 It is a curve comparison diagram of the average completion times of the same Internet of Things application before and after topological reconstruction under four methods of the present invention. Detailed implementation manners

[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0075] As Figure 1 shown, in response to the reliability challenges brought by dynamic network topologies and complex task dependencies in industrial Internet edge computing, the present invention proposes a reliability-aware dependent task offloading method based on topological reconstruction in industrial Internet edge computing, constructs a system model including an edge cloud network, an industrial cloud platform, and Internet of Things devices, and simultaneously establishes a transmission delay and reliability model quantization index system, using the Internet of Things application completion time as the core reliability evaluation index; innovatively combines the Ford-Fulkerson approximation algorithm (FFAA) with the deep Q-network (DQN). FFAA optimizes the microservice topology splitting through maximum flow analysis to reduce communication costs, and DQN dynamically adjusts the task offloading strategy between edge nodes to achieve the collaborative optimization of resource utilization efficiency and real-time performance. The specific implementation method of the present invention is as follows:

[0076] Step 1: A system model covering the edge cloud network platform, industrial cloud platform, and Internet of Things (IoT) devices was constructed. A transmission delay model and a reliability model were established, and a mathematical model for task offloading to maximize the reliability level was constructed.

[0077] The present invention constructs a system model covering the edge cloud network platform, industrial cloud platform, and IoT devices. As Figure 1 shown, the edge cloud network platform is respectively connected to the industrial cloud platform and IoT devices. The system realizes a complete chain from data collection to intelligent analysis through the closed-loop data flow of device-edge-cloud. In the system model, IoT applications are split into multiple microservices, and the processing time and communication time of each microservice directly affect the completion time of the application. Based on this, the present invention establishes a transmission delay model and a reliability model, with the completion time of the IoT application as the core reliability index.

[0078] Figure 1 The industrial Internet edge computing system model shown is a comprehensive architecture that integrates the advantages of edge computing and cloud computing, aiming to meet the requirements for real-time performance, reliability, and security in the industrial Internet. This system model provides intelligent services on the network edge side close to the data source to achieve real-time analysis of massive industrial data and real-time control tasks of the system. The system architecture includes four parts: the edge cloud network platform, industrial cloud platform, IoT applications, and IoT devices, and is composed of key components such as IoT devices, edge cloud network platform, edge servers, containers, edge networks, edge data, industrial cloud platforms, and edge access platforms. As the terminal layer, IoT devices establish connections with the edge access platform through wireless communication technologies (such as 5G, Wi-Fi, LoRa) and use lightweight protocols such as MQTT / CoAP to transmit the collected industrial field data. As the gateway between IoT devices and the edge cloud network platform, the edge access platform authenticates the accessed IoT devices, performs protocol conversion and data preprocessing, and then distributes the data to the edge cloud network platform through the edge network (including wired optical fiber and wireless Mesh network). The edge cloud network platform is composed of a cluster of edge servers deployed distributively and interconnected through a high-speed intranet. Each edge server runs multiple containers (Docker / Kubernetes), and the containers communicate through virtual network interfaces. After receiving the data stream from the access platform, the edge cloud network platform dynamically schedules the computing tasks to the optimal container according to the microservice topology relationship, and the processing results can be uploaded to the industrial cloud platform through the edge-cloud collaboration channel.

[0079] As a central hub, the industrial cloud platform is interconnected with multiple edge cloud network platforms through dedicated lines or VPNs to achieve global resource orchestration and big data analysis. After receiving the edge cloud aggregated data, the edge cloud network platform performs complex calculations such as deep learning training and business decision-making. The generated optimized models / policies are sent to the edge servers through the reverse channel to update the local processing logic, forming a closed-loop data flow of "device-edge-cloud". The communication between all components uses TLS encryption, and the network topology is uniformly managed by the SDN controller to ensure the security and real-time nature of data transmission.

[0080] The system model of the present invention realizes efficient collaborative operation based on the double-layer architecture of the edge cloud network platform and the industrial cloud platform. The edge cloud network platform consists of nodes such as base stations, servers, and gateways deployed near Internet of Things devices, providing distributed computing and real-time data processing capabilities through containerization technology to significantly reduce data transmission latency. Containerization technology in industrial Internet edge computing realizes the lightweight virtualization deployment of microservices through Docker / Kubernetes, combined with edge-optimized scheduling strategies, multi-copy fault tolerance mechanisms, and mTLS secure communication, providing an elastic and scalable operating environment for topology reconstruction and computing offloading. The industrial cloud platform, as the central hub, is responsible for processing the aggregated data uploaded by the edge nodes, performing global resource scheduling, big data analysis, and long-term storage. Internet of Things devices, as data acquisition terminals, obtain raw data such as device status and environmental parameters in the industrial field in real time through devices such as sensors and controllers, and transmit the data to adjacent edge nodes (such as base stations or edge servers) through the edge network. Internet of Things applications, as the interaction interface between the system and the physical world, trigger real-time control instructions or analysis requirements based on device data. The edge cloud network locally processes the received data, quickly responds to low-latency requirements using containerization technology, and at the same time offloads high-computation load tasks (such as large-scale data analysis) to the industrial cloud platform through the edge-cloud collaboration mechanism. After the cloud completes complex operations such as deep learning model training and global resource scheduling, the optimization results are then fed back to the edge nodes for execution.

[0081] Through the closed-loop data flow of device-edge-cloud, this system model realizes the complete chain from data acquisition, real-time response to intelligent analysis while ensuring industrial reliability. While taking into account the high-reliability requirements of the industrial Internet, it makes full use of the advantages of edge computing to achieve efficient allocation and processing of tasks. Through key technologies such as edge-cloud collaboration, redundant design, and dynamic scheduling, it significantly improves the real-time nature, reliability, and stability of the industrial Internet of Things system, providing strong support for the intelligence and automation of the industrial site.

[0082] Figure 1Among them, P heterogeneous edge servers are randomly pre-deployed into H edge cloud network platforms of the industrial Internet edge computing system. Then, M containers indexed by the set Z = {1, 2, …, M} are randomly assigned to these edge servers. In the industrial Internet edge computing system of the present invention, a container is an independent microservice running unit encapsulated by lightweight virtualization technology, serving as the basic carrier for task execution, and is dynamically scheduled and deployed by the FFAA-DQN algorithm to achieve computing offloading optimization and reliability improvement. Each Internet of Things device generates an Internet of Things application at certain times, and there are a total of I1 Internet of Things applications indexed by the set I = {1, 2, …, i, …, I1}. Each Internet of Things application contains multiple microservices that are interdependent with each other. In an Internet of Things application, a series of microservices with dependency relationships can be established as a directed acyclic graph, where each microservice is indivisible and is only processed by container j. Container j has a certain processing capacity C j (unit: instructions per second). In the established directed acyclic graph, each node represents a microservice, and they all have corresponding weights. Depending on different research questions, the meaning represented by the weights will also be different. Since the present invention uses reliability as the performance improvement index, and the completion time of the Internet of Things application is a key basis for evaluating reliability, in the present invention, the weight of each node represents the average processing time of the microservice. Each edge represents the dependency relationship between microservices, and in the present invention, it represents the communication time between microservices. Taking Figure 2 as an example, name the starting node as Start, name the ending node as End, and the intermediate process nodes can be represented by X microservices indexed by the set T = {1, 2, 3, …, X}. The start of each microservice depends on whether all its preceding microservices have been executed. For example, only after microservice Start is executed can microservices T1, T2, and T3 proceed; only after microservices Start, T1, T4, and T5 are executed can microservice T10 proceed. In an Internet of Things application, its reliability is determined by the path with the maximum sum of the average processing time, processing delay, and communication time between microservices from the starting node to the ending node.

[0083] A summary of the key symbols that appear in the process of establishing the mathematical model of the present invention is shown in Table 1.

[0084] Table 1: Key Symbols and Their Descriptions

[0085]

[0086] In the industrial Internet of Things (IIoT) edge computing environment, transmission delay is one of the key factors affecting system performance. IoT devices communicate with edge nodes via wireless networks. The offloading and processing of tasks and data require passing through multiple network nodes, and transmission delay plays a crucial role in this process. Especially in a dynamically changing network topology, task allocation and offloading strategies need to consider delay optimization to meet real-time requirements. To achieve task offloading and resource optimization in such an environment, it is particularly important to establish an accurate transmission delay model, which can help analyze the time taken for data to be transmitted in the network and thus provide a basis for task scheduling and resource allocation.

[0087] IoT devices and the edge cloud network platform use orthogonal frequency division multiple access (OFDMA) technology to achieve wireless communication. For IoT application i with input data d i The transmission delay can be calculated by Equation (1):

[0088]

[0089] Where, represents the transmission delay of data d transmitted from the IoT device generating IoT application i to edge cloud h i The transmission delay, represents the link data transmission rate between the IoT device generating IoT application i and edge cloud h. The link data transmission rate can be calculated by Equation (2):

[0090]

[0091] Here, B represents the link bandwidth between the device generating IoT application i and edge cloud h, p i represents the transmission power of the device generating IoT i, represents the channel gain between the device generating IoT i and edge cloud h, and N0 represents the noise power.

[0092] By establishing the above transmission delay model, the transmission time of IoT applications in the IIoT edge computing environment can be effectively analyzed, providing a scientific basis for task offloading and resource allocation, thereby optimizing system performance and meeting real-time requirements.

[0093] In the industrial Internet of Things (IIoT) edge computing environment, the reliability of IoT applications is a key indicator of system performance. Since the task offloading and data transmission processes may be affected by factors such as network fluctuations, node failures, and uneven loads, ensuring the reliability of the system is particularly important. To address these challenges, a reliability model needs to be established to measure and optimize the stability and performance of the system in the face of various uncertainties, so as to ensure that IoT applications can be completed efficiently and reliably in complex environments. Since the completion time of IoT applications is directly related to the satisfaction of service quality, the real-time response ability of the system, the efficiency of task processing, and the overall stability and consistency of the system, and can also indirectly reflect the degree of bandwidth consumption, thus comprehensively reflecting the key aspects of system performance. Therefore, the reliability model of the present invention will be established based on the completion time of IoT applications as the key basis.

[0094] IoT applications are usually divided into multiple collaborating microservices, which are transferred to the edge cloud network platform and processed by containers. The processing time of each microservice is determined by its instruction length and processing capacity. Specifically, for the microservice x of IoT application i, its processing time can be calculated by Equation (3):

[0095]

[0096] Divide IoT application i into |X i | microservices, and the instruction length of microservice x ∈ X i is denoted as L ix . C j represents the processing capacity of container j (unit: instructions per second).

[0097] In actual operation, the processing of microservices is not only limited by processing capacity but also affected by network latency. Therefore, the processing delay d ixj of microservices also needs to be considered. The processing delay can be regarded as a random variable, and its value is randomly selected from a preset interval. As shown in Equation (4), in the present invention, the processing delay is randomly selected from the interval [2, 6] ms:

[0098] d ixj ~Uniform(2,6)(4)

[0099] where Uniform represents the uniform distribution function, and Uniform(2, 6) means that the microservice processing delay d ixj follows a uniform distribution between 2 ms and 6 ms.

[0100] The total processing time T ixj of the xth microservice of IoT application i on the jth container is composed of the microservice processing time t ixj , the microservice processing delay d ixjis composed, and its value can be calculated by Equation (5):

[0101]

[0102] where I represents the set of Internet of Things (IoT) applications, Z represents the set of containers, and X i represents the set of microservices of IoT application i.

[0103] In the industrial Internet of Things edge computing environment, the communication time between microservices is one of the key factors affecting system performance. When an IoT application is split into multiple microservices and deployed on different edge servers or containers, the communication overhead between microservices will significantly affect the execution efficiency and reliability of the entire application. Therefore, accurately calculating the communication time between microservices is crucial for optimizing task offloading strategies and system performance.

[0104] If multiple microservices of the same IoT application are deployed on different edge servers, the communication time between these microservices cannot be ignored. The communication time between microservice x and microservice y can be calculated by Equation (6):

[0105]

[0106] where α ixy is a binary variable indicating whether microservice x and microservice y are on the same edge server. If so, its value is 0; otherwise, it is 1. β ixy is a binary variable indicating whether microservice x sends data to microservice y. If so, its value is 1; otherwise, it is 0. in xy and out yx represent the output data size from microservice x to microservice y and the output data size from microservice y to microservice x, respectively. B upload and B download represent the network upload bandwidth and download bandwidth (unit: Mb / s), respectively.

[0107] The completion time T i of IoT application i is determined by the maximum value of the transmission delay and the sum of the total processing time and communication time of all its microservices (the transmission delay is not included in the maximum value). The directed acyclic graph of IoT application i consists of multiple execution routes (i.e., W), and the execution times of these routes are different. The execution route with the longest execution time determines the completion time of the IoT application, as shown in Equation (7). Given that the data volume of the IoT application result is small, the time required to transfer the data from the edge cloud back to the IoT device can be ignored.

[0108]

[0109] Among them, W is the set of execution routes from the top-level task Start to the bottom-level task End, such as Figure 2 , including {Start, T1, T4, T10, T14, End}, {Start, T1, T5, T10, T14, End}, {Start, T1, T5, T11, T14, End}, {Start, T2, T6, T12, T15, End}, {Start, T2, T7, T12, T15, End}, {Start, T3, T8, T13, End}, {Start, T3, T9, End}; v represents the task set of a certain execution route in the set W (such as {Start, T1, T4, T10, T14, End}). |v| is the number of nodes in the task set.

[0110] In the present invention, the reliability level of the Internet of Things application reflects the performance of the system in meeting the QoS requirements. If the completion time T i of the Internet of Things application i exceeds its deadline D i requirement, it is considered that the QoS is violated. Therefore, the probability of the reliability level can be calculated by Equations (8) and (9), that is, the difference between 1 and the probability of violating the QoS:

[0111]

[0112] where I represents the total number of Internet of Things applications. The difference (T i - D i ) indicates whether the QoS requirement of the Internet of Things application i is violated. If (Ti - Di) > 0, it means that the QoS requirement is violated, and F(T i - D i ) = 1. If (Ti - Di) ≤ 0, it means that the QoS requirement is not violated, then F(T i - D i ) = 0.

[0113] By establishing the above reliability model, the reliability of the Internet of Things application in the industrial Internet edge computing environment can be accurately measured, providing a scientific basis for task offloading and resource allocation, thereby optimizing the system performance and ensuring the high reliability and stability of the system.

[0114] In a reliability-based industrial Internet of Things (IIoT) edge computing system, edge cloud nodes are interconnected via fiber optic cables, and the communication between edge clouds is considered load-independent. Specifically, when the collaborative tasks of Internet of Things application i are distributed across different edge servers, the communication time between containers mainly depends on the bandwidth of the sending container and the amount of data. Although increasing the communication bandwidth between tasks can effectively reduce communication latency, since the bandwidth of edge servers is limited, this approach of increasing bandwidth may have a negative impact on other tasks. Specifically, if the bandwidth between the collaborative tasks of a certain Internet of Things application is increased on the same edge server, the bandwidth available for other tasks on that server will be reduced, thereby lowering the reliability level of the Internet of Things applications in which other tasks reside. Therefore, in the case of limited edge cloud resources, optimizing the task offloading strategy to maximize the reliability levels of multiple Internet of Things applications becomes a key issue, and this optimization process can be modeled by Equation (10).

[0115]

[0116] Equation (11) indicates that each task can only be deployed to one container; where, x ixr represents whether the x-th microservice of Internet of Things application i is deployed to the r-th container. Since the memory capacity of a container is in terms of the number of instructions, Equation (12) is used to represent that the memory capacity of each container is greater than or equal to the total memory required by the deployed tasks, where, l ix represents the memory capacity of the x-th microservice of Internet of Things application i, and L r represents the memory capacity of the r-th container. F(i) represents the problem modeling function of the present invention, that is, to maximize the reliability level; |X i | represents the number of all microservices.

[0117] Step 2: Based on the network flow theory, model the microservice dependencies as a weighted directed acyclic graph, perform topological reconstruction on the microservices of the Internet of Things application through the Ford-Fulkerson approximation algorithm, and obtain a tightly connected microservice grouping structure formed by the minimum cut partition to reduce unnecessary communication duration; the microservice grouping structure after topological reconstruction is used as prior knowledge to input into the deep Q-network, use the deep Q-network to solve the task offloading mathematical model in Step 1, dynamically adjust the task offloading strategy, optimize resource allocation, and obtain the optimal computing offloading solution in the industrial Internet of Things edge computing environment.

[0118] First, based on the network flow theory, the microservice dependency relationship is modeled as a weighted directed acyclic graph. Through the Ford-Fulkerson approximation algorithm, the minimum cut is partitioned, and communication-intensive microservices are aggregated into homogeneous subtask groups. This process, by introducing virtual source and sink nodes, transforms the topology reconstruction problem into a maximum flow optimization problem. The minimum cut solution is achieved by iteratively finding an augmenting path in lines 8 - 14 of Algorithm 1, effectively reducing the cross-node communication cost. Meanwhile, in Step 1, a mixed-integer programming problem is established, which includes a transmission delay model (Equations (1)-(2)) and a microservice processing time model (Equations (3)-(5)). The completion time of the Internet of Things application (Equation (7)) is used as the core optimization objective, and the container resource constraints (Equations (11)-(12)) are used as hard boundary conditions. The microservice grouping structure formed by the minimum cut partition is input into the deep Q-network as prior knowledge. By designing a composite reward function that includes the immediate communication cost, long-term reliability reward (Equations (8)-(9)), and resource utilization penalty (Equations (13)-(15)), the agent is driven to learn the optimal offloading strategy in a dynamic environment. The experience replay mechanism in Algorithm 2 (lines 12 - 17) realizes the closed-loop feedback of static topology optimization and dynamic resource scheduling, enabling the system to not only maintain an efficient communication mode among microservices but also adapt to the real-time load fluctuations of edge nodes. Finally, on the premise of meeting the reliability objective of Equation (10), a multi-objective balance of computing offloading efficiency and system stability is achieved. The core output of Algorithm 1 (FFAA) is the optimal microservice grouping structure obtained through network flow analysis. Algorithm 2 (improved DQN algorithm) finally outputs three key results, which together constitute the optimal computing offloading solution in the industrial Internet edge computing environment: (1) the optimal offloading strategy π*; (2) the converged Q-value function Q*; (3) the performance metric set O.

[0119] The present invention proposes a reliability-optimized computing offloading method based on topology reconstruction. First, the topology of the microservices of the Internet of Things application is reconstructed through the Ford-Fulkerson approximation algorithm to reduce the communication cost. Then, the DQN is used to dynamically adjust the task offloading strategy and optimize the resource allocation. The specific implementation is as follows:

[0120] In modern microservice architecture, with the continuous expansion of system scale and the increase of complexity, closely communicating microservices often face problems such as high latency, bandwidth bottlenecks and waste of resources. If these microservices are deployed on different nodes, frequent cross-node communication not only reduces the response speed of the system, but also may affect the overall performance and scalability. Therefore, it has become an urgent need to improve system efficiency, reduce resource waste, and improve fault tolerance and scalability to reasonably split and merge these closely connected microservices and optimize their deployment strategies. In this context, how to intelligently identify and optimize the deployment relationship of these microservices has become a key issue to be solved. The present invention will adopt the Ford-Fulkerson approximation algorithm (FFAA) to solve these challenges in a targeted manner by calculating the maximum flow between tasks, thereby achieving more efficient microservice splitting and merging.

[0121] Based on graph theory, IoT applications can be mapped into a directed acyclic graph (DAG) to construct Figure 3 The IoT application microservice topology G = (V, E) is shown. Each node represents a microservice, and the value on the line between adjacent microservices represents the communication cost between the two, expressed as ω(e(v x ,v y )) represents, where x represents the next microservice connected to microservice y, and its value is Among them, e(v x ,v y ) represents node v x With v y The edges connected between them are nodes, i.e. microservices, so ω(e(v x ,v y )) represents the weight, i.e., the communication cost. The communication cost between microservices x and y is calculated according to formula (6). FFAA is used to obtain closely connected microservices, such as Figure 3 (Left) shows the microservice set {2, 4, 5, 8} and the set {3, 6, 7, 9}. If the closely communicating microservices are not split and merged (i.e., topology reconstruction) and are directly unloaded to the industrial Internet environment, the communication between them will consume a lot of industrial Internet communication resources. Therefore, it is necessary to analyze these microservices to realize their topology reconstruction, so as to obtain the reconstructed microservice topology structure, such as Figure 3 (right) As shown. Afterwards, the microservice characteristics and the optimization goals in the reliability model are used to decide whether to deploy the microservice locally or offload it to the edge cloud network platform or industrial cloud platform.

[0122] The Ford-Fulkerson algorithm (FFA) is a greedy algorithm used to calculate the maximum flow in a flow network. It continuously finds augmenting paths and adjusts the flow until no more flow can be increased. However, in the mobile edge computing environment, the task offloading problem in Internet of Things applications has the characteristics of dynamic network topology and task heterogeneity. The traditional FFA faces challenges in dealing with complex data dependencies and real-time requirements. Therefore, the present invention proposes a method based on FFAA: First, the dependency relationships between microservices of Internet of Things applications are regarded as a network flow model, where nodes represent microservices and the edges connecting two nodes represent the communication time between microservices. Then, by finding augmenting paths to optimize the splitting of microservices, each time a path with the maximum residual capacity is found from the source point (the start of the Internet of Things application) to the sink point (the end of the application), and the flow is increased through this augmenting path, thereby reducing the residual capacity of each edge on the path. This process is repeated until no more augmenting paths can be found. Finally, the maximum flow is obtained, which represents the optimal microservice splitting scheme, that is, which microservices should be deployed together (to reduce communication costs) and which should be deployed separately, so as to optimize resource utilization and communication efficiency. By gradually optimizing the task offloading path in a dynamically changing network topology, the computational tasks and resources are efficiently allocated among edge nodes. By regarding the task offloading problem as a flow network problem, FFAA can approximately calculate the optimal offloading path, thereby obtaining a microservice architecture with close communication, ensuring that the task completion time and bandwidth consumption are optimized, and improving the overall system performance. The pseudocode part is shown in Algorithm 1.

[0123]

[0124]

[0125] The topology reconstruction method for Internet of Things applications based on FFAA proposed by the present invention (Algorithm 1) optimizes microservice deployment through flow network modeling and dynamic segmentation mechanism. The algorithm takes the microservice topology structure G=(V, E), edge weight ω(e), and the maximum number of groups K1 as inputs, where V and E represent the set of nodes and the set of edges respectively, and the edge weight ω(e) represents the communication cost between microservices and between microservices. Since there is no specific connection relationship here, it is simplified to ω(e). In fact, it is ω(e(v x ,v y )) below, which belongs to the same concept as ω(e(v, u)) later and can be unified into ω(e(v x ,v y)) are all calculated by formula (6). The edge weight ω(e(v,u)) represents the edge weight between nodes (microservices) v and u. The maximum number of groups K1 is adjusted according to the number of edge servers P: K1 ≤ P / 2 (ensuring sufficient computing resources for each group). First, at line 1, an auxiliary flow network G′ is constructed that includes a virtual source node s and a sink node t. The source node s, which is the starting node, is the entry point of the microservice task flow, and the sink node t, which is the terminating node, is the exit point of the microservice task flow. The construction of the auxiliary flow network G' is first based on a copy of the original microservice dependency graph G (as Figure 2 shown), and then the virtual source node s and sink node t are added. The source node s is connected to all microservice nodes, and the edge capacity is the sum of the communication costs of the outgoing edges of each node (reflecting the total load of sending data). The sink node t is connected to all microservice nodes, and the edge capacity is the sum of the communication costs of the incoming edges of each node (reflecting the total load of receiving data). At the same time, the dependency edges between the original microservices are retained, and their edge weights ω(e(v,u)) are directly used as capacities. The resulting auxiliary flow network G' transforms the microservice communication optimization problem into a solvable maximum flow minimum cut problem through this special network flow structure, providing a mathematical basis for the subsequent FFAA algorithm to group and minimize the cross-node communication cost. In lines 2 - 5, the capacities of the edges are calculated. The edge capacity from the source node s to each microservice node v is initially set to the sum of the communication costs of all its outgoing edges, and the edge capacity from a node to the sink node t is the sum of the communication costs of all its incoming edges. The outgoing and incoming edges respectively characterize different data flow characteristics between microservices. The outgoing edge e(v,u) represents the communication link for microservice v to send data to the subsequent dependent node u, and its edge weight ω(e(v,u)) is calculated by formula (6), quantifying the delay cost generated by data transmission on this link. For example Figure 2 in the two outgoing edges from microservice T1 to T4 and T5, they respectively correspond to the communication overhead generated when T1 outputs the processing results to these two downstream nodes. The incoming edge e(u,v) represents the dependency relationship for microservice v to receive input data from the upstream node u, and its edge weight ω(e(u,v)) is also determined by formula (6), reflecting the delay cost of this data input link, such as Figure 2The microservice T10 receives incoming edges from T4 and T5, reflecting the data dependency constraints in task processing. This way of defining directed edges completely preserves the task execution order logic of the original microservice topology, while quantifying the costs of each communication link through edge weights. In lines 8 - 14, an s - t augmenting path in the auxiliary flow network G′ is searched iteratively, and the minimum flow of the path is calculated based on formula (6) to update the residual capacity, dynamically partitioning microservice groups with high communication density. In lines 8 - 14 of Algorithm 1, the core process of microservice dynamic grouping is implemented using an iterative network flow optimization mechanism through the Ford - Fulkerson method. First, a feasible path path from the virtual source node s to the sink node t is found in the auxiliary flow network G' based on breadth - first search (BFS). The generation condition for this path requires that the residual capacity cap(u, v) > 0 for all constituent edges (u, v). For each discovered augmenting path, its bottleneck flow min_flow = min{cap(u, v)|(u, v) ∈ path} is calculated. This minimum value essentially determines the maximum traffic load that the current path can carry. In the traffic pushing phase, the algorithm performs a capacity reduction operation of cap(u, v) -= min_flow on all forward edges along the path path, and at the same time performs a capacity compensation of cap(v, u) += min_flow on the reverse edges. This two - way update mechanism reserves the possibility for subsequent reverse traffic adjustment. At the level of microservice grouping implementation, after each traffic push, the set of reachable nodes reachable starting from the source node s is recalculated through BFS, and the intersection S of these reachable nodes and the remaining nodes remaining_nodes kThat is, a micro-service group with intensive current communication is formed. Its physical meaning is that there is still an effective communication path from the source node s to the nodes in this group in the residual network G'. This indicates that the data interaction intensity among the micro-services within the group is higher than that between groups. By iteratively executing this process, the algorithm gradually clusters the micro-service nodes with high coupling degree into groups, and simultaneously automatically meets the objective of minimizing the communication cost defined by Equation (6). The finally output group structure S = {S1,..., Sk} is the micro-service topology partitioning scheme optimized by network flow. After each iteration, the breadth-first search is used to extract the subset of nodes connected to the source node in the current residual network as the candidate group. This operation is implemented in lines 15 - 16. The breadth-first search (BFS), as the core function for graph traversal, starts from the virtual source node s, scans all the connected paths in the residual network G' where there is still available capacity (cap(u, v) > 0), and returns the set reachable of these reachable nodes. This operation automatically clusters the highly cohesive micro-service nodes by dynamically tracking the connectivity changes after the network flow is pushed: when two micro-service nodes remain connected in the residual network, it means that there is still remaining capacity in their communication link, so they are partitioned into the same group Sk; conversely, if they are disconnected due to traffic saturation, they will be assigned to different groups. This mechanism essentially realizes the engineering application of the maximum flow - minimum cut theorem, maximizing the internal communication intensity within the finally formed micro-service groups (the sum of edge weights within the group is the largest), while minimizing the inter-group communication cost (the sum of edge weights of the cut edges is the smallest), thus directly optimizing the cross-node communication delay defined by Equation (6) and providing the initial topology structure with the optimal communication efficiency for the subsequent DQN dynamic offloading strategy. In lines 16 - 20, the candidate group and the remaining nodes are intersected to generate the new group Sk until the upper limit of the number of groups is reached or all nodes are assigned. Finally, in lines 21 - 24, the unassigned nodes are incorporated into the result set S as independent groups. Incorporating the unassigned remaining nodes as independent groups into the result set S ensures that all micro-service nodes can be completely covered, avoiding task loss, and at the same time maintaining the constraint that the total number of groups does not exceed the preset value K1. This processing method is particularly applicable to the situation where there are isolated nodes or low-coupling micro-services. For example, when the communication cost of a certain micro-service with other nodes is extremely low, treating it as an independent group can prevent a sharp increase in communication overhead caused by forced merging. At line 25, the mutual exclusivity is ensured by merging overlapping groups. The merging of overlapping groups optimizes the final group structure through the following steps: first, detect all the groups in the result set S that have common nodes (such as ) Then, find the union of these groups to form new groups (S' = S1 ∪ S2). Finally, remove the atomic groups and retain the new merged groups. This process eliminates redundant partitions that may occur due to multiple traffic pushes. For example, if microservice v appears in two groups simultaneously, it indicates that it plays a bridging role in the task flow. After merging, it can more accurately reflect the actual communication topology. The finally output result set S satisfies: 1) The communication intensity between nodes within a group is higher than that between groups; 2) All groups do not overlap; 3) The total number of groups ≤ K1, providing an optimal micro-service deployment solution with the best structure for subsequent calculation offloading. This method is based on the maximum flow - minimum cut theorem, aggregating tightly coupled microservices into the same group to minimize the communication overhead across edge nodes, thereby significantly reducing the end-to-end delay of IoT applications, providing an optimized topology basis for subsequent DQN-based dynamic offloading strategies, and ultimately enhancing system reliability. Experiments show that this algorithm significantly reduces the communication cost in complex task dependency scenarios, providing theoretical support for the efficient and reliable operation of industrial Internet edge computing.

[0126] The role of the above FFAA algorithm is to identify highly coupled and closely related microservices and aggregate them, which can be regarded as a new microservice, thereby reducing communication costs, reducing the completion time of IoT applications, and indirectly improving the reliability level.

[0127] In the industrial Internet edge computing environment, the contradiction between IoT application topology reconstruction and high-reliability requirements has given rise to the need for a new type of computing offloading paradigm. Aiming at the limitations of traditional methods in coordinating microservice dependencies (such as the critical path constraint in Equation (7)) and dynamic resource states, the critical path constraint defined in Equation (7) reveals the essential bottleneck problem of the completion time of IoT applications. This constraint indicates that when there are complex dependencies between microservices (such as Figure 2 the directed acyclic graph shown), among all possible execution paths from the start node to the end node, the path with the longest total time consumption (i.e., the critical path) will directly determine the completion time of the entire application. The present invention proposes a DQN-based dynamic offloading method for IoT applications. By constructing a multi-dimensional state awareness space, deeply integrating the topology reconstruction results output by FFAA (i.e., the microservice grouping structure of the minimum cut partition) with real-time network parameters (channel gain, node load, etc.), a composite reward mechanism oriented to reliability goals is designed: taking the reliability level as the core optimization index, while constraining the communication cost and task completion time. Microservices can be understood as micro-tasks. If the task completion time is shortened, then the completion time of the entire IoT application will also be shortened. In the composite reward mechanism oriented to reliability goals, the constraint of communication cost is mainly achieved by combining the dynamic weight penalty mechanism and topology-aware reward. The communication time model in Equation (6) is embedded in the reward function. When the DQN agent deploys microservices with data dependencies to different edge nodes (i.e., α ixy = 1), the system will calculate according to the cross-node communication time Tixy Automatically generate a negative penalty term, whose penalty intensity is proportional to the amount of transmitted data in xy and out yx is proportional to, and inversely proportional to the available bandwidth B upload and B download is inversely proportional to. At the same time, by real-time monitoring the bandwidth utilization rate U of the edge node r , an additional dynamic penalty is imposed when multiple microservice groups compete for the same node bandwidth resources. This constraint mechanism forms a synergistic effect with the topology reconstruction result of the FFAA algorithm - when the agent deploys the co-group microservices (such as {2, 4, 5, 8}) divided by Algorithm 1 to the same edge node, a co-location reward will be obtained, while forced split deployment will trigger a penalty proportional to ω(e(v,u)). This dual force drives the system to naturally converge to the offloading strategy with the minimum communication cost while satisfying the critical path constraint of Equation (7).

[0128] The agent dynamically generates a microservice deployment strategy through interactive learning, and under the premise of satisfying the container memory constraint (Equation (12)), realizes the coordinated improvement of the edge node resource utilization rate and system reliability. Through the closed-loop feedback of the topology reconstruction and offloading decision engine, the framework continuously optimizes the microservice grouping granularity and deployment location, and finally achieves the optimization goal defined by Equation (10), providing an adaptive and strongly robust computing offloading solution for dynamic industrial scenarios.

[0129] To achieve efficient computing offloading of Internet of Things applications in the industrial Internet edge computing environment, the present invention constructs a DQN-based computing offloading framework, as Figure 4 shown. Through the reinforcement learning algorithm, the framework dynamically selects the optimal offloading strategy according to the current system state to improve the system reliability and resource utilization rate. In Figure 4 , during the process of splitting Internet of Things applications into multiple mutually dependent microservices and offloading them to the industrial Internet edge computing environment, the agent continuously interacts with the industrial Internet edge computing environment. For local execution, containers in edge servers in the edge cloud platform, and industrial cloud platforms, the offloading strategy will offload the groups obtained by the FFAA algorithm to the most suitable places through DQN training. Specifically, the agent observes the system state s t (including container load L r , channel gain Elements such as microservice dependencies. Since the platform in this paper is a simulation platform built based on a Python simulator, it is described in detail in the performance evaluation parameter settings. Lr represents the memory capacity of the container, which is randomly selected from [1024, 4069] MB. The channel gain range is also randomly selected, with a range of [2, 4]. The microservice dependency is the edge between two microservices, and the edge has a weight, which is the communication cost. A q-dimensional offloading action vector At = [a1, a2,..., aq] is generated through a deep Q-network, where each action component corresponds to a deployment decision of a microservice group between edge nodes, that is, an offloading strategy. The three mentioned above include: local / cloud / edge. The implementation process of the DQN generating a k-dimensional offloading action vector At = [a1, a2,..., aq] is completed through a hierarchical decision-making mechanism. The implementation process first encodes the multi-dimensional system state including container load, channel state, and microservice dependencies into the input of the neural network. After feature extraction and non-linear transformation by a three-layer fully connected network, the output layer generates a Q-value matrix for all edge nodes for each microservice group divided by the FFAA algorithm, representing the long-term revenue expectations of different deployment strategies. Based on the improved ε-greedy strategy, the system independently selects the target node with the largest Q value for each microservice group, and maintains the output dimension as q through the virtual filling or group merging mechanism to ensure compatibility with the subsequent scheduling system. During the execution phase, the Kubernetes controller converts the action vector into specific container deployment instructions, while monitoring the node resource constraints (such as the memory capacity Lr) in real time, giving priority to ensuring the low-latency deployment of microservices on the critical path, and dynamically adjusting the grouping strategy through a periodic re-evaluation mechanism, ultimately achieving the multi-objective balance of minimizing communication costs and maximizing system reliability. The entire process embeds the communication cost constraint of Equation (6) and the critical path limit of Equation (7), and drives the update of network parameters through the composite reward function designed by Equations (13)-(15), forming a closed-loop optimization system from topology awareness to resource scheduling.

[0130] The obtained interaction data <s, a, r, s′> is stored in the experience replay pool. When there is enough data in the experience replay pool, a data of size batch_size is randomly taken out from the experience pool, and the predicted value of Q is calculated using the current network (prediction network). The current network (prediction network) is the core component of the DQN. It maps the system state to a q×|Z|-dimensional Q-value matrix through a three-layer neural network, predicting the long-term revenue of each microservice group deployed to different edge nodes and driving the intelligent offloading decision. The current network calculates the Q prediction value through forward propagation: input the system state (node load, channel quality, etc.), process it through a three-layer fully connected neural network, and output a q×|Z|-dimensional matrix, where each value Q(s, a|θ) represents the expected total revenue (combining the immediate communication cost and the long-term reliability revenue) of deploying the microservice group to the node.

[0131] The Q-target value is calculated using the Q-Target network, based on the next state s t+1 and the reward r t , and the target network parameters θ with delayed update are used for forward propagation to obtain maxQ(s t+1 , a|θ-), and then combined with the Bellman equation Qtarget = rt + γ·maxQ(s t+1 , a|θ-) to calculate the final target value for stabilizing the DQN training. Then the loss function between the two is calculated. The loss function uses the mean squared error (MSE) to measure the gap between the target value and the predicted value, aiming to make the predicted value as close as possible to the target value. Gradient descent is used to update the current network parameters, calculating the gradient of the loss function with respect to the current network parameters and backpropagating to adjust the network weights according to the set learning rate, so that the predicted value gradually approaches the target value, thereby optimizing the policy. After repeating several times, the parameters of the current network are copied to the Q-Target network, significantly improving the stability and convergence of DQN training. Since the target network is used to calculate the Q-target value, if the current network is directly used to calculate the target value, it will cause the target value to fluctuate violently with the rapid update of the current network, making it difficult for the training process to converge. By fixing the parameters of the target network and updating them regularly, the relative stability of the target value can be maintained for a period of time, thereby providing a more reliable learning target for the current network and effectively alleviating the training oscillation problem caused by the frequent change of the target value. In addition, this delayed update mechanism can also reduce the overestimation bias of Q-value estimation, making the learning process smoother, and ultimately helping the agent to more reliably learn the optimal policy. Among them, s and a represent the current state and action, and the state s is composed of multi-dimensional system parameters, including: the real-time load (CPU / memory utilization) of the edge server, the container processing capacity (Cj), and the channel gain The communication cost (ω(e)) between microservices and the DAG topology of the IoT application. These parameters are continuously collected by the monitoring module of the edge cloud platform and integrated into the observation space of the agent. The action a is the microservice deployment decision generated based on the current state, specifically: according to the microservice grouping S reconstructed by the FFAA algorithm, select to assign each microservice group to the optimal edge container (meeting the memory constraint of Equation (12)) or the industrial cloud platform. r and s′ represent the reward value of the current action and the next state. When the agent executes a certain action a (i.e., the microservice deployment decision) in the current state s, the system immediately calculates the multi-dimensional composite reward value r, which comprehensively considers key indicators such as the completion time of the IoT application (negative incentive), whether the deadline is exceeded (reliability penalty), and the container resource utilization rate (load balancing reward), and automatically balances the priorities of different optimization goals through a dynamic weight mechanism. At the same time, the environment updates the entire system state according to the executed action, including recalculating the container load of the edge node, adjusting the communication cost between microservices, updating the network channel state parameters, etc., so as to generate the new state s' at the next moment. This closed-loop interaction process of "state-action-reward-new state" is continuously recorded in the experience replay pool. Among them, the Q-Target network can also be denoted as the neural network for generating the target value; similarly, the prediction network can also be denoted as Q(s,a|θ), which represents the neural network for predicting the state-action value function. Thus, the following formula can be established:

[0132] Q k+1 (s t ,a t |θ)=Q k (s t ,a t |θ)+a k E k (13)

[0133]

[0134] In Equations (13) and (14), a k and γ represent the learning rate and the discount factor respectively, s' is the predicted value of the state after executing the action a t in the k-th iteration, a' is the action with the maximum reward value in the state s', and E k is the cumulative reward value of the k-th iteration in the iterative process. a t is the action selected by the agent at time t, specifically referring to (1) the microservice grouping reconstructed by the FFAA algorithm (such as S = {S1,...,S k} decisions for deployment; (2) including allocating each microservice group to specific edge containers or industrial cloud platforms; (3) the action space design is constrained by Equation (11) to ensure that each microservice is only deployed in a single container, r t The immediate reward obtained after executing action a at time t, θ, t represent the weight parameters of the neural network for generating offloading decisions; the former and the latter represent the weight parameters of the predictive Q-network and the target Q-network, respectively.

[0135] The DQN agent randomly samples a small portion of experiences e t from the experience replay pool D = {e1, e2, …, e t} to calculate the target value t e = (s t , a t , r t , s') At the k-th time, by minimizing the objective function of Equation (15), the network parameters θ can be updated, thus realizing the update of the offloading decision.

[0136]

[0137] After repeating the above steps to obtain the best experience (s t , a t , r t , s') at time t, this experience is put into the experience replay pool as a new training sample. When the capacity of the experience replay pool is sufficient, newly generated experiences will replace the old data samples. The deep learning network learns the best experience (s t , a t , r t , s') and, over time, generates better decision outputs. Under the constraint of limited storage, the deep neural network only learns from the newly generated data samples, and this reinforcement learning mechanism will continuously improve its offloading strategy until convergence.

[0138] The pseudocode based on the above framework is shown in Algorithm 2.

[0139]

[0140]

[0141] In the industrial Internet of Things edge computing environment, among the input parameters of Algorithm 2 (the improved DQN algorithm based on the industrial Internet of Things edge computing environment), the simulation environment set E covers the edge-cloud collaborative network architecture, Internet of Things device attributes (transmission power pi, channel gain ​) Industrial application topology (microservice dependency relationship G(V,E) and its communication cost ω(e) and network dynamic parameters (bandwidth B, noise power N0); the system initial state S includes microservice instruction length L ix , container processing capacity C j , edge node load Lr and real-time link rate The exploration rate ∈ controls the exploration intensity of the intelligent agent for unknown edge node resources, and the discount factor γ weighs the optimization of immediate communication delay (the communication time T between microservices x and y for the Internet of Things application i ixy ) and long-term reliability guarantee (the objective function F(i) is to maximize the reliability level); the learning rate α adjusts the adaptation speed of the neural network to dynamic network fluctuations (such as time-varying channel gain ); the capacity K of the experience replay pool stores historical state-action-reward tuples (st, at, rt, s'), and overcomes the temporal correlation caused by industrial task dependency chains through random sampling; the target network update frequency C suppresses the Q-value drift caused by node failures or load mutations, and the number of training rounds U and the maximum number of steps R per cycle cover the entire life cycle of complex industrial tasks (from microservice splitting to application completion time T i calculation). In the output result, the optimal policy π is the microservice deployment sequence based on Ford-Fulkerson topology reconstruction (the x in formulas (11)-(12) ixr decision), Q is the converged action value function, representing the long-term reliability benefit of the offloading action a t in the state s t , and the performance metric O is based on the Internet of Things application completion time T iQuantify the reliability of the system with the deadline violation rate. The algorithm initializes the experience replay pool and the double Q-network in lines 1-3 of the code, and constructs an edge-cloud collaborative decision-making framework: create an experience replay pool to store the experience data generated during the interaction between the agent and the environment, including the current state, the actions taken, the rewards obtained, and the next state. At the same time, set the capacity of the pool to K to limit the amount of stored data. When new data arrives and the pool is full, the earliest stored data will be replaced. Then, construct two neural networks - the main Q-network and the target Q-network. The main Q-network is used to evaluate the value of each action according to the current state, that is, to predict the long-term reward that may be obtained by executing a certain action, while the target Q-network is used to provide a more stable target value to reduce the fluctuations during the learning process. Initially, copy the weights of the main Q-network to the target Q-network so that they have the same initial parameters. In this way, a foundation is laid for the agent to dynamically decide the task offloading strategy in the edge computing environment, enabling it to intelligently select whether to offload tasks to edge nodes or the cloud according to the real-time network state and resource situation during the subsequent learning process, so as to optimize the overall performance and reliability of the system. In lines 4-17, the dynamic offloading strategy is optimized through M rounds of iterative training. Among them, in lines 5-10, combined with the ε-greedy strategy and the microservice topology reconstruction result, the container deployment action a t (such as offloading communication-intensive microservice groups to the same edge server). First, according to the ε-greedy strategy, the agent randomly selects an action with probability ∈ and selects the optimal action in the current state with probability 1 - ∈. Then, using the result of microservice topology reconstruction, combine the microservice grouping with the resource status of edge nodes or the cloud to generate the specific container deployment action a t , this action determines the deployment location of each microservice grouping in the edge computing environment, thus realizing the dynamic computing offloading decision. In lines 11-12, store the state transition data to enhance the adaptability to the industrial non-stationary environment. When the agent executes the action a t in the environment, observe the reward r t given by the environment and the new state s t+1 , and then store the quadruple (φ t , a t , r t , φ t+1 ) into the experience replay pool D, where φ t is the preprocessed representation of the current state s t , and φ t+1 is the preprocessed representation of the new state s t+1Preprocessing representation. In this way, the algorithm can record the interaction experiences at different time steps and different environmental states, and then randomly sample these experiences from the experience replay pool for learning, so as to better adapt to the dynamic changes and non-stationary characteristics that may occur in the industrial environment. At lines 13 - 14, the Q-network parameters θ are updated based on the temporal difference error to minimize the communication cost T ixy and the processing delay T ixj weighted loss; by randomly sampling a batch of experience samples from the experience replay pool, calculating the target value for each sample, which consists of the immediate reward and the discounted maximum Q-value, then calculating the difference between the current Q-network prediction value and the target value as the temporal difference error, and finally using this error to update the parameters of the Q-network through gradient descent, so that the network can better predict the value of actions, thereby improving the decision-making ability of the agent in a dynamic environment. At line 15, through the periodic synchronization of the target network parameters θ-, the policy oscillation caused by sudden changes in the edge node load is suppressed, and finally the reliability improvement with the application completion time as the core is achieved. The update of the target network parameters is achieved by periodically copying the parameters of the main Q-network to the parameters of the target Q-network. Specifically, every fixed number of steps C, the current parameter values of the main Q-network are completely copied to the target Q-network, so that the parameters of the target Q-network are synchronized with the main Q-network. This periodic synchronization mechanism can ensure that the target Q-network provides relatively stable target values, thereby reducing the instability caused by frequent changes in the target values and the oscillation during the learning process during training, and helping to improve the convergence speed and stability of the reinforcement learning algorithm. Finally, at line 16, the optimal offloading strategy, the converged action value function, and the specific numerical values of the obtained performance metrics are output. The optimal offloading strategy is determined by evaluating the state-action value function Q(s,a|θ) of the main Q-network after training is completed, and at the same time, the specific numerical values of the key performance metrics during the training process are recorded. In specific implementation, first traverse all possible states s and actions a, and select the action that maximizes the value of Q(s,a|θ) as the optimal offloading strategy π. Then, the converged Q(s,a|θ) is output as the action value function Q. In addition, according to the data recorded during the experiment, performance metrics such as the average completion time Ti of the Internet of Things application and the reliability level are calculated and output, and these metrics reflect the actual effect of the algorithm in optimizing offloading decisions.

[0142] The following introduces the evaluation preconditions including the evaluation environment, parameter settings, comparison algorithms, and evaluation metrics, and then conducts comparative experiments on the completion time and reliability of Internet of Things applications obtained by four methods under different numbers of Internet of Things applications, as well as comparative experiments on the completion time of the same Internet of Things application topology before and after reconstruction under four methods.

[0143] To verify the performance of the proposed reliability computing offloading method based on IoT application topology reconstruction in the industrial Internet edge computing environment, a simulation platform was built using a Python simulator. The platform simulator constructed a simulated industrial Internet edge computing system. The experimental environment consisted of 4 edge clouds, 10 servers, 20 containers, several IoT devices, and an industrial cloud platform. The edge cloud nodes were deployed in different geographical locations and interconnected through high-speed networks. Each edge cloud node contained multiple heterogeneous edge servers, and different numbers of containers were running on these edge servers. The IoT devices communicated with the nearest edge server through wireless networks, sending data and requesting task processing. The hardware configuration of the experimental environment included high-performance servers for simulating edge cloud nodes and the industrial cloud platform, as well as multiple low-power devices for simulating IoT devices. The software environment included an operating system, network simulation tools, and a custom simulation framework for simulating the operation and data transmission of IoT applications.

[0144] In the experiment, the parameters of the configured simulation environment covered the settings of the edge cloud network platform, edge servers, containers, IoT applications, microservices, IoT devices, and the industrial cloud platform. Different edge cloud base stations were interconnected through the edge network. The number of servers in each edge cloud was randomly selected from the set [2, 4]. The processing power of the edge servers ranged from 5000 MIPS to 8000 MIPS, the memory capacity ranged from 8192 MB to 16384 MB, and the bandwidth ranged from 500 Mbps to 800 Mbps. The IoT devices were connected to the base stations through wireless networks for communication purposes. The default noise power values of the edge cloud (service network) / wireless channel interference, optical fiber transmission noise, etc. were 0.1 W, and the bandwidth range was randomly selected from the set [100, 500] Mpbs. The processing power range of each was randomly selected from the set [1000, 2000] MIPS, the memory capacity was randomly selected from the set [1024, 4069] MB, and the upload and download bandwidths were both randomly selected from the set [100, 200] Mbps. The data size of the IoT devices ranged from 5 MB to 15 MB, the transmission power ranged from 1 W to 3 W, and the channel gain ranged from 2 to 4. In addition, to evaluate whether the system met the service quality requirements, a deadline needed to be set for each IoT application, ranging from 2 to 4 seconds. Each application contained a different number of microservices, randomly generated from the range [15, 20]. The processing power of the industrial cloud platform was 20000 MIPS, the storage capacity was 10000 GB, and the bandwidth was 10000 Mbps. The instruction length range of each microservice was [500, 1200] (instructions), and the memory requirements and output data size ranges were [512, 2048] MB and [10, 30] MB respectively.

[0145] These parameter settings provide diverse test scenarios for the experiment, ensuring the broad applicability and reliability of the results.

[0146] To verify the performance of the DQN-based computing offloading method of the present invention, the following three methods were selected for comparative experiments:

[0147] (1) Random Offloading (RO): The microservices are assigned to edge nodes / edge servers or the cloud / industrial cloud platform in a completely random manner. When multiple containers meet the constraints, a virtual machine is randomly selected, and each task of the Internet of Things application is accommodated from top to bottom along its directed acyclic graph. It has strong randomness and large volatility.

[0148] (2) Offloading method based on PPO (Proximal Policy Optimization): It adopts a policy gradient optimization mechanism to directly generate actions by adjusting the parameters of the policy network. Although it can dynamically adapt to network fluctuations, it is not combined with a topology reconstruction algorithm, making it difficult to effectively handle the complex dependency relationships between microservices. Moreover, during the training process, fine-tuning of hyperparameters is required to avoid policy oscillation, resulting in low convergence efficiency in the industrial Internet of Things edge computing environment.

[0149] (3) Offloading method based on Q-learning: It relies on the iterative update of the Q-value table for the long-term reward of the state-action pair, with clear theoretical convergence. However, when facing the high-dimensional continuous states of the industrial Internet of Things (such as multi-node load, dynamic channel gain), the storage and computing overhead increase exponentially, and the fixed exploration rate mechanism is prone to falling into local optima in complex task dependency scenarios, unable to balance the immediate communication cost and the long-term reliability goal.

[0150] To comprehensively evaluate the performance of the proposed method, the following metrics were selected:

[0151] (1) Reliability level: By calculating the difference between the completion time and the deadline of the Internet of Things application, the probability of violating QoS is statistically analyzed, and thus the corresponding probability of meeting the QoS requirements is obtained to measure the reliability level of the system.

[0152] (2) Completion time of Internet of Things application: Record the total time from the start of offloading to the completion of offloading of the Internet of Things application, including the processing time and processing delay of microservices and the communication time between microservices.

[0153] The reliability level of the Internet of Things application is a key indicator to measure whether it can meet the Quality of Service (QoS) requirements. Based on the above-built industrial Internet of Things edge computing simulation environment, a comparative experiment on the reliability levels obtained by the four methods under different numbers of Internet of Things applications was carried out. Figure 5It shows the mean change in the reliability level of four computing offloading strategies (DQN, PPO, Q-learning, and RO of the present invention) under different numbers of Internet of Things (IoT) applications. The number of IoT applications increases from 5 to 25 with a step size of 5. Through comparative analysis, the performance of different strategies under dynamic network topologies and complex task dependencies is revealed.

[0154] As the number of IoT applications increases, the load on the system gradually increases, which poses higher requirements for the intelligence and adaptability of computing offloading strategies. From Figure 5 it can be seen that the computing offloading strategy based on the deep Q-network (DQN) of the present invention shows significant superiority in terms of reliability. The reliability level of the DQN strategy continuously rises as the number of IoT applications increases and approaches 1.0 (reaching 0.91) under high load conditions. This result indicates that the DQN of the present invention can dynamically adapt to environmental changes through reinforcement learning, intelligently optimize offloading decisions, and thus achieve efficient task allocation and resource management in a complex industrial Internet environment. Through continuous learning and adjustment, the DQN strategy can intelligently select the optimal offloading path according to factors such as real-time network conditions and device computing capabilities, avoiding the limitations of traditional methods. Especially in multi-device collaborative work and load balancing, the DQN shows significant advantages, which enables it to maintain a relatively high reliability level in high-load scenarios.

[0155] In contrast, the reliability level of the PPO strategy is at a medium level. Although PPO, as an advanced reinforcement learning algorithm, can adapt to dynamic environments to a certain extent, there is an obvious gap between it and the DQN in high-load scenarios. This may be because when dealing with complex task dependencies and dynamic network topologies, the stability and convergence speed of PPO's strategy update are not as good as those of the DQN. Although the Q-learning strategy is also based on reinforcement learning, its adaptability in dynamic environments is slightly weaker. Its reliability level is slightly lower than that of the DQN and PPO, and the improvement amplitude is limited as the number of IoT applications increases. This shows that the learning efficiency and decision-making ability of Q-learning are restricted to a certain extent when dealing with complex industrial Internet scenarios. Due to the lack of an intelligent optimization mechanism, the reliability level of the random offloading strategy is the lowest, and the performance improvement amplitude is small when the number of IoT applications increases. This result further proves that relying solely on random strategies is difficult to meet the requirements of high reliability in the industrial Internet environment.

[0156] It is worth noting that the reliability computing offloading method based on the topology reconstruction of Internet of Things applications proposed by the present invention, combined with the DQN strategy, can effectively improve the overall performance of the system. By introducing the topology reconstruction mechanism, optimizing the deployment strategy of microservices, and reducing the cross-node communication overhead, the DQN strategy can make full use of the reconstructed microservice deployment strategy to achieve more efficient offloading decisions. This method not only improves the reliability of the system but also significantly reduces the completion time of Internet of Things applications, further demonstrating its effectiveness and superiority in the edge computing environment of industrial Internet.

[0157] Similarly, in the above-built simulation environment, Figure 6 It shows the completion time of Internet of Things applications under four computing offloading strategies with different numbers of Internet of Things applications. Figure 6 In it, the horizontal axis represents the number of Internet of Things applications, with a range of [5, 25] and a step size of 5. The vertical axis represents the completion time of Internet of Things applications.

[0158] As the number of Internet of Things applications increases, the load of the system increases, and the completion time generally shows an upward trend. This is because more Internet of Things applications require more computing resources to be allocated, and at the same time, the communication overhead between microservices also increases. However, there are significant differences in the performance of different strategies in dealing with the increasing load. The DQN strategy of the present invention shows the lowest completion time, and as the number of Internet of Things applications increases, the growth of its completion time is relatively gentle. This indicates that DQN can effectively allocate computing resources and reduce communication overhead by dynamically learning and optimizing offloading strategies. DQN dynamically adjusts offloading decisions through reinforcement learning, and can select the optimal offloading path according to the real-time network state and task requirements, thus significantly reducing the completion time.

[0159] In contrast, the average completion time of the PPO strategy is slightly higher than that of DQN, but in some cases, it is close to the performance of DQN. This indicates that PPO also has a certain degree of adaptability in a dynamic environment, but the stability and convergence speed of its strategy update are not as good as DQN. The completion time of the Q-learning strategy is higher than that of DQN and PPO, especially when the number of Internet of Things applications is large, and its performance improvement is limited. This indicates that when dealing with complex task dependencies and dynamic network topologies, the learning efficiency and decision-making ability of Q-learning are restricted to a certain extent. The completion time of the random offloading strategy is the highest, and as the number of Internet of Things applications increases, the improvement of its performance is the smallest. Due to the nature of its randomly generated strategy, its performance is extremely unstable. This shows that the random strategy lacks an intelligent optimization mechanism and is difficult to meet the performance requirements in high-load scenarios.

[0160] Figure 7The completion time changes of four computing offloading strategies for the same Internet of Things (IoT) application before and after topology reconstruction are further analyzed. In this experiment, a randomly generated IoT application containing 25 microservices is taken as an example, and the average completion time of different methods before and after topology reconstruction is obtained after 1000 rounds of training. The experimental results show that topology reconstruction has a significant optimization effect on the completion time of IoT applications. Without topology reconstruction, due to the large communication overhead between microservices, the completion time of IoT applications is relatively long. After topology reconstruction of microservices through the Ford-Fulkerson approximation algorithm, the communication overhead is effectively reduced, resource utilization is more efficient, and the completion time of IoT applications is significantly reduced. For the DQN method, the completion time after topology reconstruction is further optimized, indicating that it can make full use of the microservice deployment strategy after reconstruction to achieve more efficient offloading decisions. The PPO and Q-learning methods also show certain performance improvements after topology reconstruction, but the improvement amplitude is relatively small, indicating that their adaptability in optimizing microservice deployment is not as good as that of DQN. Although the random offloading method also has certain improvement after topology reconstruction, the completion time is still relatively high. This shows that topology reconstruction can not only optimize the deployment of microservices, but also significantly improve the performance of offloading strategies based on deep reinforcement learning (such as DQN), further proving the effectiveness and superiority of the method combining topology reconstruction and DQN proposed in the present invention in reducing the completion time of IoT applications.

[0161] The performance of the proposed method is verified through simulation experiments. The experimental environment includes 4 edge clouds, 10 servers and 20 containers. The experimental results show that compared with the random offloading (RO), PPO-based and Q-learning-based offloading methods, the method of the present invention shows significant advantages in both the reliability level and the completion time of IoT applications. For example, in the scenario of 25 IoT applications, the reliability level of the present invention is close to 1.0, and the completion time is reduced by about 30%. In addition, topology reconstruction significantly optimizes the completion time of IoT applications, further proving the effectiveness of the method in this paper.

[0162] In view of the reliability challenges brought about by the dynamic network topology and complex task dependencies in the industrial Internet of Things (IIoT) edge computing environment, a reliability optimization computing offloading method (FFAA-DQN) that integrates the topology reconstruction of IoT applications and deep reinforcement learning is proposed. By constructing a system model that includes an edge cloud network, an industrial cloud platform, and IoT devices, FFAA and DQN are innovatively combined in depth: the former minimizes the communication cost of microservice topology reconstruction based on maximum flow analysis, and the latter dynamically optimizes the task offloading strategy through real-time network state perception. Experimental results show that the present invention significantly improves the system reliability and performance in the IIoT edge computing environment under a dynamic network environment, effectively solves the problem of collaborative optimization of resource utilization efficiency and real-time performance in complex industrial scenarios, and provides a feasible technical path for the efficient and reliable operation of edge computing.

[0163] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing, characterized in that The steps are as follows: Step 1: Construct a system model covering the edge cloud network platform, industrial cloud platform, and Internet of Things devices, establish a transmission delay model and a reliability model, and construct a task offloading mathematical model to maximize the reliability level; Step 2: Based on the network flow theory, model the microservice dependencies as a weighted directed acyclic graph, perform topological reconstruction on the microservices of the Internet of Things application through the Ford-Fulkerson approximation algorithm, obtain the minimum cut partition to form a closely related microservice grouping structure; The microservice grouping structure after topological reconstruction is input into the deep Q network as prior knowledge, use the deep Q network to solve the task offloading mathematical model in Step 1, dynamically adjust the task offloading strategy, optimize resource allocation, and obtain the optimal computing offloading solution in the industrial Internet edge computing environment.

2. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 1, characterized in that The edge cloud network platform is respectively connected to the industrial cloud platform and the Internet of Things devices; The Internet of Things devices, as the terminal layer, establish a connection with the edge access platform through wireless communication technology, and use lightweight protocols to transmit the collected industrial field data; The edge access platform, as the gateway between the Internet of Things devices and the edge cloud network platform, performs identity authentication, protocol conversion, and data preprocessing on the accessed Internet of Things devices, and then distributes the data to the edge cloud network platform through the edge network; The edge cloud network platform is composed of a distributed edge server cluster and is interconnected through an intranet; Each edge server runs multiple containers, and the containers communicate through virtual network interfaces; The Internet of Things application is split into multiple microservices. After the edge cloud network platform receives the data stream from the access platform, it dynamically schedules the computing tasks to the optimal container according to the microservice topological relationship, and the processing results are uploaded to the industrial cloud platform through the edge-cloud collaboration channel; The industrial cloud platform is interconnected with multiple edge cloud network platforms through dedicated lines or VPNs to achieve global resource orchestration and big data analysis; After receiving the edge cloud aggregated data, the edge cloud network platform performs complex calculations, and the generated optimization model / strategy is sent to the edge server through the reverse channel to update the local processing logic, forming a closed-loop data stream of "device-edge-cloud". The system model randomly pre-deploys P heterogeneous edge servers into H edge cloud network platforms of the industrial Internet edge computing system; M containers indexed by a container set Z = {1, 2, …, M} are randomly assigned to H edge servers. Each Internet of Things device generates an Internet of Things application at certain times, and there are a total of I1 Internet of Things applications indexed by a set I = {1, 2, …, i, …, I1}. Each Internet of Things application contains multiple microservices that are dependent on each other. In the Internet of Things application, a series of microservices with dependencies are established as a directed acyclic graph G = (V, E), where V and E represent the set of nodes and the set of edges, respectively; each microservice is indivisible and can only be processed by container j, and container j has a certain processing capacity C j ; In the established directed acyclic graph, each node represents a microservice, the weight of each node represents the average processing time of the microservice, and each edge represents the communication time between microservices; The reliability of an Internet of Things application is determined by the path with the largest sum of the average processing time, processing delay, and communication time between microservices from the start node to the end node.

3. The reliability perception-based dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 2, wherein The task offloading mathematical model is: taking maximizing the reliability level of multiple Internet of Things applications as the modeling function: The constraint conditions are: 1) Each task can only be deployed to one container: 2) The memory capacity of each container is greater than or equal to the total memory required by the deployed tasks: where, x ixr indicates whether the x-th microservice of the Internet of Things application i is deployed to the r-th container; l ix represents the memory capacity of the x-th microservice of the Internet of Things application i, L r represents the memory capacity of the r-th container, X i represents the microservice set of the Internet of Things application i, |X i | represents the number of all microservices; I represents the Internet of Things application set containing I1 Internet of Things applications, F(T i -D i ) is the difference function of the probability of violating QoS, T i is the completion time of the Internet of Things application i, D i is the deadline, and Z is the container set containing M containers.

4. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 3, wherein The difference function of the probability of violating QoS is Among them, the difference (T i - D i ) indicates whether the QoS requirement of the Internet of Things application i is violated; The completion time T of the Internet of Things application i i is as follows: Among them, represents the transmission delay of data d transmitted from the IoT device that generates the IoT application i to the edge cloud h i , R i h represents the link data transmission rate between the IoT device that generates the IoT application i and the edge cloud h; T ixj The total processing time of the x-th microservice of the Internet of Things application i on the j-th container is the microservice processing time, and: where t ixj is the processing time of microservice x of IoT application i, and where L ix is to divide IoT application i into |X i | microservices, and the instruction length of microservice x ∈ X i , and C j represents the processing capacity of container j; The processing delay d of the microservice ixj ~Uniform(2, 6), where Uniform represents the uniform distribution function; T ixj is the communication time between microservice x and microservice y, and Among them, α ixy indicates whether microservice x and microservice y are on the same edge server. If so, the value of α ixy is 0; otherwise, it is 1. β ixy indicates whether microservice x sends data to microservice y. If so, the value of β ixy is 1; otherwise, it is 0. in xy and out yx respectively represent the output data size from microservice x to microservice y and the output data size from microservice y to microservice x. B upload and B download respectively represent the network upload bandwidth and download bandwidth.

5. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 3 or 4, characterized in that The Ford-Fulkerson approximation algorithm performs minimum cut partitioning on a weighted directed acyclic graph, aggregates communication-intensive microservices into homogeneous sub-task groups, and transforms the topology reconstruction problem into a maximum flow optimization problem by introducing virtual source and sink nodes. The minimum cut is solved by iteratively finding an augmenting path. The deep Q-network drives the agent to learn the optimal offloading strategy in a dynamic environment by designing a composite reward function that includes immediate communication cost, long-term reliability reward, and resource utilization penalty. It uses the experience replay mechanism to achieve a closed-loop feedback for static optimization of the topology structure and dynamic resource scheduling, which can not only maintain the efficient communication mode between microservices but also adapt to the real-time load fluctuations of edge nodes. Finally, on the premise of meeting the maximized reliability level, it realizes the multi-objective balance of computing offloading efficiency and system stability, and outputs the results of the optimal offloading strategy π*, the converged Q-value function Q*, and the set of performance metrics O, which together constitute the optimal computing offloading solution in the industrial Internet of Things edge computing environment.

6. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 5, wherein, In the weighted directed acyclic graph G=(V, E), each node represents a microservice, and the value on the connection line between adjacent microservices represents the communication time between microservices, denoted by the edge weight ω(e(v, u)), where e(v, u) represents the edge connecting nodes v and u. The Ford-Fulkerson approximation algorithm optimizes the splitting of microservices by finding an augmenting path. Each time, it finds an augmenting path with the maximum residual capacity from the virtual source node to the sink node, and increases the flow through the augmenting path and reduces the residual capacity of each edge on the augmenting path. Repeat this process until no more augmenting paths can be found, obtaining the maximum flow, which represents the optimal microservice splitting scheme. The implementation method of the Ford-Fulkerson approximation algorithm is as follows: Based on the directed acyclic graph G, retain the dependency edges between the original microservices and directly use their edge weights ω(e(v, u)) as capacities to construct an auxiliary flow network G' that includes a virtual source node s and a sink node t. The source node s is the entry of the microservice task flow, and the sink node t is the exit of the microservice task flow. The source node s is connected to all nodes, and the edge capacity is the sum of the communication costs of the outgoing edges of each node. The sink node t is connected to all microservice nodes, and the edge capacity is the sum of the communication costs of the incoming edges of each node. Calculate the edge capacity: The edge capacity from the source node s to each node v is initially set to the sum of the communication costs of all outgoing edges, and the edge capacity from the node to the sink node t is initially set to the sum of the communication costs of all incoming edges. Through iterative search for the s-t augmenting path in the auxiliary flow network G', calculate the minimum flow of the augmenting path and update the residual capacity. Dynamically partition communication-intensive microservice groups.

7. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 6, wherein The method for dynamically partitioning communication-intensive microservice groups is as follows: Based on breadth-first search, find a feasible path path from the virtual source node s to the sink node t in the auxiliary flow network G'. The generation condition of the feasible path path requires that the residual capacity cap(u, v) of all constituent edges (u, v) > 0. For each feasible path path found, calculate the bottleneck flow min_flow = min{cap(u, v)|(u, v) ∈ path}; in the traffic pushing stage, perform the capacity reduction operation of cap(u, v) - min_flow on all forward edges along the feasible path path, and at the same time perform the capacity compensation of cap(v, u) + min_flow on the reverse edges; in the implementation of microservice grouping, after each traffic push, recalculate the reachable node set reachable starting from the source node s through breadth-first search, and the intersection of the reachable nodes and the remaining nodes remaining_nodes gives the grouping S k Form the current microservice grouping with intensive communication until the maximum number of groupings K is reached or the node allocation is completed; Unallocated nodes are incorporated into the result set S as independent groups, and mutual exclusivity is ensured by merging overlapping groups. When the nodes remain connected in the auxiliary flow network G', there is still remaining capacity in the communication links between the nodes, so they are divided into the same group S k ; conversely, if they are disconnected due to traffic saturation, they will be assigned to different groups; among them, The method for merging and optimizing the overlapping groups to obtain the final grouping structure is as follows: Detect all the groups in the result set S that have common nodes, take the union of the groups with common nodes to form new groups, remove the atomic groups, and retain the newly merged groups. The finally output result set S satisfies the following conditions: 1) The communication intensity between nodes within a group is higher than that between groups; 2) All groups do not overlap; 3) The total number of groups ≤ K1, where K1 ≤ P / 2 and P is the number of edge servers.

8. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 6 or 7, characterized in that In the deep Q-network, the agent observes the system state s at each decision-making moment t t , the state s t includes the container load L r , the channel gain the edges between microservices, and generates a q-dimensional offloading action vector At = [a1, a2,..., aq] through the deep Q-network, where each action component corresponds to a deployment decision of a microservice packet between edge nodes, that is, an offloading strategy; Based on the ε-greedy strategy, independently select the target node with the largest Q value for each microservice group, and maintain the output dimension as q constantly through the virtual padding or grouping merging mechanism; drive the update of network parameters through the designed composite reward function. The obtained interaction data <s, a, r, s′> is stored in the experience replay pool. After there is enough data in the experience replay pool, a batch of data of size batch_size is randomly taken from the experience pool, and the prediction network is used to calculate the predicted value of Q: Encode the state s t The prediction network maps the encoded state to a q×|Z|-dimensional Q-value matrix through a three-layer neural network to predict the long-term benefits of each microservice group deployed to different edge nodes; where s and a represent the current state and action respectively, the action a is the microservice deployment decision generated based on the current state, and r and s′ represent the reward value of the current action and the new state at the next moment respectively; Calculate the Q-target value using the Q-Target network, based on the next state s t+1 and the reward r t , and use the target network parameters θ with delayed update to perform forward propagation to obtain maxQ(s t+1 , a|θ-), and then calculate the final target value in combination with the Bellman equation; calculate the loss function between the target value and the predicted value, use gradient descent to update the current network parameters, calculate the gradient of the loss function with respect to the current network parameters, and backpropagate and adjust the network weights according to the set learning rate to make the predicted value gradually approach the target value, thereby optimizing the strategy; after repeating several times, copy the parameters of the prediction network to the Q-Target network.

9. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 8, characterized in that The Q-Target network is denoted as a neural network that generates target values; the prediction network is denoted as Q(s,a|θ), which is a neural network used to predict the state-action value function and: Q q+1 (s t ,a t |θ) = Q q (s t ,a t |θ) + a q E q ; Among them, a q and γ represent the learning rate and the discount factor respectively, s' is the predicted state value after performing action a t in the q-th iteration, a' is the action with the largest reward value in state s', E q is the cumulative reward value in the q-th iteration during the iterative process, a t is the action selected by the agent at time t, r t is the immediate reward obtained after performing action a t at time t, θ, respectively represent the weight parameters of the prediction network and the weight parameters of the Q-Target network; The agent randomly samples a small portion of experiences \(e\) from the experience replay pool \(D\) t =\(\{e_1, e_2, \ldots, e\) t \}\) and uses the experience \(e\) t =(s t , a t , r t , s') to calculate the target value At the \(q\)-th step, the network parameters \(\theta\) are updated by minimizing the composite reward function to achieve the update of the offloading decision; the composite reward function is:

10. The reliability-aware dependent task offloading method based on topology reconstruction in industrial Internet edge computing according to claim 9, characterized in that, The implementation method of the deep Q-network is as follows: Create a replay buffer to store the experience data generated during the interaction between the agent and the environment, including the current state, the actions taken, the rewards obtained, and the next state. At the same time, set the maximum capacity of the replay buffer to K; Construct two neural networks - a prediction network and a Q-Target network. Initially, copy the weights of the prediction network to the Q-Target network so that they have the same initial parameters; Use M rounds of iterative training to optimize the dynamic offloading strategy. Combine the ε-greedy strategy with the microservice topology reconstruction result to generate the container deployment action a t : According to the ε-greedy strategy, the agent randomly selects an action with probability ∈ and selects the optimal action in the current state with probability 1 - ∈. Use the result of microservice topology reconstruction to combine the microservice grouping with the resource status of the edge node or the cloud to generate the specific container deployment action a t , when the agent executes the action a in the environment t and observes the reward r given by the environment t and the new state s t+1 , store the quadruple (φ t , a t , r t , φ t+1 ) into the replay buffer, where φ t is the preprocessed representation of the current state s t , and φ t+1 is the preprocessed representation of the new state s t+1 ; Update the Q-network parameter θ based on the temporal difference error, and minimize the weighted loss of the communication cost T ixy and the processing delay T ixj ; By randomly sampling a batch of experience samples from the replay buffer, calculate the target value of each sample, calculate the difference between the predicted value of the prediction network and the target value as the temporal difference error, and use the temporal difference error to update the Q-network parameters by gradient descent; Regularly synchronize the Q-Target network parameter θ- by periodically copying the parameters of the prediction network to the parameters of the Q-Target network; Output the optimal offloading strategy, the converged action value function, and the specific numerical values of the obtained performance metrics; Determine the optimal offloading strategy by evaluating the state-action value function of the prediction network after training is completed; Traverse all possible states s and actions a, and select the action that maximizes the value of the state-action value function as the optimal offloading strategy π; output the converged state-action value function as the action value function Q.

Citation Information

Cited By

  • Dynamic topology cooperative control method and system for heterogeneous equipment

    CN120811907A

  • Drift detection method for micro-service cluster and electronic equipment

    CN120935070A

  • Drift detection method for micro-service cluster and electronic device

    CN120935070B

  • Micro-service migration method, device and equipment and computer readable storage medium

    CN121644675A

  • Cooperative reasoning method and system for power edge intelligence, and electronic equipment

    CN122240209A