Task Offloading Optimization Method for Fixed-Path AGVs in the Industrial Internet of Things Environment

A two-phase method using load balancing and DQN reinforcement learning optimizes AGV task offloading in industrial IoT environments, addressing complexity and privacy issues to minimize task completion times.

CN114201303BActive Publication Date: 2025-07-15HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111539145.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-07-15
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In the industrial Internet of Things environment, the existing task offloading scheduling methods are insufficient in complexity and expansion applicability, especially the reinforcement learning algorithm is difficult to effectively converge to the optimal value in edge computing scenarios, and requires a lot of prior knowledge, which cannot meet the privacy protection requirements.

Method used

Using a combination of load balancing algorithm and deep reinforcement learning (DQN) method, AGV task offloading is optimized through a two-stage processing scheme. The first stage ignores resource conflicts, and the second stage optimizes resource conflict nodes through Markov model, using the experience pool to break sample correlation, and realizes the optimal offloading strategy.

Benefits of technology

Under the conditions of fixed paths and resource constraints, the AGV task completion time is minimized, which meets the privacy protection needs, and has good reusability and practical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114201303B_ABST
    Figure CN114201303B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimization method for fixed-path AGV task offloading in the industrial Internet of Things environment. Based on the traditional model-free reinforcement learning method with value function update, the present invention optimizes the task offloading scheduling problem in the scenario of AGV assisting edge computing. On this basis, a load balancing algorithm and an improved DQN algorithm are combined. Finally, under the constraints of task processing time sensitivity, path, efficiency, etc., the problem of the shortest optimal offloading processing time of multiple AGVs and multiple server nodes in the Internet of Things environment is realized. The method of the present invention does not require too much prior knowledge, and the task offloading transmission in the short distance meets the requirements of data security. Moreover, the present invention has good reusability in similar application scenarios, and the practical value of the invention is relatively strong.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of edge computing, and particularly relates to a reinforcement learning method for optimizing the task offloading of fixed-path AGVs in an industrial Internet of Things environment. Background Art

[0002] At present, a large number of sensors have been deployed in the industrial Internet of Things, and it is full of time-sensitive data operations and decision-making executions. The fully cloud-centric computing mode can no longer meet the actual needs. Therefore, the edge computing mode is applied to industrial scenarios, distributing data computing tasks from the centralized cloud to edge devices closer to the data source. By being closer to the target, the edge server provides cloud computing capabilities to users at close range, which excellently solves the demand for real-time processing in industrial Internet of Things applications. Among them, the problem of offloading scheduling optimization of tasks is the current research focus. The existing task scheduling methods and models include particle swarm optimization, genetic algorithms, ant colony optimization algorithms, and game theory, etc. These methods may have good performance in certain specific scenarios, but there is still much room for improvement in the complexity of algorithm design and extended applicability.

[0003] With the development of AI, various reinforcement learning algorithms have been proven to have significant advantages in solving sequential decision-making problems. They are very suitable for dealing with the policy selection problem of complex search spaces in edge computing scenarios, and themselves only require less prior knowledge to bring better solutions to problems, while also meeting the requirements of privacy protection. Reinforcement learning can be roughly divided into two categories: model-based reinforcement learning and model-free reinforcement learning. Due to the increasing emphasis on data security, it is difficult to obtain the relevant prior knowledge of detailed data of multiple user nodes. Therefore, model-free reinforcement learning is more suitable for solving the task offloading scheduling problem under edge computing. Model-based reinforcement learning can also be further divided into two major categories. One is the policy optimization method, which does not need to maintain a value function model, but directly searches for the optimal policy. It often adopts a parameterized policy and maximizes the expected return by updating this parameter. The other is the reinforcement learning method based on value function update, generally referring to the Q-Learning algorithm. Q is a historical experience memory table related to the current state and action selection, which can represent the cumulative expectation of the benefits that can be obtained by taking actions in a certain state at a certain moment. The Q-Learning algorithm constructs an agent representing the algorithm, places it in the Markov model of the problem to be solved, and selects whether to make a new action selection by querying the accumulated learning experience or randomly select an action through a search strategy. However, in such a situation, there are two problems that are difficult to solve: the correlation between samples is too strong and the learned target depends on the target itself, so it is difficult to converge to the optimal value. In addition, due to the overly large state space and action space of the target, it will also lead to overly complex implementation. To solve such problems, this paper introduces the DQN deep reinforcement network to reduce the correlation between samples and combines the load balancing algorithm to make the training converge better. Therefore, this method proposes a reinforcement learning method for drone-assisted multi-node task offloading scheduling based on value function update. Summary of the Invention

[0004] The object of the present invention is to solve the problem of minimizing the task completion time by offloading tasks based on a given path in the edge computing scenario of an AGV cart in the industrial Internet of Things.

[0005] The described edge computing scenario mainly includes multiple AGV vehicles and multiple edge servers. Among them, each AGV carries a given task to be processed and travels along a given route. When passing by an edge server, it unloads the carried task to the edge server for processing. Each AGV has a different route but may have intersections, that is, it passes through the same edge server. Each edge server has a given processing efficiency and maximum capacity. If the unloaded tasks reach the maximum capacity, new tasks will be rejected. When the AGV vehicle reaches the target end point, all the carried tasks need to be unloaded until all tasks are processed by the server, and the entire task unloading and processing process ends. In order to reasonably allocate the tasks carried by all AGVs and make the total processing time of the tasks carried by all AGVs as short as possible, the present invention proposes a method for optimizing the task unloading of fixed-path AGVs in the industrial Internet of Things environment. This method combines the advantages of the load balancing algorithm and the reinforcement learning algorithm, and only requires very little prior knowledge and some simple interactions with the environment during the driving process to obtain learning experience, and then a near-optimal solution for the task unloading and processing completion time of the AGV can be obtained.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a method for optimizing the task unloading of fixed-path AGVs in the industrial Internet of Things environment, including the following steps:

[0007] Step 1: Multiple AGVs respectively travel from the given starting point along the planned path. During the travel, they will pass through multiple edge servers for task unloading. When each AGV reaches the target point, all tasks are unloaded. When the edge server finishes processing all the unloaded tasks, the task process ends; model this scenario.

[0008] Step 2: In order to minimize the total time consumed by the entire task unloading and processing, the model is processed in two stages: In the first stage, the resource conflict problem caused by the competition of multiple AGVs is ignored, and the optimal unloading plan of each AGV at multiple edge servers is obtained; in the second stage, on the basis of the first stage, the nodes causing resource conflicts are optimized to solve the resource conflict problem, so as to achieve the overall optimal result.

[0009] Step 3: First, perform the first-stage processing on the task unloading of the AGV. Ignore the resource conflict problem caused by the unloading of multiple AGVs. The amount of tasks unloaded by each AGV at the passing edge server is related to the processing efficiency, capacity of the edge server, and the time to reach the edge server. The weighted round-robin load balancing algorithm is used to allocate resources to each AGV according to the above conditions.

[0010] Step 4. In the second stage, process the nodes that caused resource conflicts in the first stage. Build a Markov model for the edge server nodes and AGVs that caused the conflicts. Initialize the state space as the task volumes carried by multiple AGVs, the action space as the offloading volumes of AGVs at the edge server nodes, and the reward obtained by executing this action as the reciprocal of processing this task volume. Considering that the arrival times of each AGV at the edge server nodes are different, each AGV will have an offloading priority;

[0011] Step 5. Set the limiting conditions in the application scenario as small goals for reinforcement learning, and set the total time obtained after the final scheduling process to be as small as possible as the big goal, and the big goal is achieved based on the small goals;

[0012] Step 6. At the start of the training cycle of the reinforcement learning method, the agent starts from the starting state of the Markov model, selects the next action for the agent according to the improved policy. After the agent makes an action selection, it will reach the next environmental state, and the environmental state will give corresponding rewards according to the current features;

[0013] Step 7. During the training process of the task offloading of AGVs, by using an experience pool with a fixed capacity, make full use of the advantages of off-policy, thereby disrupting the sample correlation and improving the utilization rate of samples;

[0014] Step 8. Stop training when the maximum training cycle of the algorithm is reached, output the maximum cumulative reward of the training convergence, and obtain the optimal task offloading policy.

[0015] Furthermore, in the two-stage processing scheme, the state of the deep reinforcement learning in the second stage is represented by where the vector represents the task state of the AGV at the i-th common node; the action space is the task offloading volumes of each AGV at the resource conflict nodes.

[0016] Furthermore, the task completion time consists of three parts: the movement time of the AGV, the offloading time of the AGV to offload tasks to the edge server, and the processing time of the edge server. Among them, the time for the AGV to move from the starting point to the i-th edge server is:

[0017]

[0018] And the time for the AGV to offload tasks to the edge server is:

[0019]

[0020] where s i represents the task offloading rate:

[0021] s i= W log(1 + pg i / N0)

[0022] The processing time of the edge server is:

[0023]

[0024] From this, it can be deduced that the amount of tasks assigned by the weighted round-robin load balancing algorithm is:

[0025]

[0026] Where Task and AR are the initial allocation amount and the remaining allocation amount respectively:

[0027]

[0028]

[0029] Furthermore, the AGVs arrive at each edge server at different times. Considering the processing capabilities of the edge servers, the AGVs have different offloading priorities when choosing actions;

[0030] Therefore, the corresponding reward settings are as follows:

[0031]

[0032] Where represents the offloading priority of the i-th AGV at the edge server at time t; represents the offloading amount of the i-th AGV at the edge server at time t; E t represents the processing efficiency of this edge server.

[0033] Furthermore, the total reward obtained by the agent not only needs to consider the current reward but also the future long-term rewards. Moreover, the farther the time interval, the less accurate the obtained future reward value. Therefore, the cumulative reward can be expressed as:

[0034]

[0035] Where γ is the discount factor. Using the discount factor makes the reward value with a longer time interval account for a smaller proportion in the current reward.

[0036] Furthermore, during the process of the agent transitioning from one state to another in step six, a learning experience is generated in the experience pool, which includes the characteristics of the previous state, the selected action, the obtained reward, and the next state; the agent learns based on the existing past experiences in the experience pool. Finally, when the AGV finishes offloading and enters the end state, the next training cycle begins.

[0037] Furthermore, the AGV needs to comply with the following restrictions during the unloading scheduling process:

[0038] a) The node selected by the AGV for unloading must be an edge server node covered in the AGV's path;

[0039] b) The AGV can choose to unload or not unload when passing through an edge server node. When the tasks accepted by the server node reach the capacity limit, it will reject the AGV unloading task;

[0040] c) The tasks carried by the AGV must be completely unloaded when reaching the target point, and the time for all tasks carried by the AGV to be processed should also be as short as possible.

[0041] Advantages of the present invention:

[0042] In view of the characteristics of the task unloading scheduling scenario of AGV in the industrial Internet of Things environment, the present invention improves the original unloading scheduling algorithm, and innovatively proposes a load balancing algorithm combined with a reinforcement learning algorithm to form a new unloading method. Under the constraints of fixed AGV path, latency sensitivity, resource limitation, etc., the goal of minimizing the total task completion time is achieved by reasonably selecting the unloading scheme when the AGV path is fixed. The method of the present invention does not require too much prior knowledge and does not need to deeply understand the information of edge server nodes, which meets the requirements of privacy protection. Moreover, the present invention has good reusability in similar application scenarios and strong practical value. Description of the drawings

[0043] Figure 1 Schematic diagram of the AGV unloading model provided by the embodiment of the present invention;

[0044] Figure 2 Schematic diagram of the resource conflict situation provided by the embodiment of the present invention;

[0045] Figure 3 Flowchart of the DQN algorithm provided by the embodiment of the present invention. Detailed implementation manners

[0046] The method of the present invention will be further described below in conjunction with the drawings and embodiments.

[0047] A task unloading optimization method for a fixed-path AGV in an industrial Internet of Things environment is as follows:

[0048] Step 1: In the industrial Internet of Things environment, given multiple AGVs carrying several tasks to be processed, and given the travel routes of the AGVs, multiple edge servers are evenly distributed along the travel routes of the AGVs. The tasks carried by the AGVs need to be offloaded to the edge servers for processing. When an AGV reaches the given end point, all the tasks it carries need to be offloaded. When all the tasks are processed by the edge servers, the process ends; model this scenario.

[0049] Step 2: In order to make the time for all the tasks carried by multiple AGVs to be processed as short as possible, the present invention divides the task offloading into two stages for processing. In the first stage, a weighted round-robin load balancing algorithm is used to allocate to each AGV. Then, based on the first stage, in the second stage, the deep reinforcement learning DQN algorithm is used to train the stage that causes resource conflicts in the edge servers, so as to obtain the best allocation plan to minimize the final completion time of all tasks.

[0050] Step 3: In the first stage of the present invention, a weighted round-robin load balancing algorithm is used to perform load balancing allocation for each AGV. First, ignoring the influence of the offloading of other AGVs, considering the processing efficiency, capacity of the edge servers, and the time sequence of arrival at each edge server, the tasks carried by each AGV are subjected to load balancing allocation.

[0051] Step 4: The completion time of a task consists of three parts: the movement time of the AGV, the offloading time for the AGV to offload the task to the edge server, and the processing time of the edge server. Among them, the time for the AGV to travel from the starting point to the i-th edge server is:

[0052]

[0053] And the time for the AGV to offload the task to the edge server is:

[0054]

[0055] where s i represents the task offloading rate:

[0056] s i =Wlog(1 + pg i / N0)

[0057] The processing time of the edge server is:

[0058]

[0059] From this, the task volume allocated by the weighted round-robin algorithm can be deduced as:

[0060]

[0061] where Task and AR are the initial allocation amount and the remaining allocation amount respectively:

[0062]

[0063]

[0064] Step 5: In the second stage of the present invention, the Deep Q-Network (DQN) algorithm of deep reinforcement learning is used to train the nodes causing resource conflicts. When the optimal allocation of each AGV is obtained in the first stage, the result may be that the total offloading amount at the common edge server exceeds the capacity of the edge server, which will cause a resource conflict problem at this time. Therefore, it is necessary to adjust the offloading amount of the AGV at the common edge server. Therefore, in the second stage, the DQN algorithm of deep reinforcement learning is used to train this part to obtain the optimal allocation method.

[0065] At this time, a Markov model is constructed for the edge server nodes and AGVs causing conflicts. The initial state space is the task amounts carried by multiple AGVs, the action space is the offloading amount at a certain edge server node, and the reward obtained by executing this action is the reciprocal of the processing of this task amount. Considering that the arrival times of each AGV at the edge server node are different, each AGV will have an offloading priority;

[0066] Step 6: The offloading process of the AGV at each edge server corresponds to the process of state transition of the Markov model, and each state transition of the Markov model will generate a learning unit, including the previous state of the agent, the action selected in the previous state, the reward given by the environment for this state transition, and the current state. Since the ultimate goal is to minimize the task completion time, the reward is set as the reciprocal of the task processing time to maximize the reward. The reward function is:

[0067]

[0068] Adding a weakening factor: In addition, since the total return obtained by the agent needs to consider not only the current reward but also the future long-term rewards, and the farther the time interval is, the more inaccurate the obtained future reward value is, the cumulative return can be expressed as:

[0069]

[0070] where γ is the discount factor, and the discount factor is used to make the reward value with a longer time interval account for a smaller proportion in the current return. Since both the state space and the action space are large, the experience pool of DQN is used in the present invention for experience replay, making full use of the advantages of off-policy and being able to break the correlation between data. Its target is updated as:

[0071]

[0072] Step 7: When the algorithm reaches the maximum training cycle, stop the training and output the action sequence corresponding to the maximum result, which represents the unloading action selection of the AGV in the actual scenario, so as to obtain the optimal unloading sequence of each AGV at the common edge server, that is, the optimal unloading strategy for the task unloading scheduling of the AGV.

[0073] Embodiment:

[0074] Figure 1 Schematic diagram of the AGV unloading model provided by the example of the present invention;

[0075] Figure 2 Schematic diagram of the resource conflict situation provided by the example of the present invention;

[0076] A task unloading optimization method for a fixed-path AGV in an industrial Internet of Things environment is as follows:

[0077] Step 1: First, clarify the basic information of each AGV and the edge server, including the task volume carried by the AGV, the driving speed of the AGV, the processing efficiency and capacity of the edge server, and the driving route of the AGV, etc.

[0078] Step 2: Clarify the completion time of the task, including the driving time of the AGV to reach the edge server, the unloading time of the AGV to unload the task to the edge server, and the processing time of the edge server to process the task. They are respectively:

[0079]

[0080]

[0081]

[0082] Step 3: According to the processing efficiency and capacity of the edge server and the time when the AGV reaches each edge server as weights, use the weighted round-robin algorithm for load balancing allocation of the AGV's unloading at the edge server. Since the arrival times at each edge server are different, that is, the edge server nodes that arrive first are processed first. Therefore, this part of the task is the initial allocation task, and then reallocation is performed according to the processing efficiency weight.

[0083] Step 4: The initial allocation amount of the AGV to each edge server is:

[0084]

[0085] The remaining allocation amount is:

[0086] Therefore, the allocation obtained by using the weighted round-robin load balancing algorithm is as follows:

[0087]

[0088] Step 5: The weighted round-robin algorithm can obtain the optimal allocation scheme for each AGV. However, at the common edge server, if multiple AGVs unload according to the optimal allocation scheme, the unloading volume may exceed the capacity of the edge server. Therefore, this part needs to be optimized. The present invention uses the deep reinforcement learning DQN algorithm to train this part. Figure 3 It is the flowchart of the DQN algorithm provided by the example of the present invention.

[0089] Step 6: At this time, a Markov model is constructed, and the state of the AGV is initialized, including the task volume carried by the AGV and the location of the edge server where the AGV is located. The action is the unloading volume of the AGV at the common edge server. The reward function is the reciprocal of the processing time of the tasks unloaded by the action. Initialize the maximum number of training cycles, the size of the experience pool, and the weight parameter θ.

[0090] Step 7: Execute the policy, obtain a random number. If the random number is less than ε, randomly select a node among all nodes as the action selection for the current state. The agent executes the action to move from the current state to the next state, and obtains the reward R obtained by the edge server for processing this part of the tasks t+1 , after determining the learning unit reward for the state transition in this step, update the next state, and use the update formula to update the target, where γ is the discount factor. Since the agent needs to consider not only the current reward but also the future long-term rewards in the total return, but the farther the time interval is, the less accurate the obtained future reward value is. Therefore, the discount factor is used to make the proportion of the reward value with a longer time interval in the current return smaller. θ - indicates that it is updated slower than the weights of the Q network. Finally, the complete learning unit is pushed into the experience pool. The subsequent training can be carried out by randomly taking samples from the experience pool through experience replay.

[0091] Step 8: If the agent has not reached the end state, then repeat the above steps until the task is completed and enters the end state. If the agent unloads tasks exceeding the capacity of the edge server, then a penalty reward will be given to the agent. Calculate the cumulative revenue from the start state to the end state in the current cycle. If the experience pool is full or the revenue is greater than the current maximum target reward, then update the experience pool information.

[0092] Step 9: If the number of training cycles has not reached the maximum number of cycles, repeat the periodic training until the maximum training cycle is reached. If the maximum training cycle is reached, stop the training of the reinforcement learning method. According to the greedy algorithm, starting from the initial state, select the corresponding action with the maximum reward until the end state. Record all the action selections to obtain an offloading scheduling decision sequence, and output it as the solution to the problem of minimizing the task completion time of AGV under constraints in the industrial Internet of Things environment.

[0093] It should be noted that the parts not elaborated in detail in this specification all belong to the prior art. Those skilled in the art should understand that the above examples are only for helping readers understand the principles and implementation methods of the present invention, and the scope protected by the present invention is not limited to such examples. Any equivalent replacement made on the basis of the present invention is within the scope of the rights of the present invention.

Claims

1. A task offloading optimization method for fixed-path AGVs in the industrial Internet of Things environment, characterized in that, It includes the following steps: Step 1: Multiple AGVs travel from the given starting points along the pre-planned paths. During the travel, they will pass by multiple edge servers for task offloading. When each AGV reaches the target point, all tasks are offloaded. When the edge servers finish processing all offloaded tasks, the task process ends; model this scenario; Step 2: To minimize the total time consumed for the entire task offloading process, the model is processed in two stages: In the first stage, the resource conflict problem caused by the competition of multiple AGVs is ignored, and the optimal offloading plan for each AGV at multiple edge servers is obtained; In the second stage, based on the first stage, the nodes causing resource conflicts are optimized to solve the resource conflict problem, thereby achieving the overall optimal result; Step 3: First, perform the first-stage processing on the task offloading of AGVs; ignore the resource conflict problem caused by the offloading of multiple AGVs. The amount of tasks offloaded by each AGV at the passing edge servers is related to the processing efficiency, capacity of the edge servers, and the time to reach the edge servers; adopt the weighted round-robin load balancing algorithm to allocate resources to each AGV according to the above conditions; Step 4: In the second stage, process the nodes causing resource conflicts in the first stage; construct a Markov model for the edge server nodes and AGVs causing conflicts. The initial state space is the amount of tasks carried by multiple AGVs, the action space is the offloading amount of AGVs at the edge server nodes, and the reward obtained by executing this action is the reciprocal of the amount of tasks processed. Considering that each AGV reaches the edge server nodes at different times, each AGV will have an offloading priority; Step 5: Set the limiting conditions in the application scenario as small goals for reinforcement learning, and set the total time obtained after the final scheduling process to be as small as possible as the big goal, and the big goal is achieved based on the small goals; Step 6: At the start of the training cycle of the reinforcement learning method, the agent starts from the initial state of the Markov model and selects the next action for the agent according to the improved policy. After the agent makes an action selection, it will reach the next environmental state, and the environmental state will give corresponding rewards according to the current characteristics; Step 7: During the training process of the task offloading of AGVs, by using an experience pool with a fixed capacity, make full use of the advantages of off-policy, thereby disrupting the sample correlation and improving the sample utilization rate; Step 8: Stop training when the maximum training cycle of the algorithm is reached, output the maximum cumulative reward of the training convergence, and obtain the optimal task offloading policy.

2. The task offloading optimization method for fixed-path AGV in the industrial Internet of Things environment according to claim 1, characterized in that, In the two-stage processing solution, the state of the deep reinforcement learning in the second stage is represented by , where the vector represents the task state of the AGV at the i-th common node; The action space is the amount of tasks offloaded by each AGV at the resource conflict nodes.

3. The task offloading optimization method for a fixed-path AGV in an industrial Internet of Things environment according to claim 1, characterized in that, The completion time of the task consists of three parts: the moving time of the AGV, the offloading time for the AGV to offload tasks to the edge server, and the processing time of the edge server; among them, the time for the AGV to travel from the starting point to the i-th edge server is: Among them, L j,k represents the distance between the edge server ES j and the edge server ES k ; I m represents the starting point of the m-th AGV; ES i represents the i-th edge server; P m represents the path set of the m-th AGV; And the time for the AGV to offload tasks to the edge server is: Among them, TA m,i represents the task volume unloaded by the m-th AGV at the i-th ES passed through; where s i represents the offloading rate of the task: s i = Wlog(1 + pg i / N0) where W is the channel bandwidth, p is the transmission power of the AGV, and g i is the channel gain of the i-th edge server, and N0 is the noise power; The processing time of the edge server is: Among them, E i represents the processing efficiency of the i-th ES; From this, the amount of tasks allocated by the weighted round-robin load balancing algorithm is deduced as: Among them, represents the initial allocation amount of the i-th ES; AR i represents the remaining allocation amount of the i-th ES; Where Task and AR are the initial allocation amount and the remaining allocation amount respectively: Among them represents the theoretical maximum computing time for the i-th edge server to process tasks under full load; Cap i represents the maximum task capacity of the i-th edge server; represents the total computing efficiency of all edge servers on the m-th AGV path; AR m represents the remaining unassigned task volume of the m-th AGV; RC i represents the available resources of the i-th ES.

4. The task offloading optimization method for fixed-path AGVs in the industrial Internet of Things environment according to claim 1, wherein The arrival times of AGVs at each edge server are different. Considering the processing capabilities of edge servers, AGVs have different offloading priorities when choosing actions; Therefore, the corresponding reward settings are as follows: Among them represents the offloading priority of the $i$-th AGV at the edge server at time $t$; represents the offloading volume of the $i$-th AGV at the edge server at time $t$; $E$ t represents the processing efficiency of this edge server.

5. The task offloading optimization method for a fixed-path AGV in an industrial Internet of Things environment according to claim 4, characterized in that, The total reward obtained by the agent should not only consider the current reward but also the future long-term rewards. Moreover, the farther the time interval, the less accurate the obtained future reward value. Therefore, the cumulative reward can be expressed as: where γ is the discount factor. Using the discount factor makes the reward value with a longer time interval account for a smaller proportion in the current reward.

6. The task offloading optimization method for fixed-path AGVs in the industrial Internet of Things environment according to claim 1, characterized in that In Step 6, when the agent moves from one state to another, a learning experience is generated in the experience pool. It includes the features of the previous state, the selected action, the obtained reward, and the next state. The agent learns based on the existing past experiences in the experience pool. Finally, the AGV completes the offloading and enters the end state, starting the next training cycle.

7. The task offloading optimization method for a fixed-path AGV in an industrial Internet of Things environment according to any one of claims 1-6, characterized in that, AGV needs to abide by the following restrictions during the offloading scheduling process: a) The node selected by AGV for offloading must be an edge server node covered in the AGV path; b) AGV can choose to offload or not when passing through an edge server node. When the tasks accepted by the server node reach the capacity limit, it will reject the AGV offloading task; c) The tasks carried by AGV must be completely offloaded when reaching the target point, and the time for all tasks carried by AGV to be processed should also be as short as possible.