An edge collaborative reasoning method and system for industrial inspection
Patent Information
- Application Number
- CN202611266530.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-25
AI Technical Summary
然而,视频传感设备受限于计算能力,无法单独完成复杂深度神经网络(Deep Neural Networks, DNN)推理,传统方案往往将视频数据上传至集中式服务器进行处理,这不仅占用大量网络带宽,而且存在传输延迟,难以满足工业巡检任务的低延迟要求
[0114]1、本发明的边端协同推理方法将DNN分割、无线带宽分配和边缘服务器动态优先级调度进行统一建模与联合优化,使边缘服务器队列状态、任务剩余时限和资源分配过程能够协同作用,从而降低关键巡检任务的等待时延和整体端到端处理时延,提高工业巡检任务的按时完成率和异常检测实时性;
Smart Images

Figure CN122824779A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of edge computing and artificial intelligence, and in particular to an edge-to-edge collaborative reasoning method and system for industrial inspection. Background Technology
[0002] With the development of industrial automation and intelligent manufacturing, the demand for real-time video analytics in industrial equipment inspection and production line monitoring is increasing. Industrial inspection tasks typically require the continuous acquisition of high-resolution video data and deep learning inference of equipment status to promptly detect anomalies. However, video sensing devices are limited by computing power and cannot independently perform complex deep neural network (DNN) inference. Traditional solutions often upload video data to a centralized server for processing, which not only consumes a large amount of network bandwidth but also introduces transmission latency, making it difficult to meet the low-latency requirements of industrial inspection tasks. In environments with multiple tasks and high-density equipment, task latency may lead to the failure to detect critical anomalies in a timely manner, seriously affecting production safety and efficiency.
[0003] While existing edge collaborative reasoning methods can offload some computing tasks to edge servers (ES), most methods do not fully consider the differences in task priorities, the dynamic changes in edge server queue status, and the interaction between local computing, wireless transmission, and edge computing loads, making it difficult to achieve efficient and reliable industrial inspection reasoning. Summary of the Invention
[0004] Purpose of the Invention: The purpose of this invention is to provide an edge-to-edge collaborative inference method and system for industrial inspection. In the process of continuous generation of inspection video inference tasks by multiple video sensing devices (VSDs), the coupling relationship between DNN segmentation, wireless bandwidth competition and dynamic priority scheduling of edge servers is comprehensively considered. The task processing flow and resource allocation method are dynamically optimized to reduce the end-to-end processing latency of industrial inspection tasks and improve the on-time completion rate and real-time performance of anomaly detection.
[0005] Technical solution: An edge-end collaborative reasoning method for industrial inspection, comprising the following steps:
[0006] S1. Based on the processing requirements of VSD video data collected in industrial inspection scenarios, an inspection inference task model is established, and DNN segmentation points are combined to represent the task generation time, latency tolerance, local computing load, uploaded data volume and edge computing load.
[0007] S2, based on the wireless communication process between VSD and ES, establish a wireless transmission model, divide the total system bandwidth into allocable bandwidth units, and calculate the uplink transmission rate and transmission delay of the inspection inference task.
[0008] S3, based on the continuous arrival process of inspection reasoning tasks, VSD local computing queue, ES edge computing queue, task latency tolerance, and edge-end segmentation execution process, constructs an edge-end collaborative reasoning model;
[0009] S4. Based on the low latency requirements of industrial inspection tasks, a joint optimization objective function is constructed by combining the edge-end collaborative reasoning model. By dividing the time slots, the optimization problem of solving the objective function is transformed into a partially observable Markov decision process, and a Markov model is constructed.
[0010] S5. A multi-agent proximal policy optimization algorithm enhanced by gated recurrent units is adopted to construct a GRU-MAPPO framework including VSD agent and ES agent. The Markov model is used to train each agent and update the current policy. When the preset conditions are met, the optimal policy is output.
[0011] S6 obtains the optimal global bandwidth allocation and ES dynamic priority adjustment strategy based on the optimal strategy, realizing low-latency processing of industrial inspection tasks in the edge-end collaborative reasoning process.
[0012] Furthermore, the specific steps for establishing the inspection reasoning task model include:
[0013] S11, Define the VSD collection as , Indicates the total number of VSDs;
[0014] S12 divides the system operation process into There are 1 time slot, and the time slot set is 1 Each time slot is [length missing] Then the first The start time of each time slot for:
[0015] ,
[0016] Set VSD to a fixed frame rate If an inspection inference task is generated, then the time slot length satisfies the following condition: ;
[0017] S13, let the nth... In the time slot Generated inspection reasoning task Represented as:
[0018] ,
[0019] in, Indicates the task latency tolerance. This indicates the inspection reasoning task. The total number of logical layers that can be divided into in a DNN. Indicates the time when the task was generated;
[0020] definition The set of deployed DNN logical layers is represented as , No. Each logical layer is represented as ,in This indicates the amount of feature data output by this logical layer. This indicates the computational complexity of the logic layer.
[0021] S14, for inspection reasoning tasks Set the dividing point Indicates in The number of logic layers executed locally is the set of local execution layers. and ES execution layer collection They are respectively:
[0022] ,
[0023] ;
[0024] Based on the dividing point Computational inspection reasoning task exist The compute load on the local and Elasticsearch sides are as follows:
[0025] ,
[0026] ,
[0027] in, Indicates the inspection reasoning task exist Local computing load, Indicates the inspection reasoning task Computational load on the ES side;
[0028] Then the inspection reasoning task Uploaded data volume From the dividing point A decision is expressed as:
[0029] ,
[0030] in, Indicates the dividing point The amount of input data; at that time , This indicates the amount of original input data.
[0031] Furthermore, the specific steps for establishing the wireless transmission model are as follows:
[0032] S21, Suppose the total uplink bandwidth of the system is divided into multiple bandwidth units. In the time slot The number of bandwidth units obtained is expressed as ,satisfy:
[0033] ,
[0034] in, Indicates the total number of bandwidth units in the system;
[0035] S22, let the bandwidth of each bandwidth unit be... In the time slot Any consecutive moments within place, Uplink transmission rate for:
[0036] ,
[0037] in, express The transmission power, express Uplink channel gain with base station Represents the noise power spectral density; when At that time, it was agreed ;
[0038] S23, based on the uplink transmission rate Inspection and reasoning tasks Transmission delay satisfy:
[0039] ,
[0040] in, Indicates the inspection reasoning task The moment the upload begins.
[0041] Furthermore, the specific steps for constructing the edge-to-edge collaborative reasoning model are as follows:
[0042] S31, establish a local inference queue and a transmission queue on the VSD side, and establish an edge inference queue on the ES side;
[0043] S32, After the inspection reasoning task is generated, it enters the local reasoning queue, which is processed according to the first-come-first-served rule.
[0044] Inspection and reasoning task Local inference start time With completion time They are respectively:
[0045] ,
[0046] ,
[0047] in, Indicates the inspection reasoning task The generation time, Indicates the inspection reasoning task The local inference completion time, express Computational power;
[0048] S33, after the inspection reasoning task completes local reasoning, it enters the transmission queue, and its transmission begins at... With completion time They are respectively:
[0049] ,
[0050] ,
[0051] in, Indicates the inspection reasoning task The time when the transmission is completed, For inspection reasoning tasks The transmission delay satisfies:
[0052] ;
[0053] S34, after the inspection and inference task is uploaded, it enters the edge inference queue; the edge inference queue adopts a non-preemptive dynamic priority scheduling mechanism, and once a task starts serving, it is not interrupted. When the edge server is idle, it selects the task with the highest dynamic priority from the waiting queue for service; inspection and inference task In the time slot Dynamic priority Defined as:
[0054] ,
[0055] in, Indicates time slot Priority factor, Indicates the task latency tolerance. Indicates time slot The starting time, Indicates the inspection reasoning task The time when the transmission is completed, Indicates the inspection reasoning task The moment of generation;
[0056] Set up an inspection reasoning task The starting time of edge reasoning is Then the edge reasoning completion time is:
[0057] ,
[0058] in, This indicates the computing power of Elasticsearch.
[0059] Furthermore, the specific steps for constructing the joint optimization objective function and the Markov model include:
[0060] S41, based on the inspection reasoning task End-to-end processing latency and normalized end-to-end delay Define the on-time completion rate of long-term tasks in the system. and the normalized average end-to-end latency for timely task completion The expression is as follows:
[0061] ,
[0062] ,
[0063] ,
[0064] ,
[0065] in, Indicates that the task was completed on time. hour, ;otherwise ; This represents a very small positive number introduced to prevent the denominator from being zero;
[0066] S42, with the goal of maximizing the on-time completion rate of tasks and minimizing the average end-to-end latency, construct a joint optimization objective function:
[0067] ,
[0068] in, This is a tradeoff factor used to balance the on-time completion rate of tasks with the normalized average end-to-end latency; , , They represent time slots respectively. The DNN makes decisions on segmentation points, bandwidth allocation, and priority adjustment within the system; constraint C1 limits the range of segmentation points selected by each VSD in each time slot; constraint C2 ensures that the bandwidth allocation result is a non-negative integer; constraint C3 ensures that the total bandwidth allocated to each VSD in the same time slot is equal to the total system bandwidth; constraint C4 limits the range of values for the priority adjustment parameter.
[0069] S43, if the edge-to-edge collaborative reasoning process is modeled as a partially observable Markov decision process, then the Markov model... Represented as:
[0070] ,
[0071] in, A set of intelligent agents, numbered to The intelligent agents correspond to each Number The intelligent agent corresponds to ES; Represents the global state space; , Representing intelligent agents respectively The local observation space and action space, , ; Represents the state transition probability. Represents the reward function, Indicates the discount factor;
[0072] S44, construct the global state space and the local observation space;
[0073] The global state space , time slot global state Represented as:
[0074] ,
[0075] in, Indicates ES in time slot The remaining computational load of the local inference queue at the start time. express In the time slot The remaining computational load of the local inference queue at the start time. express In the time slot The amount of data remaining in the transmission queue at the start time. express In the time slot Channel gain, This indicates the on-time completion rate of the task in the previous time slot. Indicates the transmission efficiency of the previous time slot;
[0076] The local observation space In the time slot Local observation include In the time slot Local observation and ES in time slots Local observation , respectively represented as:
[0077] ,
[0078] ,
[0079] in, express computing power This indicates the computing power of Elasticsearch (ES).
[0080] S45 constructs the action space of the intelligent agent;
[0081] Action space , Represents intelligent agents The action; Actions of the intelligent agent If we select a split point for the current task, then we have: Actions of ES agents Let the normalized bandwidth allocation vector and the normalized priority adjustment vector be, then we have ,in And satisfy Normalized priority vector And satisfy ; Indicates time slot Internal allocation The number of normalized bandwidth units, express The generated inference task is in the time slot Normalized priority factor;
[0082] S46, Construct a multi-dimensional reward function consisting of team rewards and individual rewards;
[0083] Team Rewards Defined as:
[0084] ,
[0085] in, This represents the on-time completion rate of tasks in time slot t. This represents the normalized average end-to-end delay for tasks to be completed on time within time slot t.
[0086] Individual rewards and ES individual rewards They are defined as follows:
[0087] ,
[0088] ,
[0089] in, , They represent Compared with the individual reward weight of ES, , They represent In the time slot Within the task completion rate and normalized average end-to-end latency, , They represent time slots respectively. The transmission efficiency and service efficiency of ES queues;
[0090] Total Rewards Total ES Rewards The function expressions are as follows:
[0091] ,
[0092] ,
[0093] in, and All are weighting coefficients.
[0094] Furthermore, the GRU-MAPPO framework includes a GRU module, an Actor network, and a Critic network. The Actor network is used for distributed execution of generation strategies and has a total of [number missing]. Each agent has its own output strategy; the Critic network is used for centralized training to evaluate and update the output strategies of the Actor network.
[0095] The GRU module is used to capture the effects of queue evolution, channel changes, and cross-time slot decision-making, and is positioned after the feature encoder of each agent; for any agent... Its cyclic hidden state is updated as follows:
[0096] ,
[0097] in, Indicates the observation encoder, Indicates the hidden state of the previous time slot cycle;
[0098] The Actor network is used to generate corresponding execution strategies based on the local observations of each agent; for each agent... Actor Network The input is its time slot Local observation The output of the Actor network is the agent. Action probability distribution And determine the agent based on the probability distribution of the action. Actions in the current time slot ;
[0099] The Critic network is used during the training phase to evaluate the long-term returns of each agent under the current joint policy, providing a value benchmark for policy updates; for agents... Input to the Critic network From global state With agent identity coding obtained by piecing together Among them, global state Used to characterize the overall operating status of the current system, identity code One-hot encoding is used to represent the input; based on this input The Critic network outputs an intelligent agent. The state value estimate under the current joint strategy is used to approximate its expected discounted return:
[0100] ,
[0101] in, Represents intelligent agents Expected discount return under the current joint strategy The discount factor representing the reward. Represents intelligent agents In the time slot The corresponding rewards; This represents the expectation function of the strategy.
[0102] Furthermore, the specific steps for training each agent include:
[0103] SA1 sets hyperparameters, initializes the VSD agent Actor network, ES agent Actor network, and Critic network, and initializes the cyclic hidden state. and trajectory buffer D;
[0104] SA2 initializes environment parameters at the start of each training round, including time slots. Initialize to 0, initialize each queue state to the preset value, reset the loop hidden state, and clear the trajectory buffer D;
[0105] SA3, acquires local observations of the VSD agent in the current time slot. The VSD agent is based on current local observations and the current Actor network Make decisions and determine the break-off points for the current task. And set the reasoning division method between the terminal side and the edge side for the task in the current time slot;
[0106] Obtain local observations of the ES agent in the current time slot The ES agent is based on current local observations and the current Actor network Make decisions and determine bandwidth allocation parameters. and priority adjustment parameters And set the uplink bandwidth resources and ES queue scheduling order for each VSD in the current time slot;
[0107] SA4, based on time slot length The system updates the environmental state at the scale, enabling the generation, local inference, uplink transmission, and edge inference of inspection inference tasks, and statistically analyzes the task completion status, failure status, and queue changes of the current time slot.
[0108] SA5, the VSD agent calculates the reward value for the current time slot based on the corresponding reward function. The ES agent calculates the reward value for the current time slot based on the corresponding reward function. ;
[0109] SA6 stores the current time slot's state, action, reward, and the next time slot's state into the trajectory buffer D. Once the number of samples in trajectory buffer D reaches a preset limit, the stored samples are used to test the VSD agent Actor network. ES intelligent agent Actor network and Critic Network Train the network and update the current network weights.
[0110] An edge-end collaborative reasoning system for industrial inspection scenarios is provided to execute any of the aforementioned edge-end collaborative reasoning methods for industrial inspection. The system dynamically determines the DNN segmentation point based on the task status, the computational complexity of the DNN logic layer, the amount of output feature data of the DNN logic layer, the channel status, and the queue status, so that the local computing load, the amount of uploaded data, and the edge computing load can be adjusted according to the system status. The system includes multiple VSDs, one ES, and a base station connected to the ES.
[0111] VSD is used to continuously acquire video data related to industrial field equipment, production line status or abnormal events, and execute the first few logical layers of the DNN model locally.
[0112] ES is used to receive intermediate feature data uploaded from each VSD, maintain the edge inference queue, and complete the inference of the remaining logical layers of the DNN model. Among them, ES adopts a non-preemptive dynamic priority scheduling mechanism, which determines the task scheduling order based on task latency tolerance, task generation time, upload completion time, and priority factor.
[0113] Compared with the prior art, the significant advantages of this invention are as follows:
[0114] 1. The edge-to-edge collaborative reasoning method of the present invention performs unified modeling and joint optimization of DNN segmentation, wireless bandwidth allocation and edge server dynamic priority scheduling, so that the edge server queue status, task remaining time limit and resource allocation process can work together to reduce the waiting latency of key inspection tasks and the overall end-to-end processing latency, and improve the on-time completion rate of industrial inspection tasks and the real-time performance of anomaly detection.
[0115] 2. In the edge-end collaborative reasoning method of the present invention, the collaborative reasoning process of the industrial inspection end is modeled as a partially observable Markov decision process, and a dual-type multi-agent reinforcement learning method is further proposed. Through the collaboration between video terminal equipment and edge server, the channel state, queue state and task state are dynamically adapted, and the deep neural network segmentation point, bandwidth allocation parameters and priority adjustment parameters are autonomously determined. The current strategy is continuously updated by training the agent, thereby reducing the processing latency of industrial inspection tasks and improving the system's adaptive capability.
[0116] 3. The edge-to-edge collaborative inference system of this invention combines edge inference with queue scheduling optimization to meet the low-latency requirements of industrial inspection scenarios. Under the simulation conditions of this embodiment, the edge-to-edge collaborative inference method of this invention has a more stable convergence effect compared with benchmark algorithms such as GRU-IPPO, GRU-MATD3, GRU-WQMIX and MLP-MAPPO. Compared with traditional scheduling methods such as FCFS, SJF and EDF, the average end-to-end latency is reduced by about 10%. Moreover, it can still maintain good task completion rate and latency performance under different DNN model load conditions, indicating that the edge-to-edge collaborative inference system of this invention can improve the low-latency inference capability and adaptability in multi-device concurrent scenarios. Attached Figure Description
[0117] Figure 1 This is a schematic diagram of the edge-end collaborative reasoning system architecture for industrial inspection scenarios according to the present invention;
[0118] Figure 2 This is a flowchart of the edge-to-edge collaborative reasoning method constructed in this invention;
[0119] Figure 3 This is a schematic diagram of the GRU-MAPPO heterogeneous collaborative scheduling framework constructed in this invention;
[0120] Figure 4 This is a comparison chart of cumulative reward convergence during the training process of different algorithms;
[0121] Figure 5 This is a comparison chart showing the on-time completion rate of tasks as a function of video frame rate under a dynamic queue-aware scheduling mechanism.
[0122] Figure 6 This is a comparison chart showing the average end-to-end latency as a function of video frame rate under a dynamic queue-aware scheduling mechanism.
[0123] Figure 7 This is a comparison chart of the on-time completion rate of tasks under different DNN model loads;
[0124] Figure 8 This is a comparison chart of average end-to-end latency under different DNN model loads. Detailed Implementation
[0125] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0126] like Figure 1 As shown, this invention presents an edge-to-edge collaborative inference system for industrial inspection scenarios, based on multiple Video Detectors (VSDs), one Edge Server (ES), and a base station connected to the ES. The VSDs continuously collect video data related to industrial equipment, production line status, or abnormal events, and execute the first few logical layers of the DNN model locally. The ES receives intermediate feature data uploaded by each VSD, maintains the edge inference queue, and completes the inference of the remaining logical layers of the DNN model. The ES employs a non-preemptive dynamic priority scheduling mechanism, determining the task scheduling order based on task latency tolerance, task generation time, upload completion time, and priority factors. Because only intermediate feature data corresponding to the segmentation points is transmitted between the edges, the bandwidth consumption caused by uploading the original video is reduced, and the inference latency is lowered.
[0127] The edge-to-edge collaborative inference system of the present invention dynamically determines the DNN segmentation point based on the task status, the computational complexity of the DNN logic layer, the amount of output feature data of the DNN logic layer, the channel status, and the queue status, so that the local computing load, the amount of uploaded data, and the edge computing load can be adjusted according to the system status.
[0128] like Figure 2As shown, the edge-end collaborative reasoning method of the present invention includes constructing an inspection reasoning task model, a wireless transmission model, an edge-end collaborative reasoning model, a joint optimization objective and a Markov model, and an agent training process based on GRU-MAPPO. The specific implementation steps are as follows:
[0129] Step 1: Based on the processing requirements of VSD-acquired video data in industrial inspection scenarios, establish an inspection inference task model, and combine DNN segmentation points (hereinafter referred to as segmentation points) to characterize the task generation time, latency tolerance, local computing load of the DNN logic layer, uploaded data volume, and edge computing load. The specific steps are as follows:
[0130] Step 11, define the VSD collection as , This indicates the total number of VSDs.
[0131] Step 12: Divide the operation process of the edge-to-edge collaborative reasoning system into time slot sets. , This represents the total number of time slots, with each time slot having a length of [length value missing]. Then the first The start time of each time slot for:
[0132] ,
[0133] Set VSD to a fixed frame rate If an inspection inference task is generated, then the time slot length satisfies the following condition: .
[0134] Step 13, let the nth... In the time slot Generated inspection reasoning task Represented as:
[0135] ,
[0136] in, Indicates the task latency tolerance. This indicates the inspection reasoning task. The total number of logical layers that can be divided into in a DNN. Indicates the time when the task was generated;
[0137] To adapt to chained DNN models and DAG (Directed Acyclic Graph) models containing residual connections or parallel branches, this invention uses logical layers as the unit of DNN partitioning. The set of DNN logical layers is represented as follows: ;No. Each logical layer is represented as ,in This indicates the amount of feature data output by this logical layer. This indicates the computational complexity of the logic layer.
[0138] Step 14, for the inspection reasoning task Set the dividing point Indicates in The number of logic layers executed locally is the set of local execution layers. and ES execution layer collection They are respectively:
[0139] ,
[0140] ,
[0141] Step 15, based on the dividing point Computational inspection reasoning task exist The compute load on the local and Elasticsearch sides are as follows:
[0142] ,
[0143] ,
[0144] in, Indicates the inspection reasoning task exist Local computing load, Indicates the inspection reasoning task Computational load on the ES side;
[0145] Inspection and reasoning task Uploaded data volume From the dividing point A decision is expressed as:
[0146] ,
[0147] in, Indicates the dividing point The amount of input data; at that time , This indicates the amount of original input data.
[0148] Step 2: Based on the wireless communication process between VSD and ES, establish a wireless transmission model, divide the total system bandwidth into allocable bandwidth units, and calculate the task upload rate and transmission latency. The specific steps are as follows:
[0149] Step 21: In this embodiment, the video sensing device and the edge server use OFDMA (Orthogonal Frequency Division Multiple Access) uplink transmission. Time slots are defined. Internal allocation The number of bandwidth units is The bandwidth of each bandwidth unit is The total number of available bandwidth units in the system is Then the bandwidth allocation satisfies ;
[0150] Step 22, in the time slot Any consecutive moments within place, Uplink transmission rate for:
[0151] ,
[0152] in, express The transmission power, express Uplink channel gain with base station Represents the noise power spectral density; when At that time, it was agreed ;
[0153] Step 23, based on the uplink transmission rate Inspection and reasoning tasks Transmission delay satisfy:
[0154]
[0155] in, Indicates the inspection reasoning task The moment the upload begins Indicates the inspection reasoning task Uploaded data volume.
[0156] Step 3: Based on the process of local inference on VSD, uploading intermediate data, and completing the remaining inference in ES according to the inspection inference task, construct the edge-device collaborative inference model; the specific steps are as follows:
[0157] Step 31: Establish a local inference queue and a transmission queue on the VSD side, and establish an edge inference queue on the ES side.
[0158] Step 32: After the inspection reasoning task is generated, it enters the local reasoning queue, which is processed according to the first-come, first-served rule.
[0159] Inspection and reasoning task Local inference start time With completion time They are respectively:
[0160] ,
[0161] ,
[0162] in, Indicates the inspection reasoning task The generation time, Indicates the inspection reasoning task The local inference completion time, express computing power Indicates the inspection reasoning task exist The computational load.
[0163] Step 33: After completing local inference, the inspection inference task enters the transmission queue, and its transmission begins at [time missing]. With completion time They are respectively:
[0164] ,
[0165] ,
[0166] in, Indicates the inspection reasoning task The time when the transmission is completed, For inspection reasoning tasks The transmission latency, since the upload process may span multiple time slots, satisfies:
[0167] .
[0168] Step 34: After the inspection inference task is uploaded, it enters the edge inference queue. The edge inference queue adopts a non-preemptive dynamic priority scheduling mechanism. Once a task starts serving, it is not interrupted. When the edge server is idle, it selects the task with the highest dynamic priority from the waiting queue for service. Inspection Inference Task In the time slot Dynamic priority Defined as:
[0169] ,
[0170] in, express The generated inference task is in the time slot Priority factor, Indicates the task latency tolerance. Indicates time slot The starting time; Indicates the inspection reasoning task The time when the transmission is completed, Indicates the inspection reasoning task The generation time.
[0171] Set up an inspection reasoning task The starting time of edge reasoning is , If we represent the computational power of Elasticsearch, then the edge inference completion time is:
[0172] ,
[0173] in, Indicates the inspection reasoning task The computational load in ES.
[0174] Step 4: Based on the low-latency requirements of industrial inspection tasks, a joint optimization objective function is constructed using an edge-end collaborative inference model. Through time slot partitioning, the optimization problem of solving the objective function is transformed into a Partially Observable Markov Decision Process (POMDP), and a Markov model is constructed. The specific steps are as follows:
[0175] Step 41, Set up the inspection reasoning task End-to-end processing latency and normalized end-to-end delay They are respectively:
[0176] ,
[0177] ,
[0178] Task completion on time indicator variable Defined as:
[0179] ,
[0180] The system's long-term task on-time completion rate and the normalized average end-to-end latency for timely task completion They are respectively:
[0181] ,
[0182] ,
[0183] in, Indicates that the task was completed on time. hour, ;otherwise ; This represents a very small positive number introduced to prevent the denominator from being zero.
[0184] Step 42: At the beginning of each time slot, the system jointly optimizes the DNN segmentation point, uplink bandwidth allocation, and priority adjustment parameters. The decision variables are represented as follows: , , With the goal of maximizing the on-time completion rate of tasks and minimizing the average end-to-end latency, the constructed long-term joint optimization objective function is expressed as:
[0185] ,
[0186] in, This is a weighting factor, or team reward weight, used to balance the on-time completion rate of tasks with the normalized average end-to-end latency. , , They represent time slots respectively. The DNN makes decisions on split point selection, bandwidth allocation, and priority adjustment within the system. Constraint C1 limits the range of split points selected by each VSD in each time slot; constraint C2 ensures that the bandwidth allocation result is a non-negative integer; constraint C3 ensures that the total bandwidth allocated to each VSD in the same time slot is equal to the total system bandwidth; and constraint C4 limits the range of values for the priority adjustment parameters.
[0187] Step 43: Since the video sensing device and the edge server have different observation information, action space, and control responsibilities, the edge-device collaborative reasoning process is modeled as a partially observable Markov Decision Process (POMDP). The Markov model... Represented as:
[0188] ,
[0189] in, A set of intelligent agents, numbered to The agents correspond to each VSD, numbered as follows: The intelligent agent corresponds to ES; Represents the global state space. , Indicates time slot The global state; , Representing intelligent agents respectively The local observation space and action space, In the time slot Local observation include In the time slot Local observation and ES in time slots Local observation , ; Represents the state transition probability. Represents the reward function, This represents the discount factor.
[0190] Step 44: Construct the global state space and the local observation space;
[0191] Time slot The global state is represented as:
[0192] ,
[0193] in, Indicates ES in time slot The remaining computational load of the inference queue at the start time. express In the time slot The remaining computational load of the local inference queue at the start time. express In the time slot The amount of data remaining in the transmission queue at the start time. express In the time slot Channel gain, This indicates the on-time completion rate of the task in the previous time slot. Indicates the transmission efficiency of the previous time slot;
[0194] intelligent agents in time slots Local observation and ES agents in time slots Local observation They are respectively:
[0195] ,
[0196] ;
[0197] Step 45: Construct the action space of the intelligent agent;
[0198] Action space , Indicates in time slot intelligent agent The action; Actions of the intelligent agent Select the split point for the current task, i.e. Actions of ES agents The normalized bandwidth allocation vector and the normalized priority adjustment vector are, respectively. ,in And satisfy Normalized priority vector And satisfy ; Indicates time slot Internal allocation The number of normalized bandwidth units, express The generated inference task is in the time slot The normalized priority factor.
[0199] Step 46: Construct a multidimensional reward function, which consists of team rewards and individual rewards;
[0200] Team Rewards Defined as:
[0201] ,
[0202] in, This represents the on-time completion rate of tasks in time slot t. It represents the normalized average end-to-end delay for completing a task on time within time slot t.
[0203] Individual rewards and ES individual rewards They are defined as follows:
[0204] ,
[0205] ,
[0206] in, , They represent Compared with the individual reward weight of ES, , They represent In the time slot Within the task completion rate and normalized average end-to-end latency, , They represent time slots respectively. The transmission efficiency and service efficiency of ES queues;
[0207] Total Rewards Total ES Rewards The function expressions are as follows:
[0208] ,
[0209] ,
[0210] in, and All are weighting coefficients;
[0211] Step 5, as follows Figure 3 As shown, the Gated Recurrent Unit-enhanced Multi-Agent Proximal Policy Optimization (GRU-MAPPO) algorithm is used to establish VSD agents and ES agents respectively. Each agent is trained and its current policy is updated based on a Markov model. When preset conditions are met, the optimal policy is output. The specific steps are as follows:
[0212] Step 51: Establish the GRU-MAPPO training framework.
[0213] The GRU-MAPPO network architecture includes a Gated Recurrent Unit (GRU) module, an enhanced Actor network, and a Critic network. The Actor network is used for distributed execution of the generation strategy and has a total of [number missing]. Each agent has a corresponding output strategy. The Critic network is used for centralized training to evaluate and update the output strategies of the Actor network.
[0214] The framework includes Actor Network of Intelligent Agents Actor Network of ES Intelligent Agents Critic Network and trajectory buffer D. Actor network Used to output split points Actor Network Used for output bandwidth scaling vector and priority adjustment vector Centralized Critic Network Used to evaluate the long-term benefits of each agent based on the global state during the training phase; Figure 3 In this context, MLP stands for Multilayer Perceptron. This represents the agent's current strategy. This represents the old strategy used by intelligent agents when collecting trajectory samples.
[0215] The GRU module, located after the feature encoder of each agent, is used to capture the effects of queue evolution, channel changes, and cross-time slot decisions. For any agent... Its cyclic hidden state is updated as follows:
[0216]
[0217] in, Indicates the observation encoder, Represents intelligent agents The previous time slot cycle was in a hidden state;
[0218] Actor networks are used to generate corresponding execution policies based on the local observations of each agent. For each agent... Actor Network The input is its time slot Local observation This local observation reflects the intelligent agent Local state information can be directly obtained. During the execution phase, each agent does not need to access the complete global state, but instead independently generates action policies based on its own local observations. The output of the Actor network is the agent's... Action probability distribution And determine the agent based on the probability distribution of the action. Actions in the current time slot .for Intelligent agent, actions at discrete split points The strategy adopted is expressed as follows:
[0219] ;
[0220] For an ES agent, bandwidth proportional vector action Beta-distributed output and normalized priority actions are used. Each output uses a Dirichlet distribution, i.e., discrete segmentation point actions. The strategy adopted is expressed as follows:
[0221] ,
[0222] in, Indicates that the ES agent is in the time slot Local observations This indicates the previous time slot hidden state of the ES agent; Indicates bandwidth scaling vector action The strategy adopted Indicates normalized priority actions The strategy adopted.
[0223] Centralized Critic Network Used to evaluate the long-term benefits of each agent under the current joint policy during the training phase, providing a value benchmark for policy updates. For agents The input to the Critic network consists of the global state. With agent identity coding The result is obtained by piecing together, that is Among them, the global state Used to characterize the overall operating status of the current system, identity code One-hot encoding is used to represent it, that is The One element is 1, and the rest are 0. Based on this input The Critic network outputs an intelligent agent. The state value estimate under the current joint strategy is used to approximate its expected discounted return:
[0224] ,
[0225] in, Represents intelligent agents Expected discount return under the current joint strategy The discount factor representing the reward. Represents intelligent agents In the time slot The corresponding rewards; This represents the expectation function of the strategy.
[0226] Step 52: Initialize training parameters and trajectory buffer. Set the number of training epochs. Number of training slots per round Discount Factor GAE (Generalized Advantage Estimation) parameters PPO (Proximal Policy Optimization) pruning factor Initialize the Actor network and Critic Network Initial GRU cyclic hidden states of each agent And the trajectory buffer D. The trajectory buffer D is used to store the states, local observations, actions, action probabilities, rewards, next states, recurrent hidden states, and task completion statistics in chronological order during the training process.
[0227] Step 53, in the In round training z, for time slot t, The agent based on local observations and the previous time slot's hidden state Update the current hidden state:
[0228] ,
[0229] And based on Actor Network Sampling Segmentation Points .
[0230] For an ES agent, based on local observations and the previous time slot's hidden state Update the current hidden state:
[0231] ,
[0232] And based on ES Actor network Sampling normalized bandwidth ratio vector and normalized priority adjustment vector .
[0233] Step 54, ES Actor Network The output continuous actions are converted into executable actions that satisfy the constraints. The normalized bandwidth ratio vector is then transformed to obtain the bandwidth unit allocation. Normalized priority adjustment vector is converted to Subsequently, actions are executed, and the environment updates its state according to task generation, local inference, uplink transmission, edge inference, and dynamic priority scheduling rules to obtain the global state of the next time slot. Local observation Rewards for each intelligent agent .
[0234] Step 55: Store the state transition sample obtained in the current time slot into the trajectory buffer D. Each state transition sample includes at least the global state. Local observations of each intelligent agent Execution of actions ,award Global state of the next time slot Local observation in the next time slot GRU hidden state .
[0235] Repeat steps 53 to 55 until the... Rotation training Once all time slots have been interacted, a state transition sample arranged in chronological order is formed, and a series of state transition samples corresponding to consecutive time slots constitutes trajectory data.
[0236] Step 56: For each sample in the trajectory buffer D, construct the Critic network input with agent identity encoding. Calculating time difference error based on Critic network And the advantage function is calculated using generalized advantage estimation. :
[0237] ,
[0238] in, Represents intelligent agents In the time slot The corresponding time difference error.
[0239] Simultaneously, discount rewards are calculated based on trajectory rewards. It is used to update the centralized Critic network.
[0240] Step 57: Sample mini-batch trajectory samples from the trajectory buffer D in chronological order, and update the Actor network and Critic network respectively. For the agent... strategy probability ratio Defined as:
[0241] ,
[0242] PPO trimming target Defined as:
[0243] ,
[0244] in, This represents the probability ratio of the pruned strategies. Indicates the cutting factor; This indicates the expectation of the trajectory samples. Represents intelligent agents Current strategy, Represents intelligent agents The old strategy when collecting trajectory samples.
[0245] Centralized Critic networks update by minimizing value loss. Represented as:
[0246] ,
[0247] in, Represents intelligent agents In the time slot The estimated value of the discount return, This indicates the length of the sequence used to calculate the reward. , This is the threshold parameter.
[0248] By combining the Actor pruning objective, Critic value loss, and policy entropy regularization term, the overall training objective is obtained and the network is updated accordingly.
[0249] ,
[0250] in, Indicates the overall training objective; Represents all Actor network parameters. and These are the weighting coefficients; Represents policy entropy, Represents intelligent agents The corresponding strategy.
[0251] Through the above training process, the VSD agent learns the task-level DNN segmentation strategy, and the ES agent learns the global bandwidth allocation and dynamic priority adjustment strategy, thereby achieving low-latency processing of industrial inspection tasks in the edge-to-edge collaborative inference process.
[0252] To verify the effectiveness of this invention, this embodiment constructs an edge-to-edge collaborative inference simulation environment comprising one ES and multiple VSDs. The wireless link adopts a Rayleigh fading model, and the distance between the video sensing devices (VSDs) and the base station is randomly distributed within the range of 100 m to 500 m. The local computing power of each video sensing device (VSD) is set to 1 GHz, the transmit power to 0.5 W, the computing power of the edge server (ES) is set to 10 GHz, the total system bandwidth is set to 10 MHz, the noise power spectral density is set to -174 dBm / Hz, and the task latency constraint is set to 200 ms. The hidden layer dimension of the policy network and the value network is set to 128, and the learning rate is set to... The discount factor was set to 0.995, the training length for each round was set to 200 time slots, and the average result was taken under multiple random seeds.
[0253] The simulation experiments selected MobileNetV3, YOLOv5n and ResNet50 as heterogeneous DNN inspection models, and compared the method of the present invention with benchmark methods such as GRU-IPPO, GRU-MATD3, GRU-WQMIX and MLP-MAPPO.
[0254] Figure 4The comparison results of cumulative reward changes during training using different methods are presented to verify the training convergence and policy stability of the proposed method. Specifically, the MLP-MAPPO method uses a Multi-Layer Perceptron (MLP) as the basic architecture for the Actor and Critic networks in the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. The GRU-IPPO method adds a GRU to the network structure of the Independent Proximal Policy Optimization (IPPO) algorithm. The GRU-MATD3 method adds a GRU to the network structure of the Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm. The GRU-WQMIX method adds a GRU to the network structure of the Weighted Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning (WQMIX) algorithm. In the early stages of training, the cumulative rewards of all methods are relatively low. As the training rounds increase, the agents gradually learn the collaborative relationships between task partitioning, bandwidth allocation, and queue scheduling. Compared to comparative methods that do not incorporate GRU time-series modeling or centralized value evaluation, the method of this invention can achieve higher cumulative rewards in fewer training rounds and maintain a smaller fluctuation in the later stages of training. This result demonstrates that this invention extracts queue state, link state, and task arrival change features in continuous time slots through a GRU structure, and introduces global state information during the training phase using a centralized Critic network, which helps improve the stability of multi-agent policy updates and reduce policy oscillations caused by independent learning by each agent.
[0255] Figure 5 The on-time completion rate of tasks under different video frame rates is shown. Figure 6The average end-to-end latency variation under different video frame rates is demonstrated to verify the effectiveness of the dynamic queue-aware scheduling mechanism under varying system loads. In this set of experiments, the method of this invention is compared with heuristic scheduling rules such as First Come, First Served (FCFS), Shortest Job First (SJF), and Earliest Deadline First (EDF). These rules typically consider only a single scheduling factor and struggle to simultaneously adapt to the dynamic coupling between task deadlines, edge computing queue lengths, DNN segmentation locations, and wireless transmission resources. The method of this invention can dynamically adjust scheduling priorities based on the current queue state and task urgency, and jointly optimize with DNN partitioning and bandwidth allocation. Therefore, it can maintain a high on-time task completion rate and low end-to-end latency even under high load scenarios. Especially in congested scenarios at 18 fps, the on-time task completion rate of the method of this invention is approximately 0.43, which is about 43.3% higher than the earliest deadline priority rule, about 53.6% higher than the shortest task priority rule, and more than twice that of the first-come-first-served rule, while maintaining a lower average end-to-end latency. This result demonstrates that the method of this invention can effectively alleviate edge-side queuing congestion and improve the completion capability of latency-constrained tasks even when the video frame rate increases and tasks arrive in concentrated bursts.
[0256] Figure 7 The on-time completion rate of tasks under different DNN model load conditions is shown. Figure 8 This paper demonstrates the variation in average end-to-end latency under different DNN model loads to verify the adaptability of the proposed method to heterogeneous DNN inspection tasks. In this set of experiments, MobileNetV3, YOLOv5n, and ResNet50 were selected as DNN model loads with different computational complexities. MobileNetV3 is a lightweight network, YOLOv5n has a medium-scale computational load, and ResNet50 has relatively high computational complexity. The proposed method can jointly consider DNN segmentation location, wireless bandwidth allocation, and edge-side priority scheduling, thereby dynamically balancing communication latency and computational latency. Under the ResNet50 load, the average end-to-end latency of the proposed method is approximately 126 ms, which is about 21 ms lower than both GRU-IPPO and GRU-MATD3. Under the lighter MobileNetV3 load, the performance gap between methods narrows due to the lower computational pressure of individual tasks, but the proposed method still maintains the lowest average end-to-end latency. The results demonstrate that the joint optimization method of this invention is not only applicable to lightweight DNN inspection tasks, but also maintains good latency control and task completion performance in computationally heavy heterogeneous DNN tasks.
[0257] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. Any equivalent modifications or changes made by those skilled in the art based on the content disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. An edge-end collaborative reasoning method for industrial inspection, characterized in that, Includes the following steps: S1. Based on the processing requirements of VSD video data collected in industrial inspection scenarios, an inspection inference task model is established, and DNN segmentation points are combined to represent the task generation time, latency tolerance, local computing load, uploaded data volume and edge computing load. S2, based on the wireless communication process between VSD and ES, establish a wireless transmission model, divide the total system bandwidth into allocable bandwidth units, and calculate the uplink transmission rate and transmission delay of the inspection inference task. S3, based on the continuous arrival process of inspection reasoning tasks, VSD local computing queue, ES edge computing queue, task latency tolerance, and edge-end segmentation execution process, constructs an edge-end collaborative reasoning model; S4. Based on the low latency requirements of industrial inspection tasks, a joint optimization objective function is constructed by combining the edge-end collaborative reasoning model. By dividing the time slots, the optimization problem of solving the objective function is transformed into a partially observable Markov decision process, and a Markov model is constructed. S5. A multi-agent proximal policy optimization algorithm enhanced by gated recurrent units is adopted to construct a GRU-MAPPO framework including VSD agent and ES agent. The Markov model is used to train each agent and update the current policy. When the preset conditions are met, the optimal policy is output. S6 obtains the optimal global bandwidth allocation and ES dynamic priority adjustment strategy based on the optimal strategy, realizing low-latency processing of industrial inspection tasks in the edge-end collaborative reasoning process.
2. The edge-end collaborative reasoning method for industrial inspection according to claim 1, characterized in that, The specific steps for establishing an inspection reasoning task model include: S11, Define the VSD collection as , Indicates the total number of VSDs; S12 divides the system operation process into There are 1 time slot, and the time slot set is 1 Each time slot is [length missing] Then the first The start time of each time slot for: , Set VSD to a fixed frame rate If an inspection inference task is generated, then the time slot length satisfies the following condition: ; S13, let the nth... In the time slot Generated inspection reasoning task Represented as: , in, Indicates the task latency tolerance. This indicates the inspection reasoning task. The total number of logical layers that can be divided into in a DNN. Indicates the time when the task was generated; definition The set of deployed DNN logical layers is represented as , No. Each logical layer is represented as ,in This indicates the amount of feature data output by this logical layer. This indicates the computational complexity of the logic layer. S14, for inspection reasoning tasks Set the dividing point Indicates in The number of logic layers executed locally is the set of local execution layers. and ES execution layer collection They are respectively: , ; Based on the dividing point Computational inspection reasoning task exist The compute load on the local and Elasticsearch sides are as follows: , , in, Indicates the inspection reasoning task exist Local computing load, Indicates the inspection reasoning task Computational load on the ES side; Then the inspection reasoning task Uploaded data volume From the dividing point A decision is expressed as: , in, Indicates the dividing point The amount of input data; at that time , This indicates the amount of original input data.
3. The edge-end collaborative reasoning method for industrial inspection according to claim 2, characterized in that, The specific steps for establishing a wireless transmission model are as follows: S21, Suppose the total uplink bandwidth of the system is divided into multiple bandwidth units. In the time slot The number of bandwidth units obtained is expressed as ,satisfy: , in, Indicates the total number of bandwidth units in the system; S22, let the bandwidth of each bandwidth unit be... In the time slot Any consecutive moments within place, Uplink transmission rate for: , in, express The transmission power, express Uplink channel gain with base station Represents the noise power spectral density; when At that time, it was agreed ; S23, based on the uplink transmission rate Inspection and reasoning tasks Transmission delay satisfy: , in, Indicates the inspection reasoning task The moment the upload begins.
4. The edge-end collaborative reasoning method for industrial inspection according to claim 2, characterized in that, The specific steps for constructing an edge-to-edge collaborative reasoning model are as follows: S31, establish a local inference queue and a transmission queue on the VSD side, and establish an edge inference queue on the ES side; S32, After the inspection reasoning task is generated, it enters the local reasoning queue, which is processed according to the first-come-first-served rule. Inspection and reasoning task Local inference start time With completion time They are respectively: , , in, Indicates the inspection reasoning task The generation time, Indicates the inspection reasoning task The local inference completion time, express Computational power; S33, after the inspection reasoning task completes local reasoning, it enters the transmission queue, and its transmission begins at... With completion time They are respectively: , , in, Indicates the inspection reasoning task The time when the transmission is completed, For inspection reasoning tasks Transmission delay; S34, after the inspection and inference task is uploaded, it enters the edge inference queue; the edge inference queue adopts a non-preemptive dynamic priority scheduling mechanism, and once a task starts serving, it is not interrupted. When the edge server is idle, it selects the task with the highest dynamic priority from the waiting queue for service; inspection and inference task In the time slot Dynamic priority Defined as: , in, Indicates time slot Priority factor, Indicates the task latency tolerance. Indicates time slot The starting time, Indicates the inspection reasoning task The time when the transmission is completed, Indicates the inspection reasoning task The moment of generation; Set up an inspection reasoning task The starting time of edge reasoning is Then the edge reasoning completion time is: , in, This indicates the computing power of Elasticsearch.
5. The edge-end collaborative reasoning method for industrial inspection according to claim 2, characterized in that, The specific steps for constructing the joint optimization objective function and Markov model include: S41, based on the inspection reasoning task End-to-end processing latency and normalized end-to-end delay Define the on-time completion rate of long-term tasks in the system. and the normalized average end-to-end latency for timely task completion The expression is as follows: , , , , in, Indicates that the task was completed on time. hour, ;otherwise ; This represents a very small positive number introduced to prevent the denominator from being zero; S42, with the goal of maximizing the on-time completion rate of tasks and minimizing the average end-to-end latency, construct a joint optimization objective function: , in, This is a tradeoff factor used to balance the on-time completion rate of tasks with the normalized average end-to-end latency; , , They represent time slots respectively. Within the DNN, decisions include DNN segmentation point decisions, bandwidth allocation decisions, and priority adjustment decisions; This represents the total number of bandwidth units in the system; constraint C1 is used to limit the range of partition points selected by each VSD in each time slot; constraint C2 is used to ensure that the bandwidth allocation result is a non-negative integer; constraint C3 is used to ensure that the total bandwidth allocated to each VSD in the same time slot is equal to the total bandwidth of the system; constraint C4 is used to limit the range of values for the priority adjustment parameter. S43, if the edge-to-edge collaborative reasoning process is modeled as a partially observable Markov decision process, then the Markov model... Represented as: , in, A set of intelligent agents, numbered to The intelligent agents correspond to each Number The intelligent agent corresponds to ES; Represents the global state space; , Representing intelligent agents respectively The local observation space and action space, , ; Represents the state transition probability. Represents the reward function, Indicates the discount factor; S44, construct the global state space and the local observation space; The global state space , time slot global state Represented as: , in, Indicates ES in time slot The remaining computational load of the local inference queue at the start time. express In the time slot The remaining computational load of the local inference queue at the start time. express In the time slot The amount of data remaining in the transmission queue at the start time. express In the time slot Channel gain, This indicates the on-time completion rate of the task in the previous time slot. Indicates the transmission efficiency of the previous time slot; The local observation space In the time slot Local observation include In the time slot Local observation and ES in time slots Local observation , respectively represented as: , , in, express computing power This indicates the computing power of Elasticsearch (ES). S45 constructs the action space of the intelligent agent; Action space , Represents intelligent agents The action; Actions of the intelligent agent If we select a split point for the current task, then we have: Actions of ES agents Let the normalized bandwidth allocation vector and the normalized priority adjustment vector be, then we have ,in And satisfy Normalized priority vector And satisfy ; Indicates time slot Internal allocation The number of normalized bandwidth units, express The generated inference task is in the time slot Normalization priority factor; S46, Construct a multi-dimensional reward function consisting of team rewards and individual rewards; Team Rewards Defined as: , in, This represents the on-time completion rate of tasks in time slot t. This represents the normalized average end-to-end delay for tasks to be completed on time within time slot t. Individual rewards and ES individual rewards They are defined as follows: , , in, , They represent Compared with the individual reward weight of ES, , They represent In the time slot Within the task completion rate and normalized average end-to-end latency, , They represent time slots respectively. The transmission efficiency and service efficiency of ES queues; Total Rewards Total ES Rewards The function expressions are as follows: , , in, and All are weighting coefficients.
6. The edge-end collaborative reasoning method for industrial inspection according to claim 5, characterized in that, The GRU-MAPPO framework includes GRU modules, an Actor network, and a Critic network. The Actor network is used for distributed execution of generation strategies and has a total of [number missing]. Each agent has its own output strategy; the Critic network is used for centralized training to evaluate and update the output strategies of the Actor network. The GRU module is used to capture the effects of queue evolution, channel changes, and cross-time slot decision-making, and is positioned after the feature encoder of each agent; for any agent... Its cyclic hidden state is updated as follows: , in, Indicates the observation encoder, Indicates the hidden state of the previous time slot cycle; The Actor network is used to generate corresponding execution strategies based on the local observations of each agent; for each agent... Actor Network The input is its time slot Local observation The output of the Actor network is the agent. Action probability distribution And determine the agent based on the probability distribution of the action. Actions in the current time slot ; The Critic network is used during the training phase to evaluate the long-term returns of each agent under the current joint policy, providing a value benchmark for policy updates; for agents... Input to the Critic network From global state With agent identity coding obtained by piecing together Among them, global state Used to characterize the overall operating status of the current system, identity code One-hot encoding is used to represent the input; based on this input The Critic network outputs an intelligent agent. The state value estimate under the current joint strategy is used to approximate its expected discounted return: , in, Represents intelligent agents Expected discount return under the current joint strategy The discount factor representing the reward. Represents intelligent agents In the time slot The corresponding rewards; This represents the expectation function of the strategy.
7. The edge-end collaborative reasoning method for industrial inspection according to claim 5, characterized in that, The specific steps for training each agent include: SA1 sets hyperparameters, initializes the VSD agent Actor network, ES agent Actor network, and Critic network, and initializes the cyclic hidden state. and trajectory buffer D; SA2 initializes environment parameters at the start of each training round, including time slots. Initialize to 0, initialize each queue state to the preset value, reset the loop hidden state, and clear the trajectory buffer D; SA3, acquires local observations of the VSD agent in the current time slot. The VSD agent is based on current local observations and the current Actor network Make decisions and determine the break-off points for the current task. And set the reasoning division method between the terminal side and the edge side for the task in the current time slot; Obtain local observations of the ES agent in the current time slot The ES agent is based on current local observations and the current Actor network Make decisions and determine bandwidth allocation parameters. and priority adjustment parameters And set the uplink bandwidth resources and ES queue scheduling order for each VSD in the current time slot; SA4, based on time slot length The system updates the environmental state at the scale, enabling the generation, local inference, uplink transmission, and edge inference of inspection inference tasks, and statistically analyzes the task completion status, failure status, and queue changes of the current time slot. SA5, the VSD agent calculates the reward value for the current time slot based on the corresponding reward function. The ES agent calculates the reward value for the current time slot based on the corresponding reward function. ; SA6 stores the current time slot's state, action, reward, and the next time slot's state into the trajectory buffer D. Once the number of samples in trajectory buffer D reaches a preset limit, the stored samples are used to test the VSD agent Actor network. ES intelligent agent Actor network and Critic Network Train the network and update the current network weights.
8. An edge-end collaborative reasoning system for industrial inspection scenarios, used to execute the edge-end collaborative reasoning method for industrial inspection as described in any one of claims 1-7, characterized in that, The DNN segmentation point is dynamically determined based on the task status, the computational complexity of the DNN logical layer, the amount of output feature data of the DNN logical layer, the channel status, and the queue status, so that the local computing load, the amount of uploaded data, and the edge computing load can be adjusted according to the system status; including multiple VSDs, one ES, and the base station connected to the ES; VSD is used to continuously acquire video data related to industrial field equipment, production line status or abnormal events, and execute the first few logical layers of the DNN model locally. ES is used to receive intermediate feature data uploaded from each VSD, maintain the edge inference queue, and complete the inference of the remaining logical layers of the DNN model. Among them, ES adopts a non-preemptive dynamic priority scheduling mechanism, which determines the task scheduling order based on task latency tolerance, task generation time, upload completion time, and priority factor.