Collaborative optimization system for edge inference engine
The collaborative optimization system for edge inference engines addresses device heterogeneity and dynamic resource constraints by dynamically adapting to network conditions, enhancing computational efficiency and reducing latency in edge computing environments.
Patent Information
- Application Number
- CN202510373862.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to achieve the balance of accuracy and efficiency of efficient inference tasks in an edge computing environment, and the lack of adaptive resource scheduling and task division strategies leads to imbalance inference delay and communication overhead.
A collaborative optimization system using an environment perception module, a collaborative decision-making module and a dynamic execution module is adopted to update the device capability map in real time, generate an optimization action instruction set, and perform calculation graph reconstruction and resource remapping to achieve real-time matching and dynamic adjustment of model and hardware characteristics.
It improves the utilization rate and energy efficiency ratio of computing units, reduces inference delay and communication overhead, ensures service continuity and delay certainty in complex network environments, and reduces the deployment and maintenance costs of multi-device clusters.
Smart Images

Figure CN120317364A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of collaborative optimization of edge inference engines, and particularly to a collaborative optimization system for edge inference engines. Background Art
[0002] The execution of inference tasks in an edge computing environment faces the dual constraints of device heterogeneity and resource dynamics. The current mainstream technical solutions have systematic defects in achieving efficient inference: Although model compression methods can reduce the computational load, their static optimization mechanism is difficult to adapt to the dynamically changing edge resource environment, resulting in significant accuracy degradation in complex inference tasks, and lacking the operator-level optimization ability for dedicated acceleration hardware; The distributed collaborative inference framework is limited by a fixed task partitioning strategy and cannot be adaptively adjusted according to real-time network status and device computing power fluctuations, resulting in an imbalance between communication overhead and computing benefits; Existing resource scheduling schemes generally adopt an offline configuration mode and fail to establish a real-time coupling relationship between computational load and hardware resources (such as memory bandwidth, cache capacity, power consumption threshold), which is prone to a sharp increase in inference latency in scenarios with bursty data streams. These technical shortcomings make it difficult for existing systems to meet the core requirements of inference task latency determinacy and energy efficiency ratio in scenarios such as industrial Internet of Things and mobile augmented reality. Summary of the Invention
[0003] The purpose of the present invention is to propose a collaborative optimization system for an edge inference engine to solve the technical problem of how to improve the efficiency and accuracy of edge inference.
[0004] On the one hand, a collaborative optimization system for an edge inference engine is provided, including:
[0005] An environment perception module, a collaborative decision module, and a dynamic execution module that are interconnected;
[0006] The environment perception module is configured to, when the system starts, collect the instruction emission efficiency of the computing unit, the row buffer hit rate of the memory controller, and the power consumption curve of the power management unit, construct an initial device capability map; and update the hardware state vector with a preset fixed time window to perform real-time update of the initial device capability map;
[0007] The collaborative decision module is configured to generate an optimized action instruction set through a preset policy network and output the corresponding optimized action instruction set according to the corresponding allocation scheme; wherein, the allocation scheme includes at least a lightweight heuristic rule and a complete rule; the lightweight heuristic rule is used to quickly respond to sudden state changes in a time-sensitive path and trigger a computational graph simplification operation; the complete rule is used to perform a complete policy gradient update in the main optimization path;
[0008] The dynamic execution module is used to start the computational graph reconstruction engine and the resource remapping program after receiving optimization instructions. Among them, the computational graph reconstruction engine is used to parse the mixed-precision configuration matrix output by the policy network of the engine and perform differential transformation on the original computational graph: on the premise of satisfying the inter-layer numerical stability constraint, inject parameters adapted to the current hardware characteristics into the computational flow graph; the resource remapping program is used to control the execution of the inference task within the specified resource quota through memory isolation domain configuration, computing core binding, and interrupt priority adjustment according to the allocation scheme output by the collaborative decision-making module.
[0009] Preferably, the environment perception module is specifically used to update the hardware state vector according to the following formula:
[0010]
[0011] where S h represents the hardware state vector, c μ represents the dynamic utilization rate of the computing unit, m b represents the memory bandwidth, p w represents the power consumption state of the power management unit, represents the resource type.
[0012] Preferably, it further includes obtaining the dynamic utilization rate of the computing unit by sampling through the instruction-level performance counter according to the following formula:
[0013]
[0014] where N issue (t) represents the number of valid instructions issued at time t, N max is the theoretical peak instruction throughput, d represents the computing coefficient, t0 represents the initial time, and T represents the maximum time.
[0015] Preferably, it further includes determining the memory bandwidth according to the following formula:
[0016]
[0017] where represents the number of bytes transmitted by the kth memory channel within the window period τ, i represents the current serial number of the memory channel, and w represents the maximum serial number of the memory channel.
[0018] Preferably, the collaborative decision-making module is specifically used to determine the corresponding allocation scheme through a preset decision-making process model. Among them, the decision-making process model defines the state space, determines the corresponding state through the real-time network quality index and data criticality, and determines the corresponding allocation scheme according to the determined state.
[0019] Preferably, the dynamic execution module is further configured to continuously monitor abnormal events, and when a node failure or resource limit is detected, activate a fault tolerance recovery protocol to handle the abnormal events.
[0020] Preferably, the fault tolerance recovery protocol includes freezing the current inference pipeline, rolling back to the nearest safe checkpoint; performing differential analysis on the device capability graph, and finding the optimal migration target among neighboring nodes; reconstructing the distributed computation graph according to the optimal migration target;
[0021] Wherein, the optimal migration target is a migratable target whose interruption time of executing a task is less than the threshold of the set service level agreement.
[0022] In summary, implementing the embodiments of the present invention has the following beneficial effects:
[0023] The collaborative optimization system for the edge inference engine provided by the present invention realizes real-time matching of the model structure and hardware characteristics through the joint optimization of the device capability graph and the computation graph, significantly improves the utilization rate of computing units and the energy efficiency ratio while maintaining the original accuracy, and solves the inherent contradiction between accuracy and efficiency in traditional model compression; and based on the dynamic task partitioning strategy of channel quality and computing load, effectively balances the benefits of local computing and collaborative inference, reduces the end-to-end communication overhead, and ensures the delay determinacy in a complex network environment; the combination of the dual-loop control architecture and the elastic resource scheduling mechanism realizes fast response to bursty loads, avoids a sharp increase in inference delay, and ensures service continuity in industrial scenarios. The design of the hardware abstraction intermediate representation layer supports seamless conversion of instruction sets of heterogeneous computing units, and significantly reduces the deployment and maintenance costs of multi-device clusters. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and for those of ordinary skill in the art, obtaining other drawings without creative efforts still belongs to the scope of the present invention.
[0025] Figure 1 It is a schematic diagram of a collaborative optimization system for an edge inference engine in an embodiment of the present invention.
[0026] Figure 2 It is a schematic flowchart of a collaborative optimization operation for an edge inference engine in an embodiment of the present invention. Detailed Embodiments
[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings.
[0028] As Figure 1 and Figure 2 shown, it is a schematic diagram of an embodiment of a collaborative optimization system for an edge inference engine provided by the present invention. In this embodiment, a lightweight probe program is embedded in the target edge device, and the following monitoring parameters are configured: Sampling frequency: utilization rate of the computing unit (10 kHz), memory bandwidth (1 kHz), power consumption (100 Hz), etc. Probe interface: Read voltage / frequency through the PMU register, and capture the DMA transfer status through the PCIe performance counter. Run a standard test load (such as a ResNet-50 inference task), and collect the initial hardware characteristics: H0 = [c μ = 0.82, m b = 58 GB / s, p w = 45 W], and store them in the shared memory circular buffer as a reference for subsequent optimization.
[0029] This system includes: an environment perception module, a collaborative decision-making module, and a dynamic execution module that are interconnected; the environment perception module is used to collect the instruction emission efficiency of the computing unit, the row buffer hit rate of the memory controller, and the power consumption curve of the power management unit when the system starts, and construct an initial device capability map; and update the hardware state vector with a preset fixed time window to update the initial device capability map in real time; the collaborative decision-making module is used to generate an optimized action instruction set through a preset policy network, and output the corresponding optimized action instruction set according to the corresponding allocation scheme; wherein, the allocation scheme at least includes a lightweight heuristic rule and a complete rule; the lightweight heuristic rule is used to quickly respond to sudden state changes in the time-sensitive path and trigger the calculation graph simplification operation; the complete rule is used to perform a complete policy gradient update in the main optimization path; the dynamic execution module is used to start the calculation graph reconstruction engine and the resource remapping program after receiving the optimization instruction; wherein, the calculation graph reconstruction engine is used to reconstruct the differential transformation of the original calculation graph by parsing the mixed-precision configuration matrix output by the policy network: under the premise of satisfying the inter-layer numerical stability constraint, inject the parameters adapted to the current hardware characteristics into the calculation flow graph; the resource remapping program is used to control the execution of the inference task within the specified resource quota through memory isolation domain configuration, computing core binding, and interrupt priority adjustment according to the allocation scheme output by the collaborative decision-making module.
[0030] It should be noted that the operation process of this system is embodied as a closed-loop control system at three levels: environment perception, collaborative decision-making, and dynamic execution, and continuous optimization is achieved through phased operations with precise timing. When the system starts, it first performs hardware feature baseline calibration: the environment perception module activates the embedded probe program to collect the instruction emission efficiency of the computing unit, the row buffer hit rate of the memory controller, and the power consumption curve of the power management unit with a microsecond-level time resolution, and constructs an initial device capability map. This map is transmitted to the collaborative decision-making module through the shared memory interface, triggering the weight initialization process of the offline pre-trained policy network. In this stage, a synthetic load generator is used to simulate typical edge scenarios, and the mapping relationship between the policy baseline and the hardware capabilities is established by solving the following optimization problem.
[0031]
[0032] Among them, TV(θ) is the total variation regularization term of the policy parameters, which is used to enhance the robustness of decision-making.
[0033] After entering the online inference stage, the system executes a dynamic optimization loop with a fixed time window Δt = 200ms as the cycle. At the start of each cycle, the environment perception module updates the hardware state vector. And jointly with the channel coherence time τ output by the network quality monitor. c Packet loss rate p loss And other parameters to construct a joint state space. After receiving this state input, the collaborative decision-making module generates an optimized action instruction set through the policy network π. θ (a|s t ) This process contains two parallel computational paths: in the time-sensitive path, lightweight heuristic rules are used to quickly respond to sudden state changes (such as GPU video memory exhaustion warnings), and immediately trigger the computational graph simplification operation; in the main optimization path, a complete policy gradient update is performed:
[0034]
[0035] Among them, the immediate reward r t is calculated comprehensively in three dimensions: delay improvement, energy efficiency gain, and accuracy retention. After receiving the optimization instruction, the dynamic execution module starts the computational graph reconstruction engine and the resource remapping program. The reconstruction engine parses the mixed-precision configuration matrix M output by the policy network. p ∈{0, 1} L×B (L is the number of network layers, B is the bit width level), and performs differential transformation on the original computational graph: on the premise of satisfying the inter-layer numerical stability constraint. , inject low-precision operators that adapt to the current hardware characteristics into the computational flow graph. The resource remapping program then assigns the allocation scheme A ∈ N output by the decision-making module. N×R(Where N is the number of computing units and R is the resource type), through memory isolation domain configuration, computing core binding, and interrupt priority adjustment, ensure that the inference task is executed within the specified resource quota.
[0036] In one embodiment, the environment perception module is specifically configured to update the hardware state vector according to the following formula:
[0037]
[0038] Where S h represents the hardware state vector, c μ represents the dynamic utilization rate of the computing unit, m b represents the memory bandwidth, p w represents the power consumption state of the power management unit, represents the resource type.
[0039] Where the dynamic utilization rate of the computing unit is obtained by sampling through the instruction-level performance counter according to the following formula:
[0040]
[0041] Where N issue (t) represents the number of valid instructions issued at time t, N max is the theoretical peak instruction throughput, d represents the calculation coefficient, t0 represents the initial time, and T represents the maximum time.
[0042] Determine the memory bandwidth according to the following formula:
[0043]
[0044] Where represents the number of bytes transferred by the kth memory channel within the window period τ, i represents the current sequence number of the memory channel, and w represents the maximum sequence number of the memory channel.
[0045] For the environment perception module, the module establishes a multi-dimensional quantization system for hardware characteristics and defines the device state vector as: Where c μ characterizes the dynamic utilization rate of the computing unit and is obtained by sampling through the instruction-level performance counter:
[0046]
[0047] Where N issue (t) represents the number of valid instructions issued at time t, N max is the theoretical peak instruction throughput.
[0048] The memory bandwidth m b Adopts a sliding window estimation algorithm:
[0049]
[0050] wherein is the number of bytes transferred by the k-th memory channel within the window period τ. The power consumption state p w is deduced by establishing a thermodynamic model:
[0051]
[0052] Construct an equipment capability map by integrating the above parameters to provide a quantitative basis for model optimization.
[0053] In one embodiment, the collaborative decision-making module is specifically configured to determine a corresponding allocation scheme through a preset decision process model, wherein the decision process model defines a state space, determines a corresponding state through a real-time network quality index and data criticality, and determines a corresponding allocation scheme according to the determined state.
[0054] For the collaborative decision-making module, this module constructs a Markov decision process model and defines a state space
[0055] s t =(S h , Q n , D k )
[0056] where Q n represents the network quality index, and D k is the data criticality. The policy network π(a|s) is trained by the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, and its objective function is:
[0057]
[0058] where the value function V(s) satisfies the Bellman optimal equation:
[0059]
[0060] The output of the policy network includes three decision dimensions: 1) the model structure parameter adjustment vector Δθ ∈ R d 2) the task division ratio ρ ∈ [0, 1] m 3) the resource allocation matrix A ∈ N n×r .
[0061] In one embodiment, the dynamic execution module is further configured to continuously monitor for exception events, and when a node failure or resource overlimit is detected, activate the fault tolerance and recovery protocol to handle the exception events. The fault tolerance and recovery protocol includes: freezing the current inference pipeline and rolling back to the most recent safe checkpoint; performing a differential analysis on the device capability graph to find the optimal migration target among neighboring nodes; reconstructing the distributed computation graph according to the optimal migration target; wherein the optimal migration target is a migratable target whose interruption time for task execution is less than the threshold of the set service level agreement.
[0062] For the dynamic execution module, this module realizes the transformation from strategy to execution, and defines the inference pipeline as:
[0063]
[0064] where the constraint condition R includes C-type resource limitations such as memory and power consumption. The execution engine uses a differential architecture search:
[0065]
[0066] The regularization term Reg(θ, H) enforces the alignment of the model parameters θ with the hardware characteristics H, and the specific form is:
[0067]
[0068] where, M l (H) is the hardware adaptation mask matrix, and ⊙ represents the Hadamard product. The resource scheduler solves a mixed-integer programming problem: This optimization ensures obtaining a Pareto optimal solution between discrete resource allocation (such as CPU core binding) and continuous resource allocation (such as memory quota).
[0069] Specific embodiments: Environment perception (triggered every Δt = 200 ms), hardware status update: Collect the latest hardware metrics and calculate the dynamic capability graph:
[0070] (Smoothing coefficient α = 0.7, to prevent instantaneous jitter)
[0071] Network quality assessment: Measure the round-trip delay (RTT) and packet loss rate through ICMP probe packets, and construct the network quality vector: Q n =[RTT = 28 ms, Loss = 0.3%]
[0072] Collaborative decision-making (decision-making cycle 5 Hz), policy network inference: Input the joint state s t =(H t , Q n , D k) To the TD3 policy network, output an action instruction: ```python action = {'quantization policy': {'convolutional layer': 'INT8', 'fully connected layer': 'FP16'}, 'task division': {'local execution': 70%, 'cooperating node': 30%},'resource allocation': {'GPU video memory': 3.2GB, 'NPU cores': 4}}
[0073] Policy verification, check the feasibility of the action:
[0074] if ∑ video memory requirements ≤ available video memory then execute else trigger a degradation policy
[0075] Dynamic execution (real-time response)
[0076] Computation graph reconstruction, dynamically modify the computation graph according to the quantization policy.
[0077] Exception handling process, fault detection; monitor the following exception events (detection frequency 1kHz): hardware error (ECC memory error count > threshold) - resource overrun (CPU utilization > 95% for 500ms).
[0078] In summary, implementing the embodiments of the present invention has the following beneficial effects:
[0079] The collaborative optimization system for the edge inference engine provided by the present invention, through the joint optimization of the device capability map and the computation graph, realizes the real-time matching of the model structure and hardware characteristics. While maintaining the original accuracy, it significantly improves the utilization rate of computing units and the energy efficiency ratio, and solves the inherent contradiction between accuracy and efficiency in traditional model compression; and based on the dynamic task division strategy of channel quality and computing load, it effectively balances the benefits of local computing and collaborative inference, reduces the end-to-end communication overhead, and ensures the delay determinacy in a complex network environment; the combination of the double-loop control architecture and the elastic resource scheduling mechanism realizes a fast response to sudden loads, avoids a sharp increase in inference delay, and ensures service continuity in an industrial-level scenario. The design of the hardware abstraction intermediate representation layer supports seamless instruction set conversion of heterogeneous computing units, and significantly reduces the deployment and maintenance costs of multi-device clusters.
[0080] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A collaborative optimization system for an edge inference engine, characterized in that, Including: An environment perception module, a collaborative decision-making module, and a dynamic execution module that are interconnected; The environment perception module is used to, when the system starts, collect the instruction emission efficiency of the computing unit, the row buffer hit rate of the memory controller, and the power consumption curve of the power management unit, and construct an initial device capability map; and update the hardware state vector at a preset fixed time window to perform real-time update on the initial device capability map; The collaborative decision-making module is used to generate an optimized action instruction set through a preset policy network and output the corresponding optimized action instruction set according to the corresponding allocation scheme; wherein, the allocation scheme at least includes a lightweight heuristic rule and a complete rule; the lightweight heuristic rule is used to quickly respond to sudden state changes in the time-sensitive path and trigger a computational graph simplification operation; the complete rule is used to perform a complete policy gradient update in the main optimization path; The dynamic execution module is used to, after receiving the optimization instruction, start a computational graph reconstruction engine and a resource remapping program; wherein, the computational graph reconstruction engine is used to parse the mixed-precision configuration matrix output by the policy network by the reconstruction engine and perform a differential transformation on the original computational graph: on the premise of satisfying the inter-layer numerical stability constraint, inject parameters adapted to the current hardware characteristics into the computational flow graph; the resource remapping program is used to control the execution of the inference task within the specified resource quota through memory isolation domain configuration, computing core binding, and interrupt priority adjustment according to the allocation scheme output by the collaborative decision-making module.
2. The system according to claim 1, wherein The environment perception module is specifically used to update the hardware state vector according to the following formula: Among them, S h represents the hardware status vector, c μ represents the dynamic utilization rate of the computing unit, m b represents the memory bandwidth, p w represents the power consumption status of the power management unit, represents the resource type.
3. The system according to claim 2, wherein It also includes obtaining the dynamic utilization rate of the computing unit through instruction-level performance counter sampling according to the following formula: Among them, N issue (t) represents the number of valid instructions issued at time t, N max is the theoretical peak instruction throughput, d represents the calculation coefficient, t0 represents the initial time, and T represents the maximum time.
4. The system according to claim 2, wherein It also includes determining the memory bandwidth according to the following formula: Among them, represents the number of bytes transferred by the k-th memory channel within the window period τ, i represents the current serial number of the memory channel, and w represents the maximum serial number of the memory channel.
5. The system according to claim 2, wherein The collaborative decision-making module is specifically used to determine the corresponding allocation scheme through a preset decision-making process model, wherein the decision-making process model defines the state space, determines the corresponding state through the real-time network quality index and data criticality, and determines the corresponding allocation scheme according to the determined state.
6. The system according to claim 5, characterized in that The dynamic execution module is also used to continuously monitor abnormal events, and when a node failure or resource overrun is detected, activate a fault tolerance and recovery protocol to handle the abnormal events.
7. The system according to claim 6, wherein The fault tolerance and recovery protocol includes freezing the current inference pipeline and rolling back to the nearest safe checkpoint; performing a difference analysis on the device capability map and finding the optimal migration target among neighboring nodes; Reconstructing the distributed computational graph according to the optimal migration target; wherein, the optimal migration target is a migratable target whose interruption time for executing the task is less than the threshold of the set service level agreement.