A terminal model deployment optimization method for a smart power distribution network
Patent Information
- Application Number
- CN202610815046.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-08
AI Technical Summary
[0004]本发明的目的是提供一种面向智慧配电网的终端模型部署优化方法,通过分配端边云协同的推理任务,并基于量化感知依赖图的结构化剪枝与混合位宽量化进行联合优化,以解决现有配电网终端模型部署中因忽视业务实时性约束与硬件资源限制而导致的推理精度低、时延能耗高及存储开销大的问题
[0058] 1. This invention introduces the real-time business constraints and hardware constraints from edge-cloud collaborative task allocation into model deployment, enabling pruning and quantization configuration to directly address the actual business requirements of the distribution network and the capabilities of terminal chips, avoiding blind compression that deviates from constraints. By establishing a joint optimization model aimed at minimizing inference error, energy-delay product, and storage costs, it solves the problems of low inference accuracy, high latency and energy consumption, and large storage overhead caused by neglecting real-time business constraints and hardware resource limitations in existing distribution network terminal model deployments. By introducing mask constraints to the layers to be quantized and performing bit-width traversal, it calculates the change in inference error caused by bit-width changes in each layer to be quantized, generating a hierarchical sensitivity table. This allows mixed bit-width allocation to be based on the true sensitivity of each layer to be quantized, avoiding low-sensitivity layers occupying high-bit-width resources.
Smart Images

Figure CN122331913B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a terminal model deployment optimization method for smart distribution networks, belonging to the field of smart distribution network technology. Background Technology
[0002] Monitoring data from smart distribution networks and transformer substations are characterized by high frequency, continuity, and strong time-series correlation. Traditional solutions typically upload all this data to the cloud for analysis, which not only consumes significant communication bandwidth but may also lead to uncontrollable latency and service unavailability in weak network environments. To address this, the industry has proposed offloading inference capabilities to the terminal side to improve the real-time performance and system autonomy of data processing. However, terminal devices generally face strict limitations in computing power, storage, and power consumption. Furthermore, significant differences exist among different target chips in terms of operator support, memory hierarchy, and data transfer mechanisms, making it difficult to efficiently deploy the same model across different hardware platforms.
[0003] Existing technologies suffer from the following drawbacks: First, the compression configuration and hardware mapping are mismatched, and the pruning and quantization strategies do not fully consider the operator support capabilities and memory level differences of the target chip; second, the bit width configuration and data transfer overhead are isolated, failing to achieve collaborative optimization; third, the limited computing resources on the edge make it difficult to support highly complex search and tuning processes. These drawbacks result in low inference accuracy, high latency and energy consumption, and large storage overhead in the deployed model on the edge, making it difficult to meet the dual requirements of real-time performance and reliability for smart distribution networks. Summary of the Invention
[0004] The purpose of this invention is to provide a terminal model deployment optimization method for smart distribution networks. By allocating inference tasks in a collaborative manner between the edge and cloud, and performing joint optimization based on structured pruning and hybrid bit-width quantization of quantized perception dependency graphs, this method addresses the problems of low inference accuracy, high latency and energy consumption, and large storage overhead in existing distribution network terminal model deployments caused by neglecting business real-time constraints and hardware resource limitations.
[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution:
[0006] This invention provides a method for optimizing the deployment of terminal models in smart power distribution networks, comprising:
[0007] Construct a set of inference tasks, determine the allocation method of each inference task on the terminal side, edge side, and cloud side, as well as the communication dependencies and business real-time constraints between inference tasks, and obtain inference task definition information and business constraint definition information.
[0008] Based on the inference task definition information, business constraint definition information, and hardware constraint set, graph tracing and dependency analysis are performed on the model to be deployed to the terminal to construct a quantized perception dependency graph, and a set of structured pruning parameter groups is extracted based on the quantized perception dependency graph.
[0009] The layer to be quantized is determined based on the set of parameters of structured pruning, and a mask constraint is introduced for the layer to be quantized on the target chip. Bit width traversal is performed within the preset set of candidate bit widths, and the amount of change in inference error caused by the change in bit width for each layer to be quantized is calculated.
[0010] The sensitivity of each layer to bit width variation is determined based on the change in inference error, and a hierarchical sensitivity table is generated.
[0011] Under the constraints of pruning base number and bit width range, a joint optimization model is established based on the hierarchical sensitivity table with the optimization objectives of minimizing inference error, energy extension product and storage cost, and the optimal pruning rate and optimal bit width configuration are obtained by solving the optimal pruning rate and optimal bit width configuration.
[0012] Based on the optimal pruning rate and optimal bit width configuration, an executable inference execution graph and deployment package are generated on the target chip, and the deployment package is sent to the terminal for online inference based on the inference execution graph.
[0013] Furthermore, a set of inference tasks is constructed, and the allocation method of each inference task on the terminal side, edge side, and cloud side is determined, as well as the communication dependencies and business real-time constraints between inference tasks, resulting in inference task definition information and business constraint definition information, including:
[0014] Based on the business scenarios of the distribution radio area, a set of inference tasks is constructed, and the set of inference tasks is divided into a subset of terminal-side real-time inference tasks, a subset of edge-side collaborative inference tasks, and a subset of cloud-based training tasks, which serve as inference task definition information.
[0015] Based on the set of inference tasks, the input-output dependencies, triggering conditions, and data flow between each inference task are determined, and an inference task dependency graph is generated; wherein, the triggering conditions are periodic triggering, event gating triggering, or edge node instruction triggering.
[0016] Based on the business real-time constraints, the maximum allowable end-to-end latency, output frequency, and accuracy threshold are set for each inference task in the inference task dependency graph, which serve as the business constraint definition information.
[0017] Furthermore, the set of hardware constraints includes at least one of the following: on-chip cache level parameters, memory access latency parameters, DMA transfer parameters, and the number of parallel computing units.
[0018] Furthermore, based on the inference task definition information, business constraint definition information, and hardware constraint set, graph tracing and dependency analysis are performed on the model to be deployed to the terminal to construct a quantized dependency graph. Then, based on the quantized dependency graph, a set of structured pruning parameter sets is extracted, including:
[0019] Graph tracing is performed on the model to be deployed to the terminal, and the graph tracing results are obtained by recording tensor shape, operator type and parameter sharing relationship;
[0020] Based on the graph tracing results, identify the weight quantization root node, activated quantization root node, and activated quantization termination node in the model to be deployed to the terminal, merge the quantization attachment branch and the quantization insertion branch, and deduplicate the shared weights.
[0021] Based on the quantization dependencies between layers and branches in the model to be deployed to the terminal, a quantization-aware dependency graph is generated.
[0022] Extract the set of parameter groups for structured pruning based on the quantized dependency graph, and define a gating variable for each parameter group to indicate whether the corresponding parameter group is pruned during the structured pruning process;
[0023] The quantization attachment branch refers to the bypass subnetwork introduced in the model due to the quantization operation, and the quantization insertion branch refers to the quantization operator node or dequantization operator node explicitly inserted in the model.
[0024] Furthermore, based on the parameter set of structured pruning, the layer to be quantized is determined, and a mask constraint is introduced for the layer to be quantized on the target chip. Bit width traversal is performed within a preset candidate bit width set, and the change in inference error caused by bit width variation for each layer to be quantized is calculated, including:
[0025] In the set of parameter sets for structured pruning, the layer containing the weight tensor corresponding to the parameter set that has not been pruned, indicated by the gate variable, is determined as the layer to be quantized.
[0026] The system extracts real input samples for online inference of the model to be deployed in batches from the terminal-side sliding window cache, or receives calibration samples for online inference of the model to be deployed in batches from the edge node, as on-chip evaluation data.
[0027] On the target chip, a mask constraint is introduced for the layer to be quantized to control the quantization range of the current layer to be quantized;
[0028] Iterate through each candidate bit width value within the preset candidate bit width set:
[0029] Perform forward inference under mask constraints and calculate the inference error between the output tensor with the current candidate bit width and the output tensor with the reference bit width based on on-chip evaluation data;
[0030] Based on the inference error under each candidate bit width, calculate the change in inference error of the layer to be quantized due to the change in bit width.
[0031] Furthermore, the sensitivity of each quantized layer to bit-width changes is determined based on the change in inference error, and a hierarchical sensitivity table is generated, including:
[0032] Perform a sensitivity calculation step for each layer to be quantized to obtain the sensitivity of each layer to bit width changes;
[0033] The sensitivity calculation step includes:
[0034] Obtain the inference error of the current layer to be quantized at the reference bit width, and the inference error at each candidate bit width;
[0035] Based on the inference error under the candidate bit width and the inference error under the reference bit width, calculate the sensitivity of the current quantized layer to bit width changes;
[0036] The sensitivity calculation results of each layer to be quantized are summarized in the order of network hierarchy to generate a hierarchical sensitivity table.
[0037] Furthermore, the sensitivity to changes in bit width is expressed as:
[0038] ;
[0039] In the formula, For the first The sensitivity of each layer to bit width variations. For candidate bit width, For the first The reference bit width of the layer to be quantized Indicates the first The inference error of the layer to be quantized at the candidate bit width Indicates the first The inference error of each layer to be quantized at the reference bit width; For the first The bit width of the layer to be quantized.
[0040] Furthermore, under the constraints of pruning cardinality and bit width range, a joint optimization model is established based on the hierarchical sensitivity table with the optimization objectives of minimizing inference error, energy extension product, and storage cost. The optimal pruning rate and optimal bit width configuration are then obtained, including:
[0041] Based on the hierarchical sensitivity table, obtain the sensitivity values of each layer to be quantized under different candidate bit widths;
[0042] A joint optimization model is established, the objective function of which is to minimize the weighted sum of inference error, energy extension product and storage cost, and the constraints include pruning cardinality constraint and bit width range constraint.
[0043] The pruning base number constraint is used to constrain the number of parameter groups to be pruned to meet the preset pruning base number requirement.
[0044] The bit width range constraint is used to constrain the bit width value of each layer to be quantized to be between a preset minimum bit width and a preset maximum bit width.
[0045] The joint optimization model is solved using a multi-objective optimization algorithm or a mixed-integer linear programming method, and the optimal pruning rate and optimal bit width configuration of each layer to be quantized are output.
[0046] Furthermore, the joint optimization model is expressed as:
[0047] ;
[0048] In the formula, To obtain the minimum value, For reasoning error, In order to extend the accumulation, Storage occupancy cost, of which, For weight tensors, For candidate bit width, This is the mapping strategy.
[0049] Furthermore, based on the optimal pruning rate and optimal bit width configuration, an executable inference execution graph and deployment package are generated on the target chip. The deployment package is then sent to the terminal for online inference based on the inference execution graph, including:
[0050] Based on the optimal pruning rate and optimal bit width configuration, an operator availability table is constructed; the operator availability table is used to record the kernel implementation characteristics, input / output layout and performance characteristics of each type of operator on the target chip;
[0051] Operators that are not supported by the target chip are transformed into operators that are supported by the target chip by performing operator replacement or operator decomposition operations.
[0052] Based on the optimal bit width configuration and the on-chip cache level of the target chip, the tensor layout format and weight compression format are selected.
[0053] Establish a cost model, which includes latency cost, storage cost, and data transfer cost;
[0054] The generation timing of DMA transfer descriptors and the swap-in / swap-out instruction sequence of weight tensors in the on-chip cache are determined based on the cost model.
[0055] Based on the operator availability table and the cost model, the output includes the operator execution order, parallelism, DMA transfer descriptor generation timing, and the swap-in / swap-out instruction sequence of the weight tensor in the on-chip cache, and generates an inference execution graph.
[0056] The inference execution graph, the optimal pruning rate, and the optimal bit width configuration are packaged into a deployment package, and the deployment package is sent to the terminal so that the terminal can perform online inference based on the inference execution graph.
[0057] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0058] 1. This invention introduces the real-time business constraints and hardware constraints from edge-cloud collaborative task allocation into model deployment, enabling pruning and quantization configuration to directly address the actual business requirements of the distribution network and the capabilities of terminal chips, avoiding blind compression that deviates from constraints. By establishing a joint optimization model aimed at minimizing inference error, energy-delay product, and storage costs, it solves the problems of low inference accuracy, high latency and energy consumption, and large storage overhead caused by neglecting real-time business constraints and hardware resource limitations in existing distribution network terminal model deployments. By introducing mask constraints to the layers to be quantized and performing bit-width traversal, it calculates the change in inference error caused by bit-width changes in each layer to be quantized, generating a hierarchical sensitivity table. This allows mixed bit-width allocation to be based on the true sensitivity of each layer to be quantized, avoiding low-sensitivity layers occupying high-bit-width resources.
[0059] 2. This invention extracts a set of parameter groups for structured pruning by constructing a quantization-aware dependency graph, defines a gate variable for each parameter group, and calculates the inference error change of each layer to be quantized in the bit-width traversal through mask constraints to generate a hierarchical sensitivity table. Then, a joint optimization model is established with the goal of minimizing inference error, energy extension product, and storage cost. Under the constraints of pruning cardinality and bit-width range, the optimal pruning rate and optimal bit-width configuration are solved, overcoming the local optimum problem caused by separate processing of pruning, quantization, and mapping in the prior art. Attached Figure Description
[0060] Figure 1 This is a flowchart illustrating a terminal model deployment optimization method for smart power distribution networks provided in an embodiment of the present invention. Detailed Implementation
[0061] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0062] Example 1
[0063] like Figure 1 As shown in the figure, this embodiment introduces a terminal model deployment optimization method for smart distribution networks, including:
[0064] Step 1: Construct a set of inference tasks and determine the allocation method of each inference task on the terminal side, edge side, and cloud side, as well as the communication dependencies and business real-time constraints between inference tasks, to obtain inference task definition information and business constraint definition information.
[0065] This embodiment constructs a set of inference tasks and determines the allocation method of each inference task on the terminal side, edge side, and cloud side, as well as the communication dependencies and business real-time constraints between inference tasks. This results in the inference task definition information and business constraint definition information, avoiding the situation of repeatedly changing the solution during deployment due to unclear inference task boundaries or missing communication relationships.
[0066] Step 2: Based on the inference task definition information, business constraint definition information, and hardware constraint set, perform graph tracing and dependency analysis on the model to be deployed to the terminal to construct a quantized dependency graph, and extract the set of parameter groups for structured pruning based on the quantized dependency graph.
[0067] This embodiment performs graph tracing and dependency analysis on the model to be deployed to the terminal based on the inference task definition information, business constraint definition information, and hardware constraint set to construct a quantized dependency graph. Based on the quantized dependency graph, a set of parameter groups for structured pruning is extracted so that the pruning operation preserves the inter-layer dependencies and does not damage the executable structure of the model.
[0068] Step 3: Determine the layer to be quantized based on the parameter set of structured pruning and introduce mask constraints for the layer to be quantized on the target chip. Perform bit width traversal within the preset candidate bit width set and calculate the change in inference error caused by the bit width change for each layer to be quantized.
[0069] This embodiment determines the layer to be quantized based on the parameter set of structured pruning and introduces mask constraints for the layer to be quantized on the target chip. It performs bit width traversal within a preset candidate bit width set and calculates the change in inference error caused by bit width changes for each layer to be quantized. This directly correlates the actual computing characteristics of the chip with the quantization error, which is closer to the real hardware situation than theoretical estimation.
[0070] Step 4: Determine the sensitivity of each layer to bit width changes based on the change in inference error, and generate a hierarchical sensitivity table.
[0071] In this embodiment, the sensitivity of each quantized layer to bit width changes is determined based on the change in inference error, and a hierarchical sensitivity table is generated. The hierarchical sensitivity table directly indicates which quantized layers have high bit width change costs and which quantized layers can be compressed.
[0072] Step 5: Under the constraints of pruning cardinality and bit width range, establish a joint optimization model based on the hierarchical sensitivity table with the optimization objectives of minimizing inference error, energy extension product and storage cost, and solve for the optimal pruning rate and optimal bit width configuration.
[0073] In this embodiment, under the constraints of pruning cardinality and bit width range, a joint optimization model is established based on the hierarchical sensitivity table with the optimization objectives of minimizing inference error, energy extension product, and storage occupation cost. The optimal pruning rate and optimal bit width configuration are obtained by solving the model. The solution is a balanced result of the three factors, and there will be no situation where the latency or accuracy is increased due to saving storage.
[0074] Step 6: Based on the optimal pruning rate and optimal bit width configuration, generate an executable inference execution graph and deployment package on the target chip, and send the deployment package to the terminal to perform online inference based on the inference execution graph.
[0075] This embodiment generates an executable inference execution graph and deployment package on the target chip based on the optimal pruning rate and optimal bit width configuration, and then sends the deployment package to the terminal for online inference based on the inference execution graph. This eliminates the need for the terminal to perform another model conversion or operator adaptation step, resulting in a short deployment path and a uniquely determined execution path.
[0076] Example 2
[0077] Based on the same inventive concept as Embodiment 1, this embodiment introduces the implementation steps of a terminal model deployment optimization method for smart distribution networks, including:
[0078] Step 1: Construct a set of inference tasks and determine the allocation method of each inference task on the terminal side, edge side, and cloud side, as well as the communication dependencies and business real-time constraints between inference tasks, to obtain inference task definition information and business constraint definition information.
[0079] Step 1.1: Construct an inference task set based on the business scenarios of the distribution radio area, and divide the inference task set into a subset of terminal-side real-time inference tasks, a subset of edge-side collaborative inference tasks, and a subset of cloud-based training tasks, as inference task definition information.
[0080] Step 1.2: Based on the inference task set, determine the input-output dependencies, triggering conditions, and data flow between each inference task, and generate an inference task dependency graph.
[0081] In this embodiment, the triggering condition is periodic triggering, event gating triggering, or triggering by an instruction issued by an edge node.
[0082] Step 1.3: Based on the real-time constraints of the business, set the maximum allowable end-to-end latency, output frequency and accuracy threshold for each inference task in the inference task dependency graph, as business constraint definition information.
[0083] Step 2: Based on the inference task definition information, business constraint definition information, and hardware constraint set, perform graph tracing and dependency analysis on the model to be deployed to the terminal to construct a quantized dependency graph, and extract the set of parameter groups for structured pruning based on the quantized dependency graph.
[0084] In this embodiment, the set of hardware constraints includes at least one of the following: on-chip cache level parameters, memory access latency parameters, DMA transfer parameters, and the number of parallel computing units.
[0085] Step 2.1: Perform graph tracing on the model to be deployed to the terminal, and obtain the graph tracing results by recording tensor shape, operator type and parameter sharing relationship.
[0086] Step 2.2: Based on the graph tracing results, identify the weight quantization root node, activated quantization root node, and activated quantization termination node in the model to be deployed to the terminal, merge the quantization attachment branch and the quantization insertion branch, and deduplicate the shared weights.
[0087] In this embodiment, the quantization attachment branch refers to the bypass subnetwork introduced in the model due to the quantization operation, and the quantization insertion branch refers to the quantization operator node or dequantization operator node explicitly inserted in the model.
[0088] Step 2.3: Generate a quantization-aware dependency graph based on the quantization dependencies between layers and branches in the model to be deployed to the terminal.
[0089] Step 2.4: Extract the set of parameter groups for structured pruning based on the quantized dependency graph, and define a gating variable for each parameter group to indicate whether the corresponding parameter group is pruned during the structured pruning process.
[0090] Step 3: Determine the layer to be quantized based on the parameter set of structured pruning and introduce mask constraints for the layer to be quantized on the target chip. Perform bit width traversal within the preset candidate bit width set and calculate the change in inference error caused by the bit width change for each layer to be quantized.
[0091] Step 3.1: In the set of parameter sets of structured pruning, determine the layer containing the weight tensor corresponding to the parameter set whose gate variable indicates that it has not been pruned as the layer to be quantized.
[0092] Step 3.2: Extract real input samples for online inference of the model to be deployed in batches from the terminal-side sliding window cache, or receive calibration samples for online inference of the model to be deployed in batches from the edge node, as on-chip evaluation data.
[0093] Step 3.3: Introduce a mask constraint on the target chip for the layer to be quantized to control the quantization range of the current layer to be quantized.
[0094] Step 3.4: Iterate through each candidate bit width value within the preset candidate bit width set:
[0095] Step 3.4.1: Perform forward inference under mask constraints and calculate the inference error between the output tensor with the current candidate bit width and the output tensor with the reference bit width based on the on-chip evaluation data.
[0096] Step 3.5: Calculate the change in inference error of the layer to be quantized due to the change in bit width based on the inference error of each candidate bit width.
[0097] Step 4: Determine the sensitivity of each layer to bit width changes based on the change in inference error, and generate a hierarchical sensitivity table.
[0098] Step 4.1: Perform the sensitivity calculation step for each layer to be quantized to obtain the sensitivity of each layer to bit width changes.
[0099] Step 4.1.1: Obtain the inference error of the current layer to be quantized at the reference bit width, and the inference error at each candidate bit width.
[0100] Step 4.1.2: Calculate the sensitivity of the current quantized layer to bit width changes based on the inference error under the candidate bit width and the inference error under the reference bit width.
[0101] In this embodiment, the sensitivity to bit-width changes is expressed as:
[0102] ;
[0103] In the formula, For the first The sensitivity of each layer to bit width variations. For candidate bit width, For the first The reference bit width of the layer to be quantized Indicates the first The inference error of the layer to be quantized at the candidate bit width Indicates the first The inference error of each layer to be quantized at the reference bit width; For the first The bit width of the layer to be quantized.
[0104] Step 4.2: Summarize the sensitivity calculation results of each layer to be quantized in the order of network hierarchy to generate a hierarchical sensitivity table.
[0105] Step 5: Under the constraints of pruning cardinality and bit width range, establish a joint optimization model based on the hierarchical sensitivity table with the optimization objectives of minimizing inference error, energy extension product and storage cost, and solve for the optimal pruning rate and optimal bit width configuration.
[0106] Step 5.1: Based on the hierarchical sensitivity table, obtain the sensitivity values of each layer to be quantized under different candidate bit widths.
[0107] Step 5.2: Establish a joint optimization model.
[0108] In this embodiment, the objective function of the joint optimization model is to minimize the weighted sum of inference error, energy extension product, and storage cost. The constraints include pruning cardinality constraints and bit width range constraints. The pruning cardinality constraints are used to ensure that the number of parameter groups pruned meets the preset pruning cardinality requirements. The bit width range constraints are used to ensure that the bit width of each layer to be quantized is between the preset minimum bit width and the preset maximum bit width.
[0109] In this embodiment, the joint optimization model is expressed as:
[0110] ;
[0111] In the formula, To obtain the minimum value, For reasoning error, In order to extend the accumulation, Storage occupancy cost, of which, For weight tensors, For candidate bit width, This is the mapping strategy.
[0112] Step 5.3: Solve the joint optimization model using a multi-objective optimization algorithm or a mixed-integer linear programming method, and output the optimal pruning rate and optimal bit width configuration for each layer to be quantized.
[0113] Step 6: Based on the optimal pruning rate and optimal bit width configuration, generate an executable inference execution graph and deployment package on the target chip, and send the deployment package to the terminal to perform online inference based on the inference execution graph.
[0114] Step 6.1: Construct an operator availability table based on the optimal pruning rate and optimal bit width configuration; the operator availability table is used to record the kernel implementation characteristics, input / output layout and performance characteristics of each type of operator on the target chip.
[0115] Step 6.2: Perform operator replacement or operator decomposition operations on operators not supported by the target chip to convert them into operators supported by the target chip.
[0116] Step 6.3: Based on the optimal bit width configuration and the on-chip cache level of the target chip, select the tensor layout format and weight compression format.
[0117] Step 6.4: Establish a cost model.
[0118] In this embodiment, the cost model includes latency cost, storage cost, and data transfer cost.
[0119] Step 6.5: Determine the generation timing of the DMA transfer descriptor and the swap-in / swap-out instruction sequence of the weight tensor in the on-chip cache based on the cost model.
[0120] Step 6.6: Based on the operator availability table and the cost model, output the operator execution order, parallelism, DMA transfer descriptor generation timing, and the swap-in / swap-out instruction sequence of the weight tensor in the on-chip cache, and generate the inference execution graph.
[0121] Step 6.7: Package the inference execution graph with the optimal pruning rate and the optimal bit width configuration to generate a deployment package, and send the deployment package to the terminal so that the terminal can perform online inference based on the inference execution graph.
[0122] Example 3
[0123] Based on the same inventive concept as other embodiments, this embodiment describes a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the methods of Embodiment 1 or 2 described above.
[0124] Example 4
[0125] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions that, when executed by a processor, implement the steps of the methods described in Embodiment 1 or 2 above.
[0126] In summary, this invention introduces the real-time business constraints and hardware constraints from edge-cloud collaborative task allocation into model deployment, enabling pruning and quantization configuration to directly address the actual business requirements of the distribution network and the capabilities of terminal chips, avoiding blind compression that deviates from constraints. By establishing a joint optimization model aimed at minimizing inference error, energy-delay product, and storage costs, and solving for the optimal pruning rate and optimal bit width configuration under pruning cardinality constraints and bit width range constraints, this invention solves the problems of low inference accuracy, high latency and energy consumption, and large storage overhead caused by neglecting real-time business constraints and hardware resource limitations in existing distribution network terminal model deployments. By introducing mask constraints to the layers to be quantized and performing bit width traversal, the change in inference error caused by bit width changes in each layer to be quantized is calculated, generating a hierarchical sensitivity table. This allows mixed bit width allocation to be based on the true sensitivity of each layer to be quantized, avoiding low-sensitivity layers occupying high-bit width resources.
[0127] This invention extracts a set of parameter groups for structured pruning by constructing a quantization-aware dependency graph, defines a gate variable for each parameter group, and calculates the inference error change of each layer to be quantized in the bit-width traversal through mask constraints to generate a hierarchical sensitivity table. Then, a joint optimization model is established with the goal of minimizing inference error, energy extension product, and storage cost. Under the constraints of pruning cardinality and bit-width range, the optimal pruning rate and optimal bit-width configuration are solved, overcoming the local optimum problem caused by separate processing of pruning, quantization, and mapping in the prior art.
[0128] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0129] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0130] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0131] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0132] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A terminal model deployment optimization method for smart distribution networks, characterized in that, include: Construct a set of inference tasks, determine the allocation method of each inference task on the terminal side, edge side, and cloud side, as well as the communication dependencies and business real-time constraints between inference tasks, and obtain inference task definition information and business constraint definition information. Based on the inference task definition information, business constraint definition information, and hardware constraint set, graph tracing and dependency analysis are performed on the model to be deployed to the terminal to construct a quantized dependency graph. Then, based on the quantized dependency graph, a set of structured pruned parameter sets is extracted, including: Graph tracing is performed on the model to be deployed to the terminal, and the graph tracing results are obtained by recording tensor shape, operator type and parameter sharing relationship; Based on the graph tracing results, identify the weight quantization root node, activated quantization root node, and activated quantization termination node in the model to be deployed to the terminal, merge the quantization attachment branch and the quantization insertion branch, and deduplicate the shared weights. Based on the quantization dependencies between layers and branches in the model to be deployed to the terminal, a quantization-aware dependency graph is generated. Extract the set of parameter groups for structured pruning based on the quantized dependency graph, and define a gating variable for each parameter group to indicate whether the corresponding parameter group is pruned during the structured pruning process; The quantization attachment branch refers to the bypass subnetwork introduced in the model due to the quantization operation, and the quantization insertion branch refers to the quantization operator node or dequantization operator node that is explicitly inserted in the model. The layer to be quantized is determined based on the parameter set of structured pruning, and a mask constraint is introduced for the layer to be quantized on the target chip. Bit width traversal is performed within a preset candidate bit width set, and the change in inference error caused by bit width variation for each layer to be quantized is calculated, including: In the set of parameter sets for structured pruning, the layer containing the weight tensor corresponding to the parameter set that has not been pruned, indicated by the gate variable, is determined as the layer to be quantized. The actual input samples of the model to be deployed during online inference are extracted in batches from the terminal-side sliding window cache, or the calibration samples of the model to be deployed during online inference are received in batches from the edge node as on-chip evaluation data. On the target chip, a mask constraint is introduced for the layer to be quantized to control the quantization range of the current layer to be quantized; Iterate through each candidate bit width value within the preset candidate bit width set: Perform forward inference under mask constraints and calculate the inference error between the output tensor with the current candidate bit width and the output tensor with the reference bit width based on on-chip evaluation data; Based on the inference error under each candidate bit width, calculate the change in inference error of the layer to be quantized due to the change in bit width. The sensitivity of each layer to bit width variation is determined based on the change in inference error, and a hierarchical sensitivity table is generated. Under the constraints of pruning base number and bit width range, a joint optimization model is established based on the hierarchical sensitivity table with the optimization objectives of minimizing inference error, energy extension product and storage cost, and the optimal pruning rate and optimal bit width configuration are obtained by solving the optimal pruning rate and optimal bit width configuration. Based on the optimal pruning rate and optimal bit width configuration, an executable inference execution graph and deployment package are generated on the target chip, and the deployment package is sent to the terminal for online inference based on the inference execution graph.
2. The terminal model deployment optimization method for smart distribution networks according to claim 1, characterized in that, Construct a set of inference tasks and determine the allocation method of each inference task on the terminal side, edge side, and cloud side, as well as the communication dependencies and business real-time constraints between inference tasks, to obtain inference task definition information and business constraint definition information, including: Based on the business scenarios of the distribution radio area, an inference task set is constructed, and the inference task set is divided into a terminal-side real-time inference task subset, an edge-side collaborative inference task subset, and a cloud-based training task subset, which serve as inference task definition information; Based on the set of inference tasks, the input-output dependencies, triggering conditions, and data flow between each inference task are determined, and an inference task dependency graph is generated; wherein, the triggering conditions are periodic triggering, event gating triggering, or edge node instruction triggering. Based on the business real-time constraints, the maximum allowable end-to-end latency, output frequency, and accuracy threshold are set for each inference task in the inference task dependency graph, which serve as the business constraint definition information.
3. The terminal model deployment optimization method for smart distribution networks according to claim 1, characterized in that, The set of hardware constraints includes at least one of the following: on-chip cache level parameters, memory access latency parameters, DMA transfer parameters, and the number of parallel computing units.
4. The terminal model deployment optimization method for smart distribution networks according to claim 1, characterized in that, The sensitivity of each quantized layer to bit-width changes is determined based on the change in inference error, and a hierarchical sensitivity table is generated, including: Perform a sensitivity calculation step for each layer to be quantized to obtain the sensitivity of each layer to bit width changes; The sensitivity calculation step includes: Obtain the inference error of the current layer to be quantized at the reference bit width, and the inference error at each candidate bit width; Based on the inference error under the candidate bit width and the inference error under the reference bit width, calculate the sensitivity of the current quantized layer to bit width changes; The sensitivity calculation results of each layer to be quantized are summarized in the order of network hierarchy to generate a hierarchical sensitivity table.
5. The terminal model deployment optimization method for smart distribution networks according to claim 4, characterized in that, The sensitivity to changes in bit width is expressed as: ; In the formula, For the first The sensitivity of each layer to bit width variations. For the first Candidate bit width of the layer to be quantized For the first The reference bit width of each layer to be quantized Indicates the first The inference error of the layer to be quantized at the candidate bit width. Indicates the first The inference error of the layer to be quantized at the reference bit width.
6. The terminal model deployment optimization method for smart distribution networks according to claim 5, characterized in that, Under the constraints of pruning cardinality and bit width range, a joint optimization model is established based on the hierarchical sensitivity table, with the optimization objectives of minimizing inference error, energy extension product, and storage cost. The optimal pruning rate and optimal bit width configuration are obtained by solving this model, including: Based on the hierarchical sensitivity table, obtain the sensitivity values of each layer to be quantized under different candidate bit widths; A joint optimization model is established, the objective function of which is to minimize the weighted sum of inference error, energy extension product and storage cost, and the constraints include pruning cardinality constraint and bit width range constraint. The pruning base number constraint is used to constrain the number of parameter groups to be pruned to meet the preset pruning base number requirement. The bit width range constraint is used to constrain the bit width value of each layer to be quantized to be between a preset minimum bit width and a preset maximum bit width. The joint optimization model is solved using a multi-objective optimization algorithm or a mixed-integer linear programming method, and the optimal pruning rate and optimal bit width configuration of each layer to be quantized are output.
7. The terminal model deployment optimization method for smart distribution networks according to claim 6, characterized in that, The joint optimization model is expressed as follows: ; In the formula, To obtain the minimum value, For reasoning error, In order to extend the accumulation, Storage occupancy cost, of which, For weight tensors, For the first Candidate bit width of the layer to be quantized This is the mapping strategy.
8. The terminal model deployment optimization method for smart distribution networks according to claim 7, characterized in that, Based on the optimal pruning rate and optimal bit width configuration, an executable inference execution graph and deployment package are generated on the target chip. The deployment package is then sent to the terminal for online inference based on the inference execution graph, including: Based on the optimal pruning rate and optimal bit width configuration, an operator availability table is constructed; the operator availability table is used to record the kernel implementation characteristics, input / output layout and performance characteristics of each type of operator on the target chip; Operators that are not supported by the target chip are transformed into operators that are supported by the target chip by performing operator replacement or operator decomposition operations. Based on the optimal bit width configuration and the on-chip cache level of the target chip, the tensor layout format and weight compression format are selected. Establish a cost model, which includes latency cost, storage cost, and data transfer cost; The generation timing of DMA transfer descriptors and the swap-in / swap-out instruction sequence of weight tensors in the on-chip cache are determined based on the cost model. Based on the operator availability table and the cost model, the output includes the operator execution order, parallelism, DMA transfer descriptor generation timing, and the swap-in / swap-out instruction sequence of the weight tensor in the on-chip cache, and generates an inference execution graph. The inference execution graph, the optimal pruning rate, and the optimal bit width configuration are packaged into a deployment package, and the deployment package is sent to the terminal so that the terminal can perform online inference based on the inference execution graph.
Citation Information
Patent Citations
Model optimization method, electronic equipment and storage medium
CN121745199A
Layered perceptual quantization and distributed deployment method for large language model under cloud edge collaboration
CN121887798A