Target detection method and device, terminal and storage medium
Patent Information
- Application Number
- CN202611159927.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]量化感知训练(QAT)通过在训练阶段模拟量化噪声,使模型权重对低精度计算具有鲁棒性,经QAT训练后导出的模型(如ONNX格式)的计算速度有限,量化精度和推理性能亟待提高
[0008]第四方面,本申请实施例提供一种计算机可读存储介质,其存储有计算机程序,所述计算机程序在处理器上执行时,实施上述的目标检测方法。
Smart Images

Figure CN122840118A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large model technology, and in particular to a target detection method, device, terminal and storage medium. Background Technology
[0002] With the widespread application of large models in artificial intelligence products, how to achieve high-precision and high-efficiency model deployment under resource-constrained conditions has become a key technical challenge.
[0003] Quantization-aware training (QAT) simulates quantization noise during the training phase, making the model weights robust to low-precision calculations. However, models trained with QAT (such as those in the ONNX format) have limited computation speed, and their quantization accuracy and inference performance urgently need improvement. Summary of the Invention
[0004] In view of this, embodiments of this application provide a target detection method, apparatus, terminal, and storage medium, which can achieve both efficiency and accuracy in target detection.
[0005] In a first aspect, embodiments of this application provide a target detection method, including: Quantization-aware training is performed on the target network model to obtain a model to be deployed; wherein, the intermediate representation of the model to be deployed includes node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes; After removing all node pairs from the model to be deployed, the computation graph of the model to be deployed is optimized to obtain a floating-point model; Based on the calibration dataset, post-training quantization calibration is performed on the floating-point model to generate a first quantization parameter set, and the post-training quantization calibration is performed on the model to be deployed to generate a second quantization parameter set. The parameters in the second quantization parameter set are completed using the parameters in the first quantization parameter set to generate the target hybrid quantization parameter set; A target inference engine is constructed based on the model to be deployed and the hybrid quantization parameter set, and target detection is performed based on the target inference engine.
[0006] Secondly, embodiments of this application provide a target detection device, comprising: The module for obtaining the model to be deployed is used to perform quantization-aware training on the target network model to obtain the model to be deployed; wherein, the intermediate representation of the model to be deployed includes node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes; The floating-point model construction module is used to optimize the computation graph of the model to be deployed after removing all the node pairs in the model to be deployed, so as to obtain a floating-point model. The quantization parameter set acquisition module is used to perform post-training quantization calibration on the floating-point model based on the calibration dataset to generate a first quantization parameter set, and to perform the post-training quantization calibration on the model to be deployed to generate a second quantization parameter set. The hybrid quantization parameter set generation module is used to complete the parameters in the second quantization parameter set using the parameters in the first quantization parameter set, thereby generating a target hybrid quantization parameter set. The deployment module is used to build a target inference engine based on the model to be deployed and the hybrid quantization parameter set, and to perform target detection based on the target inference engine.
[0007] Thirdly, embodiments of this application provide a terminal device, the terminal device including a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the above-described target detection method.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed on a processor, implements the target detection method described above.
[0009] The embodiments of this application have the following beneficial effects: This application obtains a floating-point model by removing node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes from the model to be deployed, enabling post-training quantization calibration to cover all network layers and solving the problem of missing quantization parameters in activation layers caused by the presence of node pairs; then, it fuses a first quantization parameter set generated based on this floating-point model with a second quantization parameter set generated based on the model to be deployed to generate a hybrid quantization parameter set covering all layers; finally, it constructs a target inference engine based on the model to be deployed and this hybrid quantization parameter set, and performs target detection based on the target inference engine. The method of this application can significantly improve the computation speed of the model to be deployed on edge devices, improve quantization accuracy and inference performance, and thus improve the efficiency of target detection. For example, current artificial intelligence products are often used in scenarios such as handling, sorting, and navigation. In these scenarios, depth maps are usually formed based on depth estimation of binocular vision, and then target detection is performed. The method of this application can improve the inference speed of the binocular depth estimation model, thereby improving the speed of object localization by the target detection module. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic diagram of the target detection architecture according to an embodiment of this application is shown; Figure 2 A first flowchart of the target detection method according to an embodiment of this application is shown; Figure 3 This paper illustrates a second flowchart of the target detection method according to an embodiment of the present application. Figure 4 A schematic diagram of the third process of the target detection method according to an embodiment of this application is shown; Figure 5 A schematic diagram of the fourth process of the target detection method according to an embodiment of this application is shown; Figure 6 The fifth flowchart of the target detection method according to an embodiment of this application is shown; Figure 7 A sixth flowchart of the target detection method according to an embodiment of this application is shown; Figure 8 A schematic diagram of a target detection device according to an embodiment of this application is shown. Detailed Implementation
[0012] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0013] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0014] In the following text, the terms "comprising," "having," and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more combinations thereof. Furthermore, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0015] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0016] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0017] The object detection method in this application is applicable to scenarios where edge-side AI products perform physical tasks, including autonomous handling, intelligent sorting, and path navigation. In these scenarios, the system typically acquires images through binocular vision, generates a depth map through disparity calculation, and uses this depth map as input to call the object inference engine constructed in this application for object detection. Furthermore, in these scenarios, model engineers have completed Quantization Awareness Training (QAT) and obtained a model to be deployed containing explicit quantization operation nodes and explicit dequantization operation nodes. The end-to-end conversion from model to executable inference engine needs to be completed locally, without relying on cloud server collaboration, involving multi-environment distribution, or requiring model reconstruction or retraining.
[0018] To better understand the embodiments of this application, some terms used in the embodiments of this application are explained below: A computation graph is a directed acyclic graph structure that a neural network model presents in an intermediate representation format (such as ONNX). It consists of nodes and directed edges connecting the nodes. Each node corresponds to a specific computational operation (such as convolution, matrix multiplication, or ReLU activation), and each directed edge represents the unidirectional flow direction of tensor data. Nodes have a type identifier, a list of input tensors, and a list of output tensors. Tensors are indexed by unique identifiers. The computation graph fully defines the forward data flow path and topological dependencies of the model.
[0019] Post-Training Quantization (PTQ) refers to the process of calculating quantization parameters (including scaling factors and zeros) based solely on the numerical distribution characteristics of the output tensors of each network layer by inputting representative calibration data into the model and performing forward inference, without modifying the model weight parameters, performing backpropagation, or adjusting the network structure. Its output is a set of quantization parameters that correspond one-to-one with the output tensors of the network layers, which is used to guide the subsequent inference engine to perform calculations in low-precision mode.
[0020] A floating-point model is an intermediate representation obtained by performing node removal and graph reconstruction operations on the original model to be deployed. Its computation graph does not contain any explicit quantization or dequantization operation nodes. The values, storage locations, and access methods of all weight parameters are exactly the same as those of the model to be deployed. The connection relationships between all layers, the dimensions, data types, and name identifiers of input and output tensors remain unchanged. All tensors are of floating-point type (such as FP32 or FP16), and there are no integer type tensors. This model is functionally completely equivalent to the model to be deployed, but eliminates the path of numerical type conversion and precision loss introduced by explicit quantization / dequantization nodes.
[0021] like Figure 1 As shown, the target detection architecture of this embodiment includes an edge computing device 100, which is configured to execute the target detection method of this embodiment; the edge computing device 100 includes two logical functional modules: The parameter preparation module 110 is used to construct a target inference engine based on the model to be deployed and the target hybrid quantization parameter set, and to perform target detection according to the target inference engine; wherein, the target inference engine is configured to receive a depth map generated by a binocular vision system as input and output the target detection result. The engine execution module 120 constructs a target inference engine based on the mixed quantization parameter set of the model to be deployed and the target, and performs target detection based on the target inference engine to obtain target detection results.
[0022] The target detection method will be described below with reference to some specific embodiments.
[0023] Figure 2 A schematic flowchart of a target detection method according to an embodiment of this application is shown. Exemplarily, the target detection method includes steps S100-S500: Step S100: Perform quantization-aware training on the target network model to obtain the model to be deployed.
[0024] The target network model is a raw neural network model that has been determined before quantization-aware training and has not undergone any quantization processing. It has a defined network architecture, floating-point precision weight parameters, and a task-oriented topology, and its computation graph does not contain any explicit quantization or dequantization operation nodes. The target network model is a full-precision floating-point model, and both weights and activations can be represented using FP32 or FP16 data types; its network architecture includes, but is not limited to, convolutional neural networks, recurrent neural networks, Transformer architectures, or combinations thereof.
[0025] It is understandable that the target network model does not directly participate in the subsequent graph structure purification, post-training quantization calibration, or inference engine construction process; it mainly serves as the input model for quantization-aware training, and after quantization-aware training, it generates a deployment model containing node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes.
[0026] Quantization-Aware Training (QAT) is a training method that explicitly models quantization errors during the model training phase. It characterizes the impact of low-precision numerical transformations on model behavior by inserting simulated quantization operations at the weights and activation outputs of specified network layers during forward propagation. During backpropagation, it maintains gradient connectivity, thereby enabling the model weights to converge to a state robust to quantization noise. After training, the model is exported as an intermediate representation containing node pairs (Q / DQ) consisting of explicit quantized linear and explicit dequantized linear nodes.
[0027] The model derived after training with quantization awareness is the model to be deployed in this embodiment. The computation graph of this model contains node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes. Each explicit quantization operation node is used to map a floating-point tensor to an integer tensor, and each explicit dequantization operation node is used to restore an integer tensor to a floating-point tensor. In each node pair, the output tensor of the explicit quantization operation node and the input tensor of the explicit dequantization operation node have the same data identifier. The node pair covers the weight parameters of all convolutional layers and other linear layers (such as fully connected layers) and their direct output tensors. The node pair does not cover the intermediate tensors of other layers (such as any nonlinear transformation layer, normalization layer, tensor fusion layer, or summation operation layer), and there are no other quantization-related nodes other than explicit quantization operation nodes and explicit dequantization operation nodes in the model to be deployed.
[0028] Step S200: After removing all node pairs from the model to be deployed, optimize the computation graph of the model to be deployed to obtain a floating-point model.
[0029] In this model, the network topology and weight parameters of the model to be deployed remain unchanged. It can be understood that, compared to the model to be deployed, all node pairs consisting of explicit quantization and dequantization operation nodes are removed from the floating-point model, and the data paths connected to the original node pairs are directly connected, thereby eliminating the numerical type conversion and precision loss paths introduced by quantization.
[0030] In some implementations, such as Figure 3 As shown, step S200 includes steps S210-S230: Step S210: Traverse each node in the computation graph of the model to be deployed and identify each node pair.
[0031] The computation graph consists of multiple nodes and directed edges connecting these nodes. Each node has a type identifier and a list of input tensors and a list of output tensors. The process of identifying node pairs is as follows: First, scan all nodes in the computation graph to find all nodes of type explicit quantization operation nodes, and then find all nodes of type explicit dequantization operation nodes. Next, for each explicit quantization operation node, check if there exists a tensor in its output tensor that also appears in the input tensor of an explicit dequantization operation node. If there is one and only one such explicit dequantization operation node, then these two nodes are identified as a paired node pair. Repeat the above process until all pairable explicit quantization operation nodes and explicit dequantization operation nodes in the computation graph have been identified, forming a complete set of node pairs.
[0032] Step S220: Connect the input tensor of each explicit quantization operation node to the output tensor of the explicit dequantization operation node paired with the current explicit quantization operation node, and remove all node pairs.
[0033] In this step, for each identified node pair, all input tensors of the explicit quantization operation node are obtained; and all output tensors of the explicit dequantization operation node are obtained. In the computation graph, each input tensor of the explicit quantization operation node is directly connected to the corresponding output tensor of the explicit dequantization operation node according to the original data dependencies; and all input and output edges of the explicit quantization operation node and the explicit dequantization operation node are disconnected; then, the explicit quantization operation node and the explicit dequantization operation node are completely removed from the computation graph. This embodiment enables data transfer that originally required explicit quantization and explicit dequantization operation nodes to be completed through direct connection, maintaining floating-point precision and the original value throughout the process.
[0034] Step S230: After removing all node pairs, an optimized computation graph is formed, and the optimized computation graph is used as a floating-point model.
[0035] It is understandable that after optimizing the original computation graph, an optimized computation graph is obtained. This optimized computation graph does not contain any explicit quantization operation nodes or explicit dequantization operation nodes; and the values, storage locations, and access methods of all weight parameters are exactly the same as those of the model to be deployed. In addition, the connection relationships between all layers are consistent with those of the model to be deployed, including the forward data flow direction and the backward gradient propagation path; all tensors are of floating-point data type, including FP32 or FP16, and there are no integer tensors; and the input nodes and output nodes are the same as those of the model to be deployed, and the dimensions, data types, and name identifiers of the input and output tensors remain unchanged.
[0036] In some implementations, after forming the optimized computation graph, the method further includes: verifying the legality and data consistency of the optimized computation graph.
[0037] Legality verification involves checking the computation graph for unconnected dangling nodes, circular dependencies, undefined tensor references, and mismatched input / output data types. Data consistency verification involves running the model to be deployed and the optimized computation graph separately using the same set of calibration input data, and comparing the numerical results on the output tensors at the same locations. If the maximum absolute error of all corresponding output tensors is less than one part per million, the data is considered consistent. The optimized computation graph, after successful verification, is the floating-point model.
[0038] It should be noted that the above steps for constructing the floating-point model can be implemented using an automated graph structure processing program (i.e., a written script). This program takes the intermediate representation file of the model to be deployed as input and performs deterministic node identification, edge reconnection, and node deletion operations on the computation graph according to the identification, connection, and removal rules defined in steps S210 to S230. After reconstruction, it automatically triggers legality and data consistency verification. The program's execution process does not rely on manual intervention, does not modify any weight parameter values, does not change the type, order, or naming of network layers, and does not introduce new nodes or connection relationships. Its sole output is an optimized computation graph file that meets all the aforementioned technical conditions. Once verified, this file is used as the floating-point model for subsequent steps.
[0039] Step S300: Perform post-training quantization calibration on the floating-point model based on the calibration dataset to generate a first quantization parameter set, and perform post-training quantization calibration on the model to be deployed to generate a second quantization parameter set.
[0040] Post-training quantization calibration refers to the process of inputting representative calibration data into the model and performing forward inference without modifying the model weight parameters, statistically analyzing the numerical distribution characteristics of the output tensors of each network layer, and calculating quantization parameters based on these distribution characteristics. Quantization parameters are used to map the floating-point numerical range to the integer representation range to support subsequent inference engine calculations in low-precision mode. This process does not involve backpropagation, gradient updates, or changes to the model structure; its output is simply a set of quantization parameters that correspond one-to-one with the output tensors of the network layers.
[0041] In this step, the calibration dataset can be selected from the same distribution as the dataset used during quantization-aware training. It can typically consist of hundreds to thousands of images or samples. These images or samples need to be preprocessed uniformly before being used for post-training quantization calibration. For example, the images can be scaled to 256×256 pixels, cropped to 224×224 pixels, and the pixel values can be normalized to the range of 0 to 1 before being used for post-training quantization calibration.
[0042] Both the first and second quantization parameter sets consist of scaling factors and their corresponding tensor identifiers. The tensor identifier is a string automatically generated based on the network layer name and data flow relationship, used to uniquely identify the output tensor. Its generation rule can be based on a prefix of the network layer's operation type and sequence number, followed by a fixed suffix; for example, the output tensor identifier of the first convolutional layer is "conv1_output", and the output tensor identifier of the first convolutional layer in the second residual block of the third stage is "layer3_2_conv1_output". This identifier is unique within the same model and can be accurately matched by the inference engine to apply the corresponding quantization parameters.
[0043] Exemplary, such as Figure 4 As shown, the floating-point model is trained and then quantized to generate a first quantization parameter set based on the calibration dataset, including steps S310-S330: Step S310: Perform post-training quantization calibration on the floating-point model and collect the first dynamic range statistics of the output tensors of each network layer. The calibration process takes a floating-point model as input and can be performed using the Post-Training Quantization (PTQ) toolchain provided by the inference framework.
[0044] The calibration dataset is input into the floating-point model batch by batch, and forward inference is performed. During inference, no actual quantization operation is performed; instead, the numerical distribution characteristics of the output tensor at each network layer are recorded. The first dynamic range statistics include, but are not limited to, the minimum, maximum, and histogram distribution of the output tensor across all calibration samples. The histogram distribution can be statistically analyzed using 2048 equal-width intervals, with each interval recording the number of floating-point values falling within that interval.
[0045] Step S320: Based on the first dynamic range statistics, determine the required first scaling factor for each network layer.
[0046] In this step, for the histogram distribution corresponding to each network layer, the scaling factor of the layer can be calculated by using the statistical distribution difference minimization algorithm. This algorithm aims to map the floating-point value range to the integer representation range. By comparing the overall difference between the floating-point histogram and the quantized integer histogram under different mapping schemes, the scaling factor corresponding to the mapping scheme that minimizes the difference is selected as the final result.
[0047] The range of integer representations is determined based on the preset integer quantization bit width; for example, when using INT8 quantization, the range of integer representations is [-128, 127]. The scaling factor is a positive real number used to characterize the linear mapping ratio from floating-point values to integers; in symmetric quantization mode, the zero point is fixed at zero, that is, the floating-point value zero is precisely mapped to the integer zero, and this zero point parameter does not participate in the calculation in this step.
[0048] In actual execution, the algorithm iterates through multiple candidate scaling factors. For each candidate scaling factor, it maps the floating-point output value of the network layer in all calibration samples to the integer representation range according to the candidate scaling factor, and calculates the distribution histogram of the mapped integers. Then, it calculates the statistical distribution difference between the integer histogram and the original floating-point histogram. The candidate scaling factor with the smallest statistical distribution difference is selected as the final scaling factor of the network layer.
[0049] For example, in this embodiment, the algorithm can be executed by the inference engine TensorRT: its built-in INT8 calibrator uses KL (Kullback-Leibler) divergence as a measure of statistical distribution difference when analyzing histograms, and determines the optimal scaling factor by minimizing KL divergence.
[0050] Step S330: Construct a set of key-value pairs from the first scaling factor and tensor identifier of each network layer, which serves as the first quantization parameter set.
[0051] It can be understood that the first quantization parameter set is organized in key-value pair format, where each key is a tensor identifier, and each value is the first scaling factor of the network layer corresponding to that tensor identifier. The number of key-value pairs in the first quantization parameter set is equal to the total number of output tensors of all weighted layers and all activated layers in the floating-point model. This set can be serialized into a binary cache file, with a file format conforming to the general specifications of inference engines. The file header contains a version identifier and verification information, followed by consecutively arranged key-value data blocks.
[0052] The first set of quantized parameters covers all weighted layers and all activation layers in the floating-point model. The weighted layers include the weight parameters of all convolutional layers and fully connected layers; the activation layers include, but are not limited to, all nonlinear transformation layers, normalization layers, tensor fusion layers, summation layers, global average pooling layers, and the output tensors of the network's final output layer; this set contains a total of 127 key-value pairs, corresponding to the output positions of all 127 network layers in the floating-point model.
[0053] Each output tensor has a unique tensor identifier, and this identifier is completely consistent with the tensor identifier corresponding to the first dynamic range statistics collected in step S310.
[0054] For example, when generating the cache file, the inference framework uses the name of the output tensor of each network layer as the key and the scaling factor calculated in step S320 as the value, and writes it into the cache file.
[0055] like Figure 5 As shown, the training and quantization calibration of the model to be deployed is performed to generate a second quantization parameter set, including steps S340-S360: Step S340: Perform post-training quantization calibration on the model to be deployed, and collect the second dynamic range statistics of the output tensors of the network layers acted upon by the quantization operation nodes and explicit dequantization operation nodes in the model to be deployed. Because the computation graph of the model to be deployed contains node pairs consisting of explicit quantization and explicit dequantization operations, the post-training quantization calibration process can only observe the output tensor of the network layer enclosed by that node pair. Therefore, the output tensor of the affected network layer is directly used as the input to an explicit quantization operation node, or its output tensor is directly generated by an explicit dequantization operation node. This type of network layer only includes convolutional layers and fully connected layers, and does not include nonlinear transformation layers, normalization layers, tensor fusion layers, or summation layers.
[0056] The second dynamic range statistics include the minimum, maximum and histogram distribution of the output tensor across all calibration samples; wherein the histogram distribution is obtained by dividing the floating-point value range into a preset number of equal-width intervals and counting the number of values falling into each interval.
[0057] Step S350: Calculate the corresponding second scaling factor based on the second dynamic range statistics.
[0058] In this step, for each observed output tensor, the same statistical distribution difference minimization algorithm as in step S320 can be used to calculate its corresponding second scaling factor. This algorithm aims to map the floating-point numerical range to the integer representation range. By comparing the overall difference between the floating-point histogram and the quantized integer histogram under different mapping schemes, the second scaling factor corresponding to the mapping scheme that minimizes this difference is selected as the result.
[0059] For example, when processing the model to be deployed, the inference framework can only perform the calculation process on the output tensors of 42 network layers due to the occlusion effect of explicit quantization operation nodes and explicit dequantization operation nodes; other inference engines also obtain the same number and coverage of scaling factor sets under the same constraints.
[0060] Step S360: Construct a set of key-value pairs from the second scaling factor and the corresponding tensor identifier of each output tensor, which serves as the second quantization parameter set.
[0061] It is understandable that the second quantization parameter set is also organized in key-value pairs, where the key is a tensor identifier and the value is a scaling factor; this set is serialized into a binary cache file with a file format compatible with the first quantization parameter set, and can be read and parsed by the same inference engine.
[0062] The second set of quantization parameters covers only the network layers in the model to be deployed that are acted upon by explicit quantization and explicit dequantization operation nodes, namely the weight parameters of all convolutional and fully connected layers and their direct output tensors; this set does not contain the output tensor identifiers of any nonlinear transformation layers, normalization layers, fusion layers or summation layers.
[0063] Step S400: Complete the parameters in the second quantization parameter set with the parameters in the first quantization parameter set to generate the target mixed quantization parameter set.
[0064] In some implementations, such as Figure 6 As shown, step S400 specifically includes steps S410-S420: Step S410: Based on the second quantization parameter set, construct the initial mixed quantization parameter set.
[0065] In this step, all key-value pairs in the second quantization parameter set are copied as is and written into a new parameter set, which is the initial mixed quantization parameter set of the initial state. This operation ensures that the scaling factors of all network layers (i.e., all convolutional layers and fully connected layers) acted upon by explicit quantization and dequantization operation nodes are fully inherited without any adjustment or overriding.
[0066] Step S420: Iterate through each parameter item in the first quantization parameter set. If the corresponding tensor identifier does not exist in the initial mixed quantization parameter set, add the current parameter item to the initial mixed quantization parameter set. Once the first quantization parameter set has been traversed, the target quantization parameter set is formed.
[0067] In this step, each key-value pair in the first quantization parameter set is traversed, and its key, i.e., tensor identifier, is extracted. It is then checked whether the tensor identifier has appeared in the initial mixed quantization parameter set constructed in step S410. If it has not appeared, the key-value pair is added to the initial mixed quantization parameter set as a whole. If it has appeared, it is skipped, and no operation is performed. The above checking and adding actions are repeated until all key-value pairs in the first quantization parameter set have been processed. At this point, the mixed quantization parameter set to be output is the target mixed quantization parameter set.
[0068] It can be understood that a portion of the key-value pairs in the target hybrid quantization parameter set are directly derived from the second quantization parameter set, corresponding to the weights of all convolutional and fully connected layers and their direct output tensors. The remaining key-value pairs are derived from the first quantization parameter set, corresponding to the output tensors of the nonlinear transformation layer, normalization layer, tensor fusion layer, summation layer, and global average pooling layer. The keys in all key-value pairs are tensor identifiers, and they are completely consistent with the tensor identifiers corresponding to the first dynamic range statistics and the second dynamic range statistics collected in steps S310 and S340.
[0069] Step S500: Construct a target inference engine based on the model to be deployed and the hybrid quantization parameter set, and perform target detection based on the target inference engine.
[0070] In this step, the model to be deployed is used as the input for the network structure definition, and the target hybrid quantization parameter set is loaded; a preset quantization precision mode (such as INT8) is enabled in the inference framework, and the target hybrid quantization parameter set is used as the source of calibration cache; the inference engine construction process is executed, which includes computation graph parsing, operator fusion optimization, and hardware kernel selection, and finally generates an executable inference engine file; then, the inference engine file, along with the dependent libraries required for running, is transferred to the edge computing device and run on it to perform object detection.
[0071] In the field of robotics, in scenarios such as handling, sorting, and navigation, target detection can take a depth map generated by a binocular vision system as input and output at least one target's spatial location information and category identifier through a target inference engine. The spatial location information includes two-dimensional image coordinates or three-dimensional spatial coordinates, and the category identifier is used to characterize the physical entity type to which the target belongs, including handling objects, sorted items, or navigation obstacles.
[0072] In some implementations, the construction of the target inference engine also includes performing layer sensitivity assessments on each network layer covered by the target hybrid quantization parameter set to identify target layers that are highly sensitive to quantization operations.
[0073] like Figure 7 As shown, the layer sensitivity assessment includes steps S10-S50: Step S10: Using the model to be deployed as input, construct a benchmark inference engine based on the quantization scaling factor in the target mixed quantization parameter set, and evaluate and obtain the benchmark accuracy index on the validation set.
[0074] The benchmark inference engine configures the first quantization precision for all candidate layers, while network layers not covered by the target mixed quantization parameter set retain their original floating-point precision.
[0075] The candidate layer is the network layer covered by the target hybrid quantization parameter set.
[0076] The validation set can be a publicly available image dataset with the same data distribution as the data used in the quantization-sensory training. This validation set does not participate in any training or calibration process; it is only used in steps S10, S20, and the final performance evaluation, and its data distribution characteristics are completely consistent with the data used in the quantization-sensory training phase.
[0077] This step takes the model to be deployed as the network definition input, loads the target hybrid quantization parameter set, and enables the first quantization precision mode of the inference engine. At this time, all network layers are configured with the first quantization precision according to the scaling factor in the parameter set. If five hundred images in the validation set can be used as input, inference can be performed on the engine, and its Top-1 classification precision can be calculated. The result is the baseline precision index.
[0078] Step S20: For each candidate layer, construct a temporary hybrid precision engine.
[0079] The temporary mixed-precision engine refers to an independently runnable inference engine instance built on the same model to be deployed as the network structure basis, the same mixed quantization parameter set as the source of quantization parameters, setting only one specified candidate layer to the second precision, while configuring all other candidate layers to the first quantization precision. This engine is only used to perform a single precision evaluation on the validation set, and is destroyed after the evaluation is completed, without participating in the final deployment. Its construction process calls the layer granularity precision configuration interface provided by the inference engine framework, without modifying the model weights, changing the computation graph topology, or generating new intermediate representation files.
[0080] Specifically, the computational precision of the current candidate layer is set to the second precision, while the remaining candidate layers maintain the first quantization precision. Network layers not covered by the target mixed quantization parameter set maintain the original floating-point precision. The current temporary mixed precision engine is evaluated on the same validation set, and the corresponding test precision index is obtained. The candidate layer is the network layer whose output tensor in the model to be deployed is covered by the mixed quantization parameter set. The candidate layer is uniquely determined by the tensor identifiers contained in the target hybrid quantization parameter set, specifically: Source layer limitation: candidate layers must belong to the actual network layers in the computation graph of the model to be deployed, including convolutional layers, fully connected layers, nonlinear transformation layers, normalization layers, tensor fusion layers, summation layers, and global average pooling layers.
[0081] The criteria for selection are limited. A network layer is included as a candidate layer only if its output tensor has a completely matching tensor identifier in the mixed quantization parameter set. For example, if conv1_output exists in the mixed quantization parameter set, the first convolutional layer is a candidate layer; if layer3_2_relu_output exists, the corresponding ReLU layer is a candidate layer.
[0082] Excluding boundary constraints, candidate layers do not include any network layers that are not present in the model to be deployed or whose output tensors are not indexed by the target hybrid quantization parameter set; in particular, input layers, intermediate debug nodes, unconnected dangling layers, and all control flow nodes that do not generate output tensors are not included. In this step, a temporary inference engine is constructed for each candidate layer. Its network definition, mixed quantization parameter set, and validation set input are exactly the same as in step S10. Only the computational precision of the current candidate layer is set to the second precision, while all other network layers are still configured to the first quantization precision according to the scaling factor in the mixed quantization parameter set. The temporary engine is run on 500 images that are exactly the same as in step S10, and its Top-1 classification precision is calculated. The result is the test precision index corresponding to the candidate layer.
[0083] Step S30: For each candidate layer, calculate the difference between the test accuracy index and the benchmark accuracy index, and use it as the accuracy recovery amount of the current candidate layer. In this step, for each candidate layer, the test accuracy index obtained in step S20 is subtracted from the baseline accuracy index obtained in step S10, and the difference is the accuracy recovery amount of the candidate layer; the difference is a non-negative real number.
[0084] Step S40: The candidate layer whose accuracy recovery amount is greater than the preset sensitivity threshold is determined as the target layer. Among them, the recovery threshold is set according to the model accuracy requirements and deployment resource constraints; this threshold can be a 0.5% accuracy improvement or other values, and its value range can be adjusted between 0.1% and 2% according to the actual task requirements.
[0085] In this step, the accuracy recovery values of all candidate layers are sorted from largest to smallest; the larger the recovery value, the more sensitive the layer is to quantization; the candidate layers with accuracy recovery values greater than the preset sensitivity threshold are determined as the target layers.
[0086] In other implementations, an upper limit on the number of layers to be skipped for quantization can be set; this upper limit can be five layers or any integer between one and ten layers; the candidate layers that are ranked in the top few positions after sorting are determined as the target layers, where the number of such positions is equal to the upper limit on the number of layers.
[0087] In step S50, when constructing the final target inference engine, for each target layer, the computational precision is set to the second precision, and the corresponding quantization scaling factor is excluded from the target mixed quantization parameter set. For non-target layers, the first quantization precision is configured according to the quantization scaling factor in the mixed quantization parameter set.
[0088] In this context, non-target layers are network layers among the candidate layers excluding the target layers. It can be understood that for non-target layers, the inference engine framework automatically configures them to the first quantization precision based on the quantization scaling factor in the hybrid quantization parameter set.
[0089] It is understandable that, during the construction of the final target inference engine, for each target layer determined in step S40, no quantization parameters are read from the mixed quantization parameter set. Instead, the layer is explicitly specified to perform calculations with the second precision through the layer precision configuration interface provided by the inference engine framework. At the same time, it is ensured that the mixed quantization parameter set does not contain the tensor identifier and its scaling factor corresponding to the target layer. For the remaining non-target layers, the scaling factor corresponding to its tensor identifier is read from the mixed quantization parameter set and configured as the first quantization precision accordingly.
[0090] In some implementations, the method of this embodiment can also be extended to a two-stage deployment and closed-loop optimization system. The first stage performs quantized perceptual training and derives the model to be deployed; the second stage executes steps S200 to S500 to complete engine construction. The system automatically records the layer sensitivity score, parameter variability, final accuracy, and latency during each deployment process, and optimizes the preset sensitivity threshold or layer limit in subsequent deployments based on historical data. This closed-loop optimization does not change the basic process of steps S100 to S500, and is only used to improve long-term deployment efficiency. This embodiment removes all node pairs consisting of explicit quantization and dequantization operation nodes from the model to be deployed, reconstructs the computation graph to obtain a floating-point model, and maintains the original network topology and weight parameters unchanged, thereby supporting complete statistics on the output tensors of all network layers. Furthermore, it generates a first quantization parameter set based on the floating-point model and a second quantization parameter set based on the model to be deployed. Using the second quantization parameter set as a basis, it determines whether to add parameters from the first quantization parameter set based on the existence of tensor identifiers, and then fuses them to generate a hybrid quantization parameter set covering all layers. Next, it identifies target layers highly sensitive to quantization operations through layer sensitivity assessment. When constructing the target inference engine, it configures a second precision for the target layer and a first quantization precision for the remaining network layers based on the hybrid quantization parameter set. This allows it to retain the adaptability of the weight layers obtained from QAT training without changing the original structure and weights of the model to be deployed, while also supplementing the quantization parameters of activation layers missing due to node structure limitations. This supports the inference framework in constructing a target inference engine based on the target hybrid quantization parameter set and performing target detection based on this target inference engine.
[0091] Figure 8 A schematic diagram of a target detection device according to an embodiment of this application is shown. Exemplarily, the target detection device includes: The module 10 for obtaining the model to be deployed is used to perform quantization-aware training on the target network model to obtain the model to be deployed; wherein, the intermediate representation of the model to be deployed includes node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes. The floating-point model building module 20 is used to optimize the computation graph of the model to be deployed after removing all node pairs in the model to be deployed, so as to obtain the floating-point model. The quantization parameter set acquisition module 30 is used to perform post-training quantization calibration on the floating-point model based on the calibration dataset to generate the first quantization parameter set, and to perform post-training quantization calibration on the model to be deployed to generate the second quantization parameter set. The hybrid quantization parameter set generation module 40 is used to complete the parameters in the second quantization parameter set using the parameters in the first quantization parameter set, thereby generating a target hybrid quantization parameter set. Deployment module 50 is used to build a target inference engine based on the model to be deployed and the hybrid quantization parameter set, and to perform target detection based on the target inference engine.
[0092] In some implementations, after removing all node pairs from the model to be deployed, the computation graph of the model to be deployed is optimized to obtain a floating-point model, including: traversing each node in the computation graph of the model to be deployed and identifying each node pair; connecting the input tensor of each explicit quantization operation node to the output tensor of the explicit dequantization operation node paired with the current explicit quantization operation node, and removing all node pairs; after removing all node pairs, an optimized computation graph is formed, and the optimized computation graph is used as the floating-point model.
[0093] In some implementations, both the first quantization parameter set and the second quantization parameter set consist of scaling factors and their corresponding tensor identifiers; the tensor identifier is a string generated based on the network layer name and data flow relationship to uniquely identify the output tensor.
[0094] The process involves performing post-training quantization calibration on the floating-point model based on a calibration dataset to generate a first quantization parameter set. This includes: performing post-training quantization calibration on the floating-point model and collecting first dynamic range statistics of the output tensors of each network layer; determining the required first scaling factor for each network layer based on the first dynamic range statistics; and constructing a set of key-value pairs from the first scaling factors and tensor identifiers of each network layer as the first quantization parameter set. The number of key-value pairs in the first quantization parameter set is equal to the total number of output tensors of all weighted layers and all activated layers in the floating-point model.
[0095] In some implementations, performing post-training quantization calibration on the model to be deployed to generate a second quantization parameter set includes: performing post-training quantization calibration on the model to be deployed, collecting second dynamic range statistics of the output tensors of the network layers acted upon by the quantization operation nodes and explicit dequantization operation nodes of the model to be deployed; calculating the corresponding second scaling factor based on the second dynamic range statistics; and constructing a key-value pair set of the second scaling factor of each output tensor and the corresponding tensor identifier as the second quantization parameter set.
[0096] In some implementations, the parameters in the second quantization parameter set are completed using the parameters in the first quantization parameter set to generate a target mixed quantization parameter set. This includes: constructing an initial mixed quantization parameter set based on the second quantization parameter set; traversing each parameter item in the first quantization parameter set, and if the corresponding tensor identifier does not exist in the initial mixed quantization parameter set, adding the current parameter item to the initial mixed quantization parameter set until the first quantization parameter set is traversed, thus forming the target mixed quantization parameter set.
[0097] In some implementations, the construction of the target inference engine further includes: performing layer sensitivity assessment on each network layer covered by the target mixed quantization parameter set to identify target layers highly sensitive to quantization operations; the layer sensitivity assessment includes: using the model to be deployed as input, constructing a benchmark inference engine based on the quantization scaling factor in the target mixed quantization parameter set, and evaluating and obtaining a benchmark accuracy metric on a validation set; wherein the benchmark inference engine configures a first quantization accuracy for all candidate layers, and the candidate layers are network layers covered by the target mixed quantization parameter set; for each candidate layer, a temporary mixed accuracy engine is constructed; wherein only the operational accuracy of the current candidate layer is set to the second accuracy, and the remaining candidate layers are set to the second accuracy. The selected layer is kept at the first quantization precision; the current temporary mixed precision engine is evaluated on the same validation set and the corresponding test precision index is obtained; for each candidate layer, the difference between the test precision index and the benchmark precision index is calculated as the precision recovery amount of the current candidate layer; the candidate layers with a precision recovery amount greater than the preset sensitivity threshold are determined as target layers; when constructing the final target inference engine, for each target layer, the computational precision is set to the second precision, and the corresponding quantization scaling factor is excluded from the target mixed quantization parameter set; for non-target layers, the first quantization precision is configured according to the quantization scaling factor in the mixed quantization parameter set; where non-target layers are network layers in the candidate layers excluding the target layers.
[0098] In some implementations, after forming the optimized computation graph, the method further includes: verifying the legality and data consistency of the optimized computation graph.
[0099] It is understood that the apparatus in this embodiment corresponds to the target detection method in the above embodiments, and the options in the above embodiments are also applicable to this embodiment, so they will not be described again here.
[0100] This application also provides a terminal device, exemplary of which includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to enable the terminal device to perform the functions of the various modules in the above-described target detection method or target detection device.
[0101] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0102] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.
[0103] This application also provides a computer-readable storage medium for storing the computer program used in the aforementioned terminal device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0105] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0106] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0107] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A target detection method, characterized in that, include: Quantization-aware training is performed on the target network model to obtain a model to be deployed; wherein, the intermediate representation of the model to be deployed includes node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes; After removing all node pairs from the model to be deployed, the computation graph of the model to be deployed is optimized to obtain a floating-point model; Based on the calibration dataset, post-training quantization calibration is performed on the floating-point model to generate a first quantization parameter set, and the post-training quantization calibration is performed on the model to be deployed to generate a second quantization parameter set. The parameters in the second quantization parameter set are completed using the parameters in the first quantization parameter set to generate the target hybrid quantization parameter set; A target inference engine is constructed based on the model to be deployed and the hybrid quantization parameter set, and target detection is performed based on the target inference engine.
2. The target detection method according to claim 1, characterized in that, After removing all node pairs from the model to be deployed, the computation graph of the model to be deployed is optimized to obtain a floating-point model, including: Traverse each node in the computation graph of the model to be deployed and identify each node pair; Connect the input tensor of each explicit quantization operation node to the output tensor of the explicit dequantization operation node paired with the current explicit quantization operation node, and remove all node pairs. After removing all the node pairs, an optimized computation graph is formed, and the optimized computation graph is used as the floating-point model.
3. The target detection method according to claim 1, characterized in that, Both the first quantization parameter set and the second quantization parameter set consist of scaling factors and their corresponding tensor identifiers; The tensor identifier is a string generated based on the network layer name and data flow relationship to uniquely identify the output tensor; Based on the calibration dataset, post-training quantization calibration is performed on the floating-point model to generate a first quantization parameter set, including: Perform the post-training quantization calibration on the floating-point model and collect the first dynamic range statistics of the output tensors of each network layer; Based on the first dynamic range statistics, a first scaling factor is determined for each of the network layers. The first scaling factor and the tensor identifier of each network layer are constructed into a key-value pair set, which serves as the first quantization parameter set; wherein the number of key-value pairs contained in the first quantization parameter set is equal to the total number of output tensors of all weighted layers and all activated layers in the floating-point model.
4. The target detection method according to claim 3, characterized in that, The step of performing the post-training quantization calibration on the model to be deployed to generate a second quantization parameter set includes: Perform the post-training quantization calibration on the model to be deployed, and collect the second dynamic range statistics of the output tensors of the network layers acted upon by the quantization operation nodes and explicit dequantization operation nodes of the model to be deployed; Based on the second dynamic range statistics, the corresponding second scaling factor is calculated; The second scaling factor of each output tensor and the corresponding tensor identifier are used to construct a set of key-value pairs, which serve as the second quantization parameter set.
5. The target detection method according to claim 4, characterized in that, The step of supplementing the parameters in the second quantization parameter set with the parameters in the first quantization parameter set to generate the target hybrid quantization parameter set includes: Based on the second set of quantization parameters, an initial mixed quantization parameter set is constructed; Iterate through each parameter item in the first quantization parameter set. If the corresponding tensor identifier does not exist in the initial mixed quantization parameter set, add the current parameter item to the initial mixed quantization parameter set. Continue until the first quantization parameter set is traversed, and then the target mixed quantization parameter set is formed.
6. The target detection method according to claim 1, characterized in that, When constructing the target inference engine, the method further includes: performing layer sensitivity evaluation on each network layer covered by the target hybrid quantization parameter set to identify target layers that are highly sensitive to quantization operations; The layer sensitivity assessment includes: Using the model to be deployed as input, a benchmark inference engine is constructed based on the quantization scaling factor in the target hybrid quantization parameter set, and the benchmark accuracy index is evaluated and obtained on the validation set; wherein, the benchmark inference engine configures a first quantization accuracy for all candidate layers, and the candidate layers are the network layers covered by the target hybrid quantization parameter set; For each candidate layer, a temporary mixed precision engine is constructed; wherein, only the computational precision of the current candidate layer is set to the second precision, while the other candidate layers are kept at the first quantization precision; the current temporary mixed precision engine is evaluated on the same validation set and the corresponding test precision index is obtained; For each candidate layer, the difference between the test accuracy index and the benchmark accuracy index is calculated as the accuracy recovery amount of the current candidate layer; Candidate layers whose accuracy recovery amount is greater than a preset sensitivity threshold are identified as target layers; When constructing the final target inference engine, for each target layer, the computational precision is set to the second precision, and the corresponding quantization scaling factor is excluded from the target mixed quantization parameter set; for non-target layers, the first quantization precision is configured according to the quantization scaling factor in the mixed quantization parameter set; wherein, the non-target layer is the network layer in the candidate layer excluding the target layer.
7. The method according to claim 2, characterized in that, The process of forming the optimized computation graph also includes: verifying the legality and data consistency of the optimized computation graph.
8. A target detection device, characterized in that, include: The module for obtaining the model to be deployed is used to perform quantization-aware training on the target network model to obtain the model to be deployed; wherein, the intermediate representation of the model to be deployed includes node pairs consisting of explicit quantization operation nodes and explicit dequantization operation nodes; The floating-point model construction module is used to optimize the computation graph of the model to be deployed after removing all the node pairs in the model to be deployed, so as to obtain a floating-point model. The quantization parameter set acquisition module is used to perform post-training quantization calibration on the floating-point model based on the calibration dataset to generate a first quantization parameter set, and to perform the post-training quantization calibration on the model to be deployed to generate a second quantization parameter set. The hybrid quantization parameter set generation module is used to complete the parameters in the second quantization parameter set using the parameters in the first quantization parameter set, thereby generating a target hybrid quantization parameter set. The deployment module is used to build a target inference engine based on the model to be deployed and the hybrid quantization parameter set, and to perform target detection based on the target inference engine.
9. A terminal device, characterized in that, The terminal device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the target detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed on a processor, implements the target detection method according to any one of claims 1-7.