Pedestrian detection method and system based on deep learning automatic pruning

By modeling the DNN model as a graph structure and using deep reinforcement learning to generate pruning strategies, the problem of insufficient generalization ability of existing automated pruning methods for various DNN architectures is solved, and efficient computational complexity and inference delay are achieved, while maintaining the accuracy of the model.

CN120164191AActive Publication Date: 2025-06-17OCEAN UNIV OF CHINA +1

Patent Information

Application Number
CN202510644843.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing automated pruning methods mainly focus on structured and unstructured pruning, lack the ability to generalize various deep neural network (DNN) architectures, and fail to effectively integrate pruning pattern information and topological information to improve pruning effect.

Method used

By modeling the DNN model into a graph structure, integrating pruning pattern information, and generating pruning strategies using deep reinforcement learning (DRL), automated pruning of multiple DNN architectures are achieved. This method combines the graph embedding representation and pruning reward function, and adaptively adjusts the pruning strategy to achieve the optimal reduction effect.

Benefits of technology

Automatic pruning of various DNN architectures including CNN and Transformer has been achieved, which significantly reduces computational complexity (the FLOPs can be reduced by 90%) and inference delay (the latency can be reduced by more than 50% on some edge devices), while maintaining the accuracy of the model and improving computing efficiency and hardware acceleration compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164191A_ABST
    Figure CN120164191A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, in particular to a pedestrian detection method and system based on deep learning automatic pruning. The method comprises the steps of obtaining pedestrian image data; preprocessing the acquired pedestrian image data; constructing a deep neural network DNN model; pruning the DNN model based on a mode pruning strategy; and deploying the pruned DNN model to an automatic driving vehicle for pedestrian detection. Compared with a traditional pruning method, the scheme has the advantages that the calculation efficiency and the hardware acceleration compatibility are remarkably improved while the model accuracy loss is extremely small and even partial recovery can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and in particular to a pedestrian detection method and system based on deep learning automatic pruning. Background Art

[0002] In recent years, with the continuous development of deep learning technology and the continuous improvement of GPU performance, deep neural networks (DNNs) have developed towards deeper levels and more parameters to learn higher-level and multi-dimensional data semantics, achieving breakthrough progress in fields such as image recognition, image classification, audio and video processing, and promoting the development of fields such as intelligent transportation, e-commerce platforms, intelligent healthcare, and digital cities. However, in practical applications, especially on edge computing devices, there is often a contradiction between limited hardware resources, small memory, and the large number of layers and parameters in deep neural network models. This is because edge devices generally use low-power processors to balance performance and cost. For example, for the VGG-16 network with 138 million parameters, it requires approximately 1.6 billion floating-point operations per second during inference and occupies approximately 528MB of memory space, which poses a challenge to the model deployment on edge computing devices. For instance, to run the VGG-16 network model on an edge device equipped with a Hi3536AV100 SoC chip and 512MB of memory, it is necessary to use model lightweighting techniques to adjust or design the network structure or increase the memory size to complete the deployment.

[0003] With the increasing demand for deploying deep neural network models on edge devices, various model lightweighting techniques have emerged in an endless stream (knowledge distillation, model compression, NAS, parameter quantization, tensor decomposition, etc.). Among them, model compression, especially model pruning methods, has become an essential model lightweighting strategy. Model pruning methods can be divided into three types according to the pruning granularity (as Figure 1 shown): unstructured pruning (or fine-grained pruning), semi-structured pruning (or pattern-based pruning), and structured pruning (or coarse-grained pruning). For example, taking a convolutional layer containing k convolutional kernels with a kernel size of 3×3 as an example, unstructured pruning can delete weights at any position in the weight tensor of a filter, resulting in an irregular non-zero weight distribution. This irregularity requires specialized software and hardware support to be effectively accelerated. In contrast, structured pruning deletes entire filters, enabling acceleration using conventional deep learning frameworks. However, due to its relatively coarse granularity, structured pruning usually has difficulty achieving the best balance between compression rate and accuracy.

[0004] To find the best pruning strategy, automated methods such as AutoML (Automated Machine Learning) have recently attracted attention. These methods use graph learning techniques, especially graph neural networks (GNNs), to represent or embed DNNs. Then the embedded data is input into a deep reinforcement learning (DRL) algorithm to find the optimal pruning strategy, thereby achieving higher compression efficiency. However, despite the significant progress made by existing automatic pruning methods, they mainly focus on structured and unstructured pruning, and pattern-based pruning methods are limited and face three main challenges: (1) They mainly deal with convolutional neural networks (CNNs) and lack the generalization ability for various DNN architectures, such as the encoder and decoder in ViT-B / 16; (2) They are mainly designed for specific operators with fixed sizes (such as 3×3 convolutions) in CNNs, which may limit their effectiveness for other shapes; (3) Previous work has not considered integrating pruning pattern information and topological information into the DNN model and combining it with automatic search to improve the final pruning effect. Summary of the Invention

[0005] To solve the above-mentioned problems, the present invention provides a pedestrian detection method and system based on deep learning automated pruning.

[0006] In a first aspect, a pedestrian detection method based on deep learning automated pruning provided by the present invention adopts the following technical solution: A pedestrian detection method based on deep learning automated pruning includes: Obtain pedestrian image data; Preprocess the obtained pedestrian image data; Construct a deep neural network DNN model; Prune the DNN model based on a pattern pruning strategy; Deploy the pruned DNN model to an autonomous vehicle for pedestrian detection.

[0007] Further, pruning the DNN model based on the pattern pruning strategy includes initializing a pruning pattern library and evaluating the DNN model compression parameters; modeling the DNN model as a graph structure and obtaining a graph embedding representation; using DRL to generate a pruning strategy based on the graph embedding representation; and performing compilation optimization on the DNN model using a compiler after executing the pruning strategy to obtain a pruned and optimized DNN model.

[0008] Further, initializing the pruning pattern library and evaluating the DNN model compression parameters includes designing a pattern library graded by sparsity rate in combination with the distribution law of weight parameters. The sparsity rate of the patterns in the pattern library decreases uniformly from 100% to 0%, and a sparsity interval Δ is set to control the sparsity granularity. When Δ = 9, 10 sparsity rates are generated, corresponding to 10 types of patterns. By calculating the compression parameters of the DNN model, it is determined whether the constraints are satisfied. If the constraints are not satisfied, the pruning strategy is continued to be searched.

[0009] Further, modeling the DNN model as a graph structure and obtaining the graph embedding representation includes extracting the network topology structure information from the DNN model, integrating the pruning pattern information into the graph structure, and obtaining the graph embedding representation of the model through the GNN method. The graph constructed by the DNN is denoted as where is the node set,[[]] is the edge set, and the subset represents the nodes associated with the weight tensors, is the input tensor node set, is the output node set, and the subsets and correspond to the edges linked to . The nodes in different layers of the DNN are further divided into finer-grained sets.

[0010] Further, using DRL to generate a pruning strategy based on the graph embedding representation includes passing the graph embedding obtained from the graph encoder as the environmental state to the Agent. The direct output obtained by the Agent is the probability distribution of the patterns in the pattern library. Therefore, its action space is continuous. Then, by calculating the action values, sampling is performed on the selected target pattern to obtain the pruned DNN model. Among them, the function is used as a multi-layer perceptron to be responsible for extracting graph representation information from the graph representation embedding, and the activation function maps the output of to the interval to determine the selection probability of each pattern, that is, the action space. Finally, the categorical distribution function is applied to sample from the probability distribution to assign patterns to different weight tensors in the DNN, generating a newly selected pattern library for pruning.

[0011] ​​Further, the method of generating a pruning strategy from graph embedding representations using DRL further includes setting a reward function by combining model compression metrics and model inference accuracy, expressed as: where is a learnable parameter. When is greater than a set threshold, the Agent is encouraged to adopt an aggressive compression strategy, prioritizing model compression over accuracy. Conversely, when is less than the set value, the Agent is incentivized to take a conservative approach.

[0012] Further, after performing the pruning strategy on the DNN model, compiler-based compilation optimization includes, after the pruned DNN model meets the constraint conditions, performing a pruning operation on the model, replacing the model parameters with 0 according to the assigned pattern to obtain the pruned DNN model. At the same time, since the inference accuracy of the pruned DNN model will decrease, the pruned DNN model is fine-tuned to restore its inference accuracy.

[0013] Further, after performing the pruning strategy on the DNN model, compiler-based compilation optimization further includes using the TVM open-source compiler for graph optimization, performing operator fusion, memory optimization, and parallelism adjustment; adopting an auto-tuning strategy to adjust the kernel size, parallel strategy, and data caching method for the target hardware platform to fully reduce the inference latency; compiling to generate the final low-level code or executable file, and supporting hot updates or online deployment mechanisms for easy testing and optimization in various hardware environments.

[0014] Further, deploying the pruned model to an autonomous vehicle for pedestrian detection includes deploying the pruned DNN model in the Perception Module, which is responsible for extracting environmental information from sensor data and making detection results.

[0015] In a second aspect, a pedestrian detection system based on deep learning automated pruning includes: A data acquisition module configured to acquire pedestrian image data; A preprocessing module configured to preprocess the acquired pedestrian image data; A model construction module configured to construct a deep neural network (DNN) model; A pruning module configured to prune the DNN model based on a pattern pruning strategy; A detection module configured to deploy the pruned DNN model to an autonomous vehicle for pedestrian detection.

[0016] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted to be loaded and executed by a processor of a terminal device to implement the pedestrian detection method based on deep learning automated pruning as described above.

[0017] Fourthly, the present invention provides a terminal device including a processor and a computer-readable storage medium. The processor is configured to implement each instruction, and the computer-readable storage medium is configured to store a plurality of instructions adapted to be loaded and executed by the processor to implement the pedestrian detection method based on deep learning automated pruning as described above.

[0018] In summary, the present invention has the following beneficial technical effects: Firstly, it can achieve automatic pruning for various DNN architectures including CNN, Transformer, etc., with good generality. Secondly, by using a graph structure to represent network topology and weight information and embedding pruning patterns into edge features, it can finely control the pruning granularity and effectively delete local computing modules. Then, it adopts deep reinforcement learning to automatically search for pruning strategies, enabling the pruning process to adaptively adjust according to the pruning reward function to achieve the optimal pruning effect, thereby significantly reducing FLOPs (for example, up to 90% reduction) and greatly reducing the inference latency (more than 50% reduction on some edge devices). Finally, compared with traditional pruning methods, the proposed scheme has a minimal loss of model accuracy or even partial recovery, while significantly improving the computing efficiency and hardware acceleration compatibility.

[0019] The present invention has the following advantages: 1) Using a hierarchical sparsity rate can cover a larger search space, and the size of the search space can be controlled by adjusting Δ; 2) Unifying unstructured pruning and structured pruning simplifies the design work of pruning strategies; 3) It is not limited to a specific model or a single operation structure and is applicable to various complex networks (including residual networks, Transformer, etc.); 4) Adopting deep reinforcement learning for strategy search can automatically discover redundant parts in the network without manual design of pruning schemes; 5) The pruned network can be optimized at the underlying layer with the help of advanced compilers, especially suitable for the low-latency operation requirements on edge devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is the overall framework diagram of the present invention.

[0021] Figure 2 is the framework diagram of modeling DNN as a graph and obtaining the graph embedding representation in the present invention.

[0022] Figure 3 is the framework diagram of the deep reinforcement learning pruning strategy search module.

[0023] Figure 4It is a framework diagram for performing pattern pruning and fine-tuning on the DNN execution mode.

[0024] Figure 5 It is a framework diagram for compiling, optimizing and deploying the pruned model.

[0025] Figure 6 It is a flowchart of the deep learning model automatic pattern pruning method of the present invention for heterogeneous edge devices. Specific implementation mode

[0026] The present invention will be further described in detail below with reference to the accompanying drawings.

[0027] Embodiment 1 Refer to Figure 1 , a pedestrian detection method based on deep learning automatic pruning in this embodiment includes: Obtain pedestrian image data; Preprocess the obtained pedestrian image data; Construct a deep neural network DNN model; Prune the DNN model based on the pattern pruning strategy; Deploy the pruned DNN model to an autonomous vehicle for pedestrian detection.

[0028] Specifically: S1. Obtain data, Use the publicly available pedestrian detection datasets COCO and Caltech Pedestrian Dataset to train the model, ensuring that the datasets contain a large number of pedestrian images and annotations.

[0029] COCO dataset: It contains approximately 330,000 images and 80 categories, including pedestrians (category ID is 1). Each object has a corresponding bounding box, and the annotation includes the category label (pedestrian) and the position coordinates.

[0030] Caltech Pedestrian dataset: It includes approximately 10,000 frames of images, containing scenes of pedestrians and backgrounds. Each frame of image annotates the position and category of pedestrians. Each pedestrian position is annotated by a rectangular box (bounding box).

[0031] S2. Preprocess the obtained data, Pedestrian detection models (such as YOLO, Faster R-CNN, etc.) usually require the input image size to be fixed. The following processing method is adopted: Resize: Resize the image to the input size required by the network. For example, the standard input size of YOLO is 416×416 or 608×608 pixels, and Faster R-CNN usually requires an input of 800×800 pixels.

[0032] Maintain aspect ratio: When resizing the image, it is best to maintain the aspect ratio of the image to avoid affecting image features due to distortion. Usually, the image can be padded to a fixed size (by adding black borders).

[0033] Normalization: Scale the pixel values of the image to a unified range, usually 0 to 1 or -1 to 1. This can accelerate training and avoid the problem of gradient vanishing / explosion. Use 0-1 normalization.

[0034] Data augmentation: A commonly used technique in deep learning to increase the diversity of the training set and improve the generalization ability of the model. Use augmentation methods such as rotation, cropping, horizontal flipping, and color jittering.

[0035] Data annotation format conversion: YOLO format, each image corresponds to a text file, which contains the class ID of each target (pedestrian), the center coordinates of the bounding box (normalized values relative to the image width and height), and the aspect ratio; PascalVOC format, each image corresponds to an XML file, which contains the class of each target, the bounding box coordinates, etc.

[0036] S3. Build a deep neural network DNN model; Among them, 1. The construction process of the YOLOv3 model: (1) Network structure (Darknet-53 backbone + FPN), Backbone network: Darknet-53 (including residual connections), output feature maps of 3 scales (such as 13×13, 26×26, 52×52).

[0037] Feature Pyramid Network (FPN): Fuse multi-scale features to improve the detection ability of small targets.

[0038] (2) Input and output, Input: The image is resized to a fixed size (such as 416×416).

[0039] Output: Each grid predicts B bounding boxes, and each box contains: Coordinates : The offset of the center point relative to the grid, and the scaling of the width and height relative to the anchor box.

[0040] Confidence: .

[0041] Class probability: (Softmax or Sigmoid output).

[0042] (3) Key formulas Bounding box prediction: (1) where : Coordinates of the top - left corner of the grid; : Width and height of the anchor box; : Network prediction value; : Sigmoid function.

[0043] Loss function (multi - task loss): (2) where : Weight coefficient; : Mean squared error; : Binary cross - entropy; : Multi - class cross - entropy.

[0044] (4) Post - processing (NMS) Filter out low - score boxes by confidence, and retain boxes with IoU lower than the threshold.

[0045] 2. Faster R - CNN construction process: (1) Network structure Backbone network: VGG16 / ResNet, etc., to extract feature maps.

[0046] Region Proposal Network (RPN): Slide a window on the feature map to generate anchor boxes (Anchors, e.g., 9 scales / width - height ratios). Output: Class scores (foreground / background) of the anchor boxes and bounding box offsets.

[0047] RoIPooling / RoIAlign: Unify the proposed regions into fixed - size features.

[0048] Classification and regression head: Predict specific classes and accurate box coordinates.

[0049] (2) Key formulas RPN loss function: (3) where : Probability that the anchor box is foreground; : True label (0 / 1); : Predicted offset; : True offset.

[0050] SmoothL1 loss: (4) Rol Pooling: Divide the proposed regions of different sizes into a grid of H×W, and perform max pooling within each grid.

[0051] Final detection loss: (5) (3) Training process, End-to-end training: The RPN and the detection network share features and are alternately optimized.

[0052] Anchor box matching strategy: Samples with IoU > 0.7 with the ground truth box are positive samples, and < 0.3 are negative samples.

[0053] S4. Prune the DNN model based on the pattern pruning strategy, including the following steps: Step 1, Initialize the pruning pattern library, The direct object of pattern pruning is the weight parameters of the DNN. Parameters with large values have a greater impact on the model performance, and vice versa, those with smaller values have a smaller impact on the model performance. The parameter value distributions in different weight tensors are different, but the shapes (patterns) formed by the parameters with larger weights are regular: certain specific shapes will appear repeatedly; multiple kernels in the same filter show similar patterns. The present invention designs a pattern library graded by sparsity rate in combination with the distribution law of weight parameters. The sparsity rate of the patterns in the pattern library decreases uniformly from 100% to 0%. Set a sparsity interval to control the sparsity granularity. For example, when, 10 sparsity rates can be generated, and there are correspondingly 10 types of patterns.

[0054] Among them, when , the pattern library contains pruning patterns with 10 sparsity rates. For each sparsity rate, start searching from the smallest number of patterns (i.e., at the beginning, each sparsity rate has only one pattern), and through multiple iterations, gradually increase the number of patterns until a balance is reached between the number of patterns and the pruning effect, and ensure that the pattern shape is consistent with the shape of the weight tensor of the DNN (for example, for a 3×3 convolution kernel, the pattern size is a 3×3 matrix), and automatically generate multiple binary matrices with different pruning shapes; in the pattern library, each pattern is stored in the form of a binary matrix, where the value of 1 indicates retention and 0 indicates pruning. The maximum number of each pattern depends on the number of weights pruned by that sparsity rate. For example, for a 3×3 pattern with a 55.56% sparsity rate, 5 weights need to be pruned, then the total number of patterns at this sparsity rate is pieces.

[0055] Step 2, Evaluate the compression parameters of the DNN model, First, calculate the compression parameters of the DNN model (such as inference accuracy, FLOPs, and pruning rate) to determine whether they meet the constraints. If they do not meet the constraints, continue to search for pruning strategies (jump to step 3); otherwise, perform pruning on the DNN (jump to step 5).

[0056] Among them, set the pruning rate constraint hyperparameter to 60%. Use MACs (Multiply-Accumulate Operations) to calculate the computational amount of zero parameters and all parameters of the DNN model, and calculate the proportion of the computational amount of zero parameters to the computational amount of all parameters. When this proportion is greater than or equal to 60%, it means that the constraint conditions are met; otherwise, they are not met.

[0057] Step 3, model the DNN as a graph and obtain the graph embedding representation. Constructing the DNN as a graph structure and obtaining the graph embedding is the core part of this method. First, extract the network topology structure information from the model and incorporate the pruning pattern information into the graph structure, and then obtain the graph embedding representation of the model through the GNN method. The graph constructed by the DNN can be expressed as , where is the node set,[[]] is the edge set. The subset represents the nodes associated with the weight tensors,[[]] is the input tensor node set,[[]] is the output node set. The subsets and correspond to the edges linked to . Nodes in different layers of the DNN can be further divided into finer-grained sets. For example, the nodes in the first layer are included in . Next, take two typical DNN architectures: CNN and Transformer as examples to illustrate the graph construction process.

[0058] (1) CNN graph construction process.

[0059] Given a regular CNN , use a set of weights to parameterize the architecture of the CNN, where represents the weight tensor corresponding to the th convolutional kernel in the th convolutional layer,[[]] is the number of convolutional layers,[[]] is the th layer,[[]] is the number of convolutional kernels in the and .

[0060] The specific graph construction process is as follows: First, map the weight set of the layer to the node set , where the node represents the weight of the th convolutional kernel. At the same time, map the input and output feature maps of the layer to the nodes and respectively. Then map the dependency relationship of the convolutional operation to the connection relationship between the node set and the node sets to form the edge set in the graph and . After completing the above steps, the graph is obtained, and the process is as shown. Since the parameters of the CNN model are mainly located in the convolutional layer, and the final fully connected layer is connected to the output layer, the focus in the graph construction process is on the convolutional layer. After constructing the basic structure of the graph Figure 3 , integrate the pruning pattern

[0061] (as shown by PatternsLibrary determined by the Agent in ) into . Through research, it is found that fusing key information into edge features produces better results than fusing it into node features. Therefore, the following fusion method is adopted: fuse the pattern Figure 2 into the edge features of the edge set , and at the same time fuse the weights corresponding to the node set into their node features. In addition, the node features in the node sets and and the edge feature embeddings in the edge set and are assigned using random initialization. In addition, in CNN architectures such as ResNet, there are residual connections. These residual connections establish additional paths between layers, allowing information to flow across layers without being interrupted by pruning. Therefore, when constructing the graph, represent the residual connections as additional edges that link the nodes corresponding to the layers involved in the residual path. This ensures that the graph

[0062] accurately reflects the complete structure of the CNN, including these key connections, which is crucial for maintaining the performance of the CNN model during pruning. (2) Transformer graph construction process.

[0063]

[0064] For the DNN composition containing the Transformer structure, take the Transformer encoder as an example. To model the Transformer encoder assuming the input embedding of the th encoder is , where indicates that the input sequence consists of tokens, and is the embedding dimension. Then, the Q, K, V matrices can be expressed as , and , where is the learnable weight matrix, and is the respective output dimension. The attention scores are expressed as , and the output is .

[0065] Although the Transformer model structure is significantly different from that of CNN, a graph can also be constructed for it to achieve pattern pruning. For the th encoder, map , , to nodes respectively. Similarly, the input and output of the th encoder are mapped to and respectively. Then, the node set together with the associated edge set constitute the target graph . In addition, since the Transformer model contains multiple encoder and decoder blocks, and each encoder and decoder block contains a feed-forward network (MLP) that holds a large proportion of the model's parameters, this also needs to be considered when constructing the graph.

[0066] All in all, the process of constructing the DNN into a graph and fusing the pattern can be expressed as: (6) After the graph construction is completed, we need to encode the graph and extract its representative embedding , and this process can be expressed as: (7) Step 4, use DRL to generate pruning strategies, The reinforcement learning agent calculates the sampling probability of each pattern in the pattern library based on the graph embedding (State), assigns a new pruning pattern to the DNN, feeds back the reassigned pruning pattern (Action) to the environment, and re-evaluates whether the model meets the constraints; updates the pruning pattern library and gradually increases the number of patterns at each sparsity rate.

[0067] Environmental state The DRL training objective is to combine the model parameters of the DNN with the topological structure information to determine the pruning patterns of its operators. Therefore, we use the graph embedding obtained from the graph encoder as the environmental state and pass it to the agent. Since the pruning patterns applied to the DNN change after each iteration, it is necessary to reconstruct the graph to update the embedding representation of the DNN.

[0068] Action space The direct output obtained by the agent is the probability distribution of the patterns in the pattern library , so its action space is continuous: (8) where is the total number of predefined patterns. The action value is calculated using formula (9), and then the selected target pattern is sampled to obtain the pruned DNN.

[0069] (9) where represents the obtained graph representation embedding. The function is a multi-layer perceptron responsible for extracting graph representation information from . The activation function maps the output of to the interval to determine the selection probability of each pattern (i.e., the action space).

[0070] Finally, the categorical distribution function is applied to sample from the probability distribution to assign patterns to different weight tensors in the DNN. This will generate a newly selected pattern library for pruning.

[0071] Specifically, assume there are 10 patterns in the current pattern library (10 sparsity rates, each sparsity rate contains one pattern), sort these patterns in ascending order of sparsity rate and number them (0 - 9). Traverse the weight tensors of the DNN. For each weight tensor, the agent applies the categorical distribution function Categorical to sample from the probability distribution Perform medium sampling (as shown in Equation 10) to obtain a pattern number , that is, apply the pattern numbered to the current weight tensor, record the pattern information in the DNN meta-information, and at the same time, record the usage count of each pattern while traversing the weights. After the traversal is completed, the pruning pattern library needs to be updated: calculate the sparsity rate with the highest usage count of patterns, and then add a pattern that does not exist in the current pattern library at this sparsity rate. If the number of patterns at this sparsity rate in the pattern library has reached the upper limit, add a pattern with the second highest usage count of patterns at the sparsity rate, and so on.

[0072] (10) Reward function . The design of the reward function is crucial for the training of DRL. Generally, a higher pruning rate will lead to a more severe decline in the inference accuracy of the model. As the training progresses, the Agent tends to reduce the pruning rate to achieve better inference accuracy. Therefore, we design the reward function by combining a model compression metric (taking FLOP as an example, but other metrics such as multiply-accumulate operation MAC can be applied similarly), and the model inference accuracy, as follows:[[]] (11) where is a learnable parameter. When is large, encourage the Agent to adopt a more aggressive compression strategy, giving priority to model compression rather than accuracy. On the contrary, when is small, the Agent will be motivated to adopt a more conservative approach.

[0073] In addition, Equation (11) can be optimized by replacing because both FLOPs and inference latency are constraints for pruning and can be used as part of the reward function.

[0074] Step 5, perform pattern pruning and fine-tuning on the DNN When the pruning model meets the constraint conditions, perform pruning on the model (that is, replace the model parameters with 0 according to the assigned pattern) to obtain the pruned DNN. The inference accuracy of the pruned model will decline, and fine-tune the pruned DNN to restore its inference accuracy.

[0075] ​Among them, according to the pruning pattern output by the policy module, a corresponding binary matrix is used for local pruning of each convolutional layer or operation unit, and the Hadamard product operation is directly performed on the weight matrix, so as to achieve local parameter zeroing; after the pruning operation is completed, the entire model is retrained using a preset fine-tuning training scheme (such as SGD or Adam optimizer, setting appropriate learning rate and decay strategy) to ensure that the performance of the pruned model is restored to close to or exceeding the original level.

[0076] The specific steps are as follows: (1) Pruning pattern application, Input: The weight matrix of the original DNN model , and the pattern library output by the pruning policy module (including the binary mask matrix).

[0077] Operation: 1) Traverse each weight tensor in the DNN (such as the convolutional kernel or the weight of the fully connected layer), and select the corresponding binary mask matrix from the pattern library according to the pattern number assigned by the Agent (for the case where the convolutional kernel size is ).

[0078] 2) For each weight tensor , perform the element-wise Hadamard product operation, and the formula is as follows: (12) Where represents element-wise multiplication, and the weights corresponding to the positions with a value of 0 in the mask matrix are set to zero to achieve pruning.

[0079] (2) Global sparse constraint check, Calculate the overall sparsity rate of the pruned model: (13) If the sparsity rate does not reach the preset threshold (for example, 60%), return to step 4 to readjust the pruning strategy; otherwise, enter the fine-tuning stage.

[0080] (3) Fine-tuning to restore accuracy, Optimizer setting: Use the Adam optimizer, and set the initial learning rate to , and decay to 0.5 times the original every 10 epochs.

[0081] Loss function: Use the multi-task loss function, combining categorical cross-entropy (CE) and bounding box regression loss (SmoothL1), and the formula is as follows: (14) Training strategy: 1) Freeze the sparse weight structure after pruning and only fine-tune the retained weights.

[0082] 2) Use a subset (20% - 30%) of the original training data for fast fine-tuning and iterate for 5 - 10 epochs.

[0083] 3) Monitor the validation set accuracy. If the accuracy recovers to more than 98% of the original model, terminate the fine-tuning; otherwise, extend the training period.

[0084] (4) Pruning stability verification Conduct robustness tests on the fine-tuned model, including adding noise, occlusion simulation, etc., to ensure that pruning does not introduce sensitive vulnerabilities.

[0085] Step 6, compile and optimize the pruned model After going through the above steps, the DNN still relies on large deep learning frameworks such as PyTorch or TensorFlow for inference. However, deploying the pruned model on smartphones or edge devices equipped with MCUs requires an AI compilation framework. Use it to convert the deep learning model into efficient, low-level executable code that can run on different hardware platforms. This compilation process applies various optimization strategies to improve execution performance. The proposed model design method can better facilitate the role of the AI compilation framework and be better applied in practice.

[0086] Among them, export the intermediate representation of the pruned model and use the TVM open-source compiler for graph optimization, perform operator fusion, memory optimization, and parallelism adjustment; adopt the AutoScheduler to adjust the kernel size, parallel strategy, and data caching method for the target hardware platform to fully reduce the inference latency; compile and generate the final low-level code or executable file, and support hot update or online deployment mechanisms for easy testing and optimization in various hardware environments.

[0087] The specific steps are as follows: (1) Intermediate representation (IR) export Convert the fine-tuned DNN model to a general intermediate representation (ONNX format) and strip the framework dependencies.

[0088] (2) TVM compiler graph optimization 1) Operator fusion: Identify consecutive operations (such as Conv - BN - ReLU) in the computation graph and merge them into a single composite operator to reduce memory access overhead.

[0089] 2) Memory optimization: Analyze the tensor life cycle, reuse the memory space, and reduce the peak memory occupancy.

[0090] 3) Parallelism adjustment: Automatically divide the data parallel dimension (such as batch size or number of channels) according to the number of parallel computing units of the target hardware (such as GPU or NPU).

[0091] (3) Auto Tuning (AutoTVM), 1) Definition of parameter search space: Kernel Size: Sample candidate configurations within a legal range (such as 3×3, 5×5).

[0092] Data tiling strategy: Adjust the data block division method to match the hardware cache hierarchy.

[0093] 2) Tuning based on cost model: Use reinforcement learning or genetic algorithms to evaluate the inference latency and resource consumption of different configurations, and select the Pareto optimal solution.

[0094] (4) Low-level code generation, Compile the optimized computational graph into efficient low-level code (such as CUDA kernels or ARM assembly) for the target hardware platform.

[0095] Support output in the format of dynamic libraries (.so) or executable files (.bin), and adapt to the runtime environment of edge devices.

[0096] (5) Deployment and hot update mechanism, 1) Online deployment: Push the compiled model to the computing unit of the autonomous vehicle through OTA (Over-the-Air Technology).

[0097] 2) Hot update: Design a lightweight incremental update protocol to replace only the changed weight blocks in the pruned model, reducing the transmission overhead.

[0098] 3) Performance monitoring: Collect metrics such as inference latency and memory occupancy in real time on the in-vehicle terminal, and feedback them to the cloud for subsequent optimization and iteration. Generally, in the present invention, the deep neural network model is first converted into a graph structure, and the topological structure information and pruning mode information of the neural network are incorporated into the graph structure. Then, through the graph encoder integrating the pruning mode and the policy search based on deep reinforcement learning, a pruning mode adapted to the network structure is automatically generated, and the mode is applied to the original model. After multiple iterations, when the pruning strategy can meet the pruning constraint conditions, pruning operations are performed on the model, and then after fine-tuning, a pruned model that not only meets the accuracy requirements but also significantly reduces the computational amount is obtained. Finally, it is optimized and deployed to edge devices through a compiler.

[0099] S5. Deploy the pruned model to an autonomous vehicle, where the hardware platform selection is as follows: Select a computing platform suitable for autonomous vehicles (taking NVIDIA Jetson Orin NX 8GB as an example), and use the GPU acceleration supported by it for optimized inference.

[0100] Deployment process: 1) Load the quantized DNN model into the in-vehicle computing platform.

[0101] 2) Collect image data through sensors such as cameras, and input the images into the pedestrian detection model.

[0102] 3) The model will output the detection boxes and class information of pedestrians, and transmit them to other parts of the perception module (such as data fusion and decision-making modules).

[0103] Testing: Verify the accuracy, real-time performance, and robustness of the model in a test environment (simulator or actual road test) to ensure that it can work stably in various complex scenarios (such as light changes, occlusions, dynamic pedestrians, etc.).

[0104] Continuous optimization and maintenance: 1) Model update: The model can be updated regularly by collecting new data and performing online learning to ensure high-precision pedestrian detection in different environments.

[0105] 2) Performance monitoring: Perform performance monitoring through the in-vehicle computing platform to ensure that the inference time and resource consumption of the model are within an acceptable range.

[0106] Embodiment 2 This embodiment provides a pedestrian detection system based on deep learning automated pruning. The following further describes the present invention in conjunction with the attached Figures 1-5 illustrations and embodiments: (I) System architecture, The deep learning model automated pattern pruning and deployment system for heterogeneous edge devices includes a pruning pattern library generation module 100, a graph construction and encoding module 101, a deep reinforcement learning pruning strategy search module 102, an execution pruning and fine-tuning module 103, and a compilation and deployment module 104. As Figure 1 shown below, the following specifically describes each part: The pruning pattern library module 100: According to the network parameter statistical information and weight distribution rules, various pruning patterns are formed, covering multi-level pruning strategies from high sparsity rates to low sparsity rates, aiming to balance the pruning granularity and the preservation of key network structures.

[0107] Graph Construction and Encoding Module 101: Convert the pre-trained DNN model into a graph structure representation, abstractly describe the relationships between layers, weight dependencies, and operation processes in the model; encode the constructed graph to extract global topological information and local structural features, providing accurate environmental state information for the decision-making of subsequent pruning strategies, such as Figure 2 。

[0108] Deep Reinforcement Learning Pruning Strategy Search Module 102: Based on the graph embedding information obtained from graph construction and encoding, conduct strategy search for pruning patterns to automatically determine which pruning pattern to adopt for each layer; use the method of deep reinforcement learning to continuously adjust the pruning strategy according to the feedback of the model performance after pruning (such as FLOPs reduction rate, accuracy retention degree, inference latency), achieving an automatic balance between compression effect and model performance, such as Figure 3 。

[0109] Pruning and Fine-tuning Execution Module 103: According to the best pruning strategy obtained from deep reinforcement learning search, apply the predefined pruning pattern to the DNN weights to form a pruned network structure; conduct fine-tuning training on the pruned network to restore or further improve the model accuracy, ensuring that the performance of the compressed model does not decline in actual tasks, such as Figure 4 。

[0110] Compilation and Deployment Module 104: Compile and optimize the pruned and fine-tuned deep neural network model into low-level code to generate efficient execution code suitable for target hardware (such as edge devices, embedded systems); achieve fast inference of the model on the actual deployment platform, reduce latency, and fully utilize the acceleration characteristics of the hardware, such as Figure 5 。

[0111] A computer-readable storage medium storing multiple instructions suitable for being loaded and executed by a processor of a terminal device for the described pedestrian detection method based on deep learning automated pruning.

[0112] A terminal device including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions suitable for being loaded and executed by the processor for the described pedestrian detection method based on deep learning automated pruning.

[0113] The above are all preferred embodiments of the present invention. Without limiting the protection scope of the present invention accordingly, therefore: All equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A pedestrian detection method based on deep learning automatic pruning, characterized in that: include: Obtain pedestrian image data; Preprocessing the acquired pedestrian image data; Build a deep neural network DNN model; Prune the DNN model based on the pattern pruning strategy; The pruned DNN model is deployed on an autonomous vehicle for pedestrian detection.

2. The pedestrian detection method based on deep learning automatic pruning according to claim 1, characterized in that: The method prunes the DNN model based on the pattern pruning strategy, including initializing the pruning pattern library and evaluating the DNN model compression parameters; Model the DNN model as a graph structure and obtain the graph embedding representation; Use DRL to generate pruning strategies based on graph embedding representation; After executing the pruning strategy on the DNN model, the compiler is used to perform compilation optimization to obtain the pruned and optimized DNN model.

3. The pedestrian detection method based on deep learning automatic pruning according to claim 2 is characterized in that: The method of initializing the pruning pattern library and evaluating the compression parameters of the DNN model includes designing a pattern library graded by sparsity rate in combination with the distribution law of the weight parameters. The sparsity rate of the patterns in the pattern library uniformly decreases from 100% to 0%. The sparsity interval Δ is set to control the sparsity granularity. Δ=9 generates 10 sparsity rates corresponding to 10 types of patterns. The compression parameters of the DNN model are calculated to determine whether the constraints are met, and if the constraints are not met, the pruning strategy is continued to be searched.

4. The pedestrian detection method based on deep learning automatic pruning according to claim 3 is characterized in that: The method of modeling the DNN model as a graph structure and obtaining a graph embedding representation includes extracting network topology information from the DNN model, integrating pruning pattern information into the graph structure, obtaining a graph embedding representation of the model through a GNN method, and obtaining a graph embedding representation of the model through a DNN. Constructed graph Expressed as ,in is a node set, is an edge set, a subset represents the node associated with the weight tensor, is the set of input tensor nodes, is the set of output nodes, subset and Corresponds to link to The nodes in different layers of the DNN are further divided into more fine-grained sets.

5. The pedestrian detection method based on deep learning automatic pruning according to claim 4 is characterized in that: The method uses DRL to generate a pruning strategy based on the graph embedding representation, including the graph embedding obtained from the graph encoder As the environment state is passed to the Agent, the direct output obtained by the Agent is the pattern in the pattern library The probability distribution of , so its action space Continuous is continuous, and then by calculating the action value, the selected target mode Sampling is performed to obtain the pruned DNN model; wherein, the function As a multi-layer perceptron, it is responsible for extracting graph representation information from graph representation embedding. The activation function will The output is mapped to interval, determine the selection probability of each mode, that is, the action space, and finally, apply the classification distribution function from the probability distribution Sampling in the DNN, assigning patterns to different weight tensors in the DNN, producing a library of newly selected patterns Used for pruning.

6. The pedestrian detection method based on deep learning automatic pruning according to claim 5, characterized in that: The method of using DRL to generate a pruning strategy based on a graph embedding representation also includes setting a reward function by combining a model compression metric and a model reasoning accuracy, which is expressed as: in is a learnable parameter. When it is greater than the set threshold, the Agent is encouraged to adopt an aggressive compression strategy, giving priority to model compression rather than accuracy. On the contrary, when When it is less than the set value, the agent will be motivated to take a conservative approach.

7. The pedestrian detection method based on deep learning automatic pruning according to claim 6, characterized in that: After executing the pruning strategy on the DNN model, the compiler is used to perform compilation optimization, including performing a pruning operation on the model after the pruned DNN model meets the constraint conditions, replacing the model parameters with 0 according to the assigned pattern to obtain the pruned DNN model, and at the same time, because the reasoning accuracy of the pruned DNN model will decrease, the pruned DNN model is fine-tuned to restore its reasoning accuracy.

8. The pedestrian detection method based on deep learning automatic pruning according to claim 7, characterized in that: After executing the pruning strategy on the DNN model, the compiler is used to perform compilation optimization, and the TVM open source compiler is used to perform graph optimization, perform operator fusion, memory optimization and parallelism adjustment; an automatic tuning strategy is used to adjust the kernel size, parallel strategy and data caching method for the target hardware platform to fully reduce the inference delay; Compilation generates the final low-level code or executable file, and supports hot update or online deployment mechanism, which facilitates testing and optimization in various hardware environments.

9. The pedestrian detection method based on deep learning automatic pruning according to claim 8, characterized in that: The pruned model is deployed on the autonomous driving car for pedestrian detection, including deploying the pruned DNN model in the perception module, which is responsible for extracting environmental information from sensor data and making detection results.

10. A pedestrian detection system based on deep learning automatic pruning, characterized in that: include: The data acquisition module is configured to acquire pedestrian image data; A preprocessing module is configured to preprocess the acquired pedestrian image data; The model building module is configured to build a deep neural network DNN model; The pruning module is configured to prune the DNN model based on the pattern pruning strategy; The detection module is configured to deploy the pruned DNN model to the autonomous vehicle for pedestrian detection.

Citation Information

Patent Citations

  • Agricultural pest identification method based on deep learning of binary convolutional neural network

    CN108304844A

  • Model optimization algorithm for target detection YOLOv3 based on deep learning

    CN112001477A

  • Pedestrian detection method based on channel pruning yov5s

    CN115527086A

  • Structured pruning method for deep pedestrian search model

    CN117217282A

  • Pedestrian detection method based on knowledge migration pruning model

    CN117315722A

Cited By

  • Multi-scene image style migration and edge calculation optimization method

    CN120953099A