Automatic driving target detection method based on self-distillation auxiliary dynamic neural network

By adopting a dynamic neural network architecture based on self-mutual distillation in the autonomous driving system, the problem of object detection in an environment without teacher guidance and resource-constrained is solved, and efficient classification performance and inference speed are achieved.

CN120032333AActive Publication Date: 2025-05-23NANJING TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510120242.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-23
Estimated Expiration
2045-01-25

AI Technical Summary

Technical Problem

Existing dynamic neural networks are difficult to achieve efficient object detection without teacher guidance and edge network support in autonomous driving systems, especially on resource-constrained mobile platforms.

Method used

Using a dynamic neural network architecture based on self-mutual distillation assistance, through the coordinated optimization of the mixed distillation network and the policy network, the self-learning and self-enhancement of the detection network are achieved, and the network structure is dynamically adjusted to adapt to different environments and goals.

Benefits of technology

Without teacher guidance, the classification performance and inference speed of the detection network are significantly improved, the image processing capability of different instances is adapted to the inference delay, and the detection accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032333A_ABST
    Figure CN120032333A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving target detection method based on a self-distillation auxiliary dynamic neural network. Collected images serve as input of a target detection model to achieve target classification. The target detection model is based on a self-mutual distillation assisted dynamic neural network and is composed of a strategy network, a detection network and a mixed distillation network; the detection network is formed by connecting a group of multi-branch residual blocks in series; the detection network takes the collected image as input and outputs a target classification result; the strategy network generates corresponding routing vectors for different inputs, and determines to open and close residual blocks in the detection network; the mixed distillation network comprises two sub-networks; the two sub-networks are differentiated from the detection network in an initial training stage, and parameters of the two sub-networks are randomly initialized; the mixed and distilled sub-networks are fused into a fusion network; and through knowledge distillation, the knowledge of the fusion network is compressed and migrated back to the detection network, and one-time closed-loop optimization is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of artificial intelligence technology in autonomous driving, and specifically to an autonomous driving target detection method based on self-distillation assisted dynamic neural network. Background Art

[0002] Deep Neural Network (DNN) plays a vital role in autonomous driving systems, especially in key links such as perception, decision-making and execution. Autonomous driving systems need to be able to process sensor data in real time and respond quickly. This requires neural network models to take into account both accuracy and real-time performance. Dynamic neural networks[1] and knowledge distillation[2][3][4] can effectively improve the reasoning speed of neural networks while maintaining model accuracy. Dynamic neural networks can dynamically adjust their structure and parameters based on input data to improve computational efficiency and representation capabilities. Knowledge distillation improves the performance of student models by transferring knowledge from large and complex teacher models to small student models. Combining knowledge distillation and dynamic neural networks can further improve the performance and adaptability of dynamic networks.

[0003] Most existing dynamic neural network solutions rely on pre-trained detection networks based on traditional supervised learning, and their upper limit of classification ability (when performing full network reasoning) has been solidified. Under the knowledge distillation framework, the student model of the detection network can break through this performance bottleneck without increasing the model size by learning the generalization ability and feature representation of the teacher model. However, in many cases, the detection model deployed on self-driving cars or drones cannot obtain guidance from the teacher model, nor can it obtain model downloads and updates from the edge network. In order to adapt to this harsh environment, the detection network as a student must have the ability of self-learning and self-enhancement. This puts higher requirements on the design of dynamic neural network distillation methods and the coordinated optimization of strategy-detection networks, which is reflected in the following aspects:

[0004] (1) Endogenous knowledge transfer and complementarity. In many cases, mobile platforms such as vehicles or drones are limited by environmental factors and their own computing resources. They are unable to download or update models from the edge, nor can they find a suitable teacher model for support. As emphasized in references [5] and [6], it is difficult to select a suitable teacher model with traditional knowledge distillation. From the perspective of vehicle mobility and security threats, mutual learning between vehicles is also difficult to achieve in reality. This drives us to explore a self-learning dynamic neural network solution to adapt to multi-classification tasks on mobile platforms without teacher guidance and infrastructure assistance to update models.

[0005] (2) Coordination between flexibility and reusability of detection network structure. Dynamic reasoning requires that the detection network ensures the rationality and effectiveness of the structure while maintaining flexibility. In order to improve reasoning efficiency and scalability, the design of dynamic networks often needs to be modular so that the same network modules can be reused on different tasks and datasets. LC-Net[7] is a reusable framework that reduces the reasoning cost of deep neural networks in resource-constrained scenarios by dynamically skipping redundant layers and channels. However, there are large differences in the computing resources of different models of autonomous vehicles and drone platforms. The input instances of mobile platforms also vary greatly in complexity. This requires that the detection network must be customizable and able to quickly customize a model that is suitable for itself. Therefore, it is worth further exploring the design of a flexible and reusable modular structure and ensuring the stability of its reasoning performance under different instances.

[0006] (3) The trade-off between inference cost and classification accuracy. Dynamic networks need to adjust their structures or parameters during the inference phase to adapt to changes in the clarity and category of the input image. Alex et al. [8] introduced an additional network layer (called a "gating mechanism") to determine the switching of each layer of the network. This gating mechanism evaluates the current input and past calculation results, and then decides whether to continue to perform more calculation steps. However, policy network reasoning also requires computational costs, and overly complex routing decisions will affect model efficiency. Researchers need to develop new optimization techniques to train networks that contain discrete decisions. Summary of the invention

[0007] For environments without teacher guidance and infrastructure to assist in updating models, the present invention uses images collected by autonomous driving vehicles (through on-board image acquisition equipment) as input to the on-board target detection model to achieve target detection. The architecture of the target detection model of the present invention is based on a self-distillation-assisted dynamic neural network. The technical solution of the present invention is as follows.

[0008] A self-distillation-assisted dynamic neural network-based autonomous driving target detection method uses the images collected by the autonomous driving mobile platform as the input of the mobile platform's target detection model to achieve target recognition and classification;

[0009] The architecture of the object detection model is based on a self-distillation-assisted dynamic neural network, which consists of a policy network, a detection network and a hybrid distillation network;

[0010] The detection network is composed of a series of multi-branch residual blocks. It takes the collected image as input and outputs the target classification result. The policy network is strengthened through back propagation. The policy network generates a routing vector based on the input to determine the routing selection of the detection network.

[0011] The hybrid distillation network contains two subnetworks; the two subnetworks are differentiated from the detection network in the initial stage of training and their parameters are randomly initialized; the subnetworks that have undergone hybrid distillation are fused into a fusion network; through knowledge distillation, the knowledge of the fusion network (i.e., the parameter weights and feature representations in the model) is compressed and migrated back to the detection network, completing a closed-loop optimization.

[0012] The mobile platform includes a vehicle or a drone.

[0013] Compared with the prior art, the present invention proposes a target detection model that uses an instance-adaptive dynamic neural network architecture based on hybrid distillation, which takes into account the detection accuracy, reasoning speed and scalability of multi-instance classification tasks under resource-constrained conditions. According to the task features extracted by the policy network, the detection network dynamically selects network routes to cope with the diversity and changes of the environment and targets in motion.

[0014] The main contributions and innovations of the present invention include:

[0015] First, the present invention designs a training and parameter fusion framework based on hybrid distillation, which includes two stages: mutual distillation and self-fusion distillation. For the former, two sub-networks with the same structure enhance the diversity of each other's parameters in the form of mutual distillation. For the latter, the enhanced two sub-networks are fused into a more powerful fusion model. The knowledge of the model is distilled and transferred back to the detection network, improving its classification performance without increasing the size of the detection network.

[0016] Second, in order to realize dynamic reasoning on the detection network, the present invention develops a lightweight policy network based on curriculum learning. For different inputs, the routing vector it generates is mapped to a detection network composed of multi-branch residual blocks, which determines the opening and closing of the residual blocks, thereby realizing the dynamicization of the model. Different from the traditional "one-step-one-decision" routing mode, the policy network proposed in the present invention outputs a routing vector that matches the task attributes at one time during the policy generation stage, without participating in subsequent reasoning, while reducing the reasoning cost of the policy and detection network.

[0017] Third, the experimental results on the CIFAR and ImageNet datasets confirm that compared with traditional mutual learning and ensemble learning, the hybrid distillation proposed in the present invention is more effective in improving the classification performance of the detection network. The proposed dynamic neural network method can maintain the sparsity of the network channel when processing images containing different instances. At the same reasoning cost, the detection accuracy of the proposed method is higher than that of early withdrawal and random depth. When the accuracy level is consistent, the reasoning delay is reduced by 36% and 42% compared with early withdrawal and pruning. Compared with route generation that relies solely on curriculum learning (CL), the proposed joint training brings 17% detection accuracy and 17% reasoning speed improvement. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Representation of instance-adaptive dynamic neural network framework enabled by hybrid distillation;

[0019] Figure 2 represents the inter-distillation framework;

[0020] Figure 3 represents the self-fusion distillation framework;

[0021] Figure 4 represents a hybrid distillation framework;

[0022] Figure 5 Represents a multi-branch residual structure;

[0023] Figure 6 Representation strategy-detection network joint training framework;

[0024] Figure 7 Indicates representative picture instances;

[0025] Figure 8 Routing visualization showing a representative picture.

[0026] Figure 1 to Figure 6 Main terminology translation involved:

[0027] Policy network based on CL: Policy network based on CL.

[0028] Multi-branch residual detection network: Multi-branch residual detection network.

[0029] Hybrid distillation network: Hybrid distillation network.

[0030] Feature Fusion Module: Feature fusion module. DETAILED DESCRIPTION

[0031] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0032] 1 Overview

[0033] For resource-constrained vehicle target classification tasks, this paper proposes a detection model based on instance-adaptive dynamic neural network with hybrid distillation, which can adapt to harsh environments without teacher network guidance and edge network support. In the scenario without supervision and guidance of large-scale teacher models, this solution combines diversity enhancement strategies (mutual distillation and self-fusion distillation) to integrate and refine sub-network knowledge with the same structure as the detection network into the detection network, and further integrates dynamic routing to achieve dual optimization of detection accuracy and reasoning speed.

[0034] The present invention designs a weakly supervised joint training framework to achieve efficient collaboration between the strategy network and the detection network by jointly fine-tuning the two.

[0035] Experimental results verify that the proposed scheme improves classification accuracy and reduces model inference cost compared with other schemes, and visualization experiments verify the effectiveness of the dynamic strategy.

[0036] The subsequent content is arranged as follows:

[0037] Section 2 introduces the instance-adaptive dynamic neural network architecture assisted by hybrid distillation, the detection network and policy network it contains, and proposes a joint training method.

[0038] Section 3 describes the simulation experiments for performance evaluation.

[0039] Section 4 concludes the study and discusses future prospects.

[0040] 2 System Architecture and Solution Design

[0041] This section introduces the design ideas and implementation details of the instance-adaptive dynamic neural network assisted by hybrid distillation. Figure 1 As shown in the figure, the proposed framework consists of a policy network based on curriculum learning (CL), a multi-branch residual detection network and a hybrid distillation network. The hybrid distillation network contains two sub-networks (called subnet-1 and subnet-2), which are differentiated from the detection network in the initial stage of training and are randomly initialized with parameters. The sub-networks that have undergone hybrid distillation are fused into a more powerful model. Through knowledge distillation, the knowledge of the fused model is compressed and transferred back to the detection network to complete a closed-loop optimization. The scheme relies on a closed-loop optimization model of self-differentiation, self-supervision, and self-enhancement to improve classification performance without the need for a teacher model or infrastructure to assist in updating parameters.

[0042] 2.1 Hybrid Distillation Detection Network

[0043] In the initial stage, the dynamic neural network of the mobile platform only contains a policy network and a detection network. In the training preparation stage, the detection network differentiates into two sub-networks with the same structure as itself, and then starts the training based on hybrid distillation. The main process includes:

[0044] Mutual distillation: The two sub-networks learn from each other using the mutual distillation method to enhance each other’s parameter diversity, thereby improving the classification performance of the two sub-networks.

[0045] Self-fusion distillation: Use autoencoders to fuse two enhanced sub-networks to form a more powerful fusion network. Since the parameters of the two sub-networks are fused, its size and computational complexity are greatly improved.

[0046] Endogenous knowledge transfer: The fusion network compresses and transfers knowledge back to the detection network of the dynamic neural network through the self-fusion distillation method.

[0047] Next, the implementation details of mutual distillation and self-fusion distillation are introduced respectively.

[0048] (1) Inter-distillation

[0049] The two sub-networks in the mutual distillation are denoted as θ 1 and θ 2 ,like Figure 2 As shown. Let z c Represents the predicted probability of the network on category c. Network model θ 1 The prediction results are expressed as

[0050]

[0051] Where τ represents the distillation temperature. Network model θ 1 The regularized cross entropy of the generated soft labels and the true labels is calculated as

[0052]

[0053] Dependence (2), traditional supervised learning enables the network to correctly predict instance labels. In order to improve θ 1 To improve the generalization performance of 2 , is θ 1 Provide training experience. The prediction results of these two networks are expressed as σ 1 (z c ) and σ 2 (z c ), their matching degree is quantified by KL divergence. From σ 1 (z c ) to σ 2 (z c ) is calculated as

[0054]

[0055] Network θ 1 and θ 2 The total loss function is expressed as

[0056]

[0057] and

[0058]

[0059] In this way, the subnetwork θ 1 and θ 2 It not only grasps the correct true labels, but also learns the probability estimates of the parallel network.

[0060] (2) Self-fusion distillation

[0061] The two sub-networks after mutual distillation are used to generate a fusion model. A self-fusion distillation method is introduced to maximize the absorption of the knowledge of the two sub-networks. Figure 3 As shown, θ 1 and θ 2 The feature maps of the fusion model are connected in series to form a fusion model. Compared with a single sub-network, this fusion model has more diverse parameters and stronger classification performance. Then, the classification knowledge of the fusion model is refined and instilled into the detection network S 0 The feature maps of the Lth convolutional layers of the two sub-networks (F 1L and F 2L ) are connected, and the result after concatenation is recorded as F e . Autoencoder ω 1 Encode the concatenated feature maps into a meaningful compact feature map, denoted as F f , whose size is the same as the feature map of the Lth layer of the detection network.

[0062] F e The output result is represented as F f After being passed to the ensemble classifier, we use the true labels y and z e The classifier is trained under supervision. The soft output of the fusion classifier is represented as z f Based on equations (1), (2) and (3), the training objective of the fusion classifier is to minimize

[0063] L fusion =L CE (y,z f )+τ 2 D KL (z e ||zf ) (6)

[0064] Self-fusion distillation includes two stages: compressing the fused model knowledge and transferring the high-value knowledge. The output of the last layer of the detection network is a feature map, represented by F 0L We encourage F 0L From F f Then, we calculate the feature map of the de-decoded detection network output and the sub-network concatenated feature map F e The cross entropy loss is used to quantify the effectiveness of the feature map learned by the detection network. Regarding the feature map output by the detection network, the optimization goal is to minimize

[0065]

[0066] where ||·|| 2 is the norm, r(·) is used to align F 0L and F e The channel size.

[0067] Based on formula (7), the final goal of the detection network is expressed as minimizing

[0068]

[0069] Among them, z 0 is the soft output of the detection network, λ f is the weight factor of the fusion feature matching degree.

[0070] 2.2 Detection Network and Policy Network

[0071] Detection network:

[0072] Figure 4 The detection network in is composed of a set of multi-branch residual blocks connected in series. Figure 5 As shown in the figure, a 1×1 convolution kernel is used to reduce the dimensionality of the input, and a multi-branch structure is used to extract data features. Branch 1 and branch 2 use convolution kernels of sizes 3×3 and 5×5, respectively. The residual block can obtain a sum matrix by summing the feature matrices obtained from each branch, and then continue to use a 1×1 convolution kernel to transform the dimension of the sum matrix. Then, using the quick connection in the residual network, the matrix is ​​superimposed with the downsampled original features to obtain the output of the multi-branch residual block. When y i When input to the i-th multi-branch residual block, the output of the residual block is y i+1 =F i (y i )+y i , as the input of the next residual block.

[0073] For a residual network, skipping a residual block does not cause too much accuracy loss. Even when some residual blocks are removed, low-dimensional feature information can still be partially preserved [9]. Compared with single-path static networks (such as AlexNet

[10] and VGGNet

[11] ), there are many optional paths in the proposed detection network. When the residual module receives the "skip" instruction, the convolution kernel in the residual block will not participate in reasoning, which is equivalent to y i+1 =y i .

[0074] Policy Network:

[0075] To address the above issues, a CL-based policy network is introduced to gradually determine the action sequence through exploratory search.

[0076] At the beginning of training, CL lets the model start learning from easy samples, and gradually advances to complex samples and knowledge in the middle and late stages of training. CL assigns different weights to training samples according to complexity. In the initial training phase, simple samples are assigned the highest weights, and the weights of more difficult samples are gradually increased. For training samples, each optimization of CL uses different weights for weighting. Assuming that the original distribution of training samples is P(z), the weight assigned to each sample during the λth optimization is 0≤W λ (z)≤1, where 0≤λ≤1 and W 1 (z) = 1. In the λth optimization, the sample distribution is expressed as

[0077]

[0078] where ∫Q λ (z)dz=1. When λ=1, Q 1 (z) = P(z). Inequality (10) is used to ensure that Q λ (z) The information entropy increases monotonically;

[0079]

[0080] Inequality (11) is used to ensure that W λ (z) Monotonically non-decreasing.

[0081]

[0082] 2.3 Joint Training with Weak Supervision

[0083] Hybrid distillation enhances the diversity of detection network parameters and the upper limit of classification accuracy. The vectors generated by the policy network determine the detection network routing, and the classification results of the detection network are used to reinforce the policy network through back propagation. The model trained by CL learns a variety of routing strategies, but the detection accuracy is slightly lower than that of performing complete model reasoning. This drives us to explore a weakly supervised (no teacher guidance) joint training framework. The purpose is to improve the synergy of dynamic routing and detection networks and balance detection accuracy and reasoning speed. Figure 6 The proposed joint training workflow is given, including:

[0084] Knowledge transfer: In the initial stage of training, the detection network differentiates into two sub-networks with the same structure, and uses a hybrid distillation network to enhance the diversity of detection network parameters.

[0085] Policy generation: The CL-based policy network is responsible for extracting features of input instances and outputting routing vectors. The CL training requires the joint participation of the policy and detection networks.

[0086] Link jump: The strategy vector guides the multi-branch residual network to achieve dynamic jumps through shortcut connections;

[0087] Dynamic reasoning: Residual blocks in the detection network are selectively skipped to achieve dynamic reasoning;

[0088] Co-training: The detection network fine-tunes its parameters based on the classification results and inputs the results into the reward function in CL. The output reward value is back-propagated to the policy network, prompting it to learn a more appropriate route.

[0089] Different from random depth

[12] , the selection of residual blocks in the proposed scheme is controlled by the policy network to enhance the matching degree of the same detection task instance. For an input image, the policy network outputs all residual network routes at once, that is, predicts all actions of the detection network, which is essentially a single-step Markov decision process given an input state. Let s k ∈[0,1] is the kth element in the strategy vector s, and its value represents the probability that residual block k is turned on. Given an image x and a pre-trained detection network consisting of K multi-branch residual blocks, the strategy for selecting residual blocks is defined as a K-dimensional Bernoulli distribution, i.e.

[0090]

[0091] The policy network here is described as a function f(x; W) about the image x and the weight W. The image x is activated by the activation function σ(x) = 1 / (1+e -x ) The output after processing is represented as

[0092] s=f(x;W) (13)

[0093] A lightweight Resnet-8 model is used to build the policy network. The inference cost is largely determined by the number of convolution kernels. Due to the small number of convolution kernels, the inference cost generated by the policy network accounts for only 8% of the overall network. The policy network generates an action vector u based on s to determine the block involved in the inference. 0-1 variable u k Corresponding to the kth element in u, u k =1(u k =0) represents turning on (off) the kth residual block.

[0094] We customize a reward function to quantify the benefits of the action vector u and guide the policy network to find a high-precision and low-cost route. The training goal around this reward function is to minimize the length of the detection network route (the use of blocks) while ensuring that the prediction results are consistent with the true label. The reward function is defined as

[0095]

[0096] in, Represents the proportion of residual blocks that are enabled in the entire detection network. When the result is predicted correctly, the shorter the route length, the more positive rewards are given to encourage the policy network to skip more residual blocks. γ is used to penalize predictions with low similarity to the teacher network to balance inference speed and detection accuracy. Selecting more residual blocks is beneficial to improve detection accuracy, but will reduce inference speed. We maximize the reward expectation, i.e.

[0097]

[0098] To obtain the best policy network parameters for training.

[0099] In summary, the routing generated by formula (13) determines which blocks in the detection network perform forward propagation to predict the classification results; the policy network calculates the reward value based on whether the prediction is correct and the number of blocks used.

[0100] As a traditional reinforcement learning method, the policy gradient method samples from multiple distributions to obtain policy gradients, and is often used to search for the gradient that can maximize formula (15). Different from this, the present invention collects training policy samples from k-dimensional Bernoulli distribution. For u k ∈{0,1}, the policy gradient is calculated as

[0101]

[0102] In the mini-batch samples, Monte Carlo sampling is used to obtain the expected gradient of Equation (16).

[0103] After hybrid distillation training, the detection network improves the upper limit of detection accuracy and parameter diversity. The dynamic model framework can learn a variety of routing strategies after CL training, but the detection accuracy is slightly lower than that of performing complete model reasoning. To this end, after CL training, the policy network and the multi-branch residual network must be jointly fine-tuned to balance detection accuracy and reasoning speed. The execution of joint training is summarized as Algorithm 1. The policy network sets the first Kh variables in the policy vector s to 1, and gradually increases h to start CL training (lines 9-12). After CL training, the model explores a variety of routing strategies (line 17). To further optimize the model, the policy network and the detection network are jointly trained (lines 18-21).

[0104]

[0105]

[0106] 3 Experimental design and result analysis

[0107] In order to evaluate the detection and reasoning performance of the proposed method for different image complexities and instances, this experiment selected three authoritative datasets: CIFAR-10

[13] , CIFAR-100

[13] , and IMAGE-NET

[14] . The CIFAR dataset contains 60,000 32×32 RGB images, of which 50,000 and 10,000 images are used for training and testing. The ImageNet dataset contains 1.2 million training images of 1,000 categories, of which 50,000 are used as validation sets to test the top-1 accuracy.

[0108] The proposed method is implemented on the PyTorch platform, and the ADAM optimizer is used to train the model. The model training server is equipped with an Intel i9-13900k CPU and an NVIDIA GeForce RTX 4090 GPU. The distillation temperature τ in formula (1) is set to 2, and the hyperparameter λ in formula (8) is f is set to 10. The learning rate is 1×10-3. In the policy network (CL) training, the batch size is set to 2048. In the joint training, the batch size is adjusted to 256 and the learning rate is adjusted to 1×10-5.

[0109] Two multi-branch residual detection networks with similar convolution kernel numbers to Resnet50

[15] and Resnet110

[15] are constructed, named Inception15 and Inception54. As the basic student detection networks, they are composed of 15 and 54 multi-branch residual blocks respectively. We use them to build detection networks and test the effectiveness of dynamic routing.

[0110] 3.1 Analysis of the effectiveness of mixed distillation

[0111] The first set of experiments analyzes the role of hybrid distillation in promoting the classification accuracy of the detection network. Table 1 shows the results of the hybrid distillation experiment on CIFAR-100. The experiment compares the initial detection network without distillation (named: nodistillation), the sub-networks after mutual distillation (named: subnet-1 and subnet-2), the network after self-fusion (named: fusion) and the detection network enhanced by hybrid distillation (named: hybrid distillation) to verify the effectiveness of the proposed hybrid distillation. Among them, the classification accuracy of fusion is better than that of subnet-1 and subnet-2, which shows that fusing the feature maps of the two subnets is of great help to the classification accuracy. However, this improvement requires retaining all subnet parameters, which increases the storage and inference costs. Therefore, the knowledge in the fusion is compressed and migrated back to the detection network through knowledge distillation. On the Resnet20, Resnet50 and DenseNet40 models, hybrid distillation improves the classification accuracy by 2.44%, 1.81% and 2.59% respectively compared with nodistillation.

[0112] Table 1 Performance gains brought by mixed distillation

[0113]

[0114] This experiment selected traditional knowledge distillation (named: KD), mutual learning (named: DML) and ensemble learning (named: ONE) as benchmark methods and compared them with the hybrid distillation method proposed in this invention. As shown in Table 2, under ResNet-32, the classification accuracy of the proposed hybrid distillation method is 71.70%, which is higher than 70.32%, 71.15% and 71.07% of KD, DML and ONE, respectively, proving the superiority of hybrid distillation.

[0115] Table 2 Comparison of mixed distillation and benchmark distillation methods on CIFAR100

[0116]

[0117] Further experiments are conducted on the ImageNet dataset to demonstrate the effectiveness of the proposed method for complex image inputs. As shown in Table 3, the hybrid distillation on ResNet50 achieves a top-1 accuracy of 70.9%, which is better than DML and ONE. Compared with direct supervised training with ResNet50, the proposed hybrid distillation method can obtain a gain of 1.2%. This gain is better than other distillation models. The proposed scheme has achieved satisfactory results on CIFAR and Imagenet datasets of different complexity, indicating that the proposed hybrid distillation method has the potential to be applied to datasets containing other types of instances.

[0118] Table 3 Comparison of hybrid distillation and baseline distillation methods on ImageNet

[0119]

[0120] 3.2 Dynamic Routing Effectiveness Analysis

[0121] The second set of experiments compares the proposed dynamic routing with early exit

[16] and random depth

[17] . Table 4 lists the results of Iception15 and Iception54 on CIFAR10 and CIFAR100 datasets, where Acc represents the model detection accuracy and L represents the average residual block usage of the model.

[0122] Assume that the average length of the output route of the proposed method is L 1 The number of residual blocks used for the early-falloff and random-depth models is set to That is L 1 To ensure that the upper limit of the baseline model’s detection ability is not weaker than the proposed method. When running the early exit model, keep the previous When running the random depth model, randomly select residual block and keep it turned on.

[0123] On the CIFAR-10 dataset, the average length of the dynamic routing after CL training for Inception15 is 10.6, achieving an average accuracy of 88.1%. Compared with early exit and stochastic depth, it is improved by 72.9% and 68.8% respectively. It is worth noting that when running Inception54, nearly 15% of the images use less than 10 blocks, and some even less than 3. These results confirm that the proposed CL method not only improves the classification accuracy but also significantly reduces the inference cost. Neither pruning, stochastic depth nor early exit can achieve fine-grained dynamic adjustment of the neural network. Next, we examine the improvement in performance by joint training. On the CIFAR10 dataset, the Inception15 and Inception54 models after joint training not only improve the detection accuracy by 3.5% and 17.1% compared with the models relying solely on CL, but also reduce the average routing length by 3.2 and 3.9, demonstrating the efficiency and practicality of joint training.

[0124] Table 4 Influence of Routing Strategies on Detection Accuracy

[0125]

[0126] The proposed policy network outputs a complete routing vector each time. The model inference process does not need to refer to the intermediate output results, which helps to reduce the policy execution cost. To verify this inference, in this section, the proposed dynamic routing strategy is compared with the benchmark inference method that makes a decision step by step (named Single), which uses a traditional reinforcement learning training policy network. To ensure fairness, all methods are configured with the same number of residual blocks to observe the difference in inference speed of different methods under the same accuracy.

[0127] Table 5 Influence of Routing Strategies on Inference Speed

[0128]

[0129] Table 5 summarizes the average inference latency and acceleration effect of different models on the CIFAR-10 data. When achieving the same detection accuracy as Iception15, the proposed scheme Proposed improves the detection speed by 14.6% compared with Full-Net. On the other hand, Single reduces the detection speed by 28.9% compared with Full-Net, because Single uses a step-by-step decision-making inference method, and multi-step inference brings additional computational load, resulting in negative acceleration. These results confirm that the "one-time generation of routing" of the proposed policy network significantly improves the model inference speed.

[0130] 3.3 Visual Analysis of Dynamic Routing

[0131] This set of experiments aims to visualize the routing choices of dynamic neural networks during inference. This paper demonstrates the dynamic routing strategy of the Iception15 model on ImageNet. Figure 7 As shown in the figure, the present invention selects 8 representative image examples from the Image-Net dataset, and the ID and real label of the image are shown below the image. The above 8 image examples are input into our dynamic neural network model, and the routing vector generated by the policy network is visualized as Figure 8 Among them, the horizontal axis represents the number of multi-branch residual blocks, and the vertical axis is the image ID; gray (blank) represents the residual blocks that participate (do not participate) in routing.

[0132] Depend on Figure 8 It can be seen that, facing different input image instances, the dynamic neural network can give a variety of routing strategies, which proves the effectiveness and generalization of the dynamic routing strategy of the present invention.

[0133] 4 Conclusion

[0134] The present invention proposes an instance-adaptive dynamic neural network method based on hybrid distillation, which can dynamically select the optimal routing path in a multi-branch residual network according to the specific features of the input instance.

[0135] In the early stage of training, the present invention adopts a hybrid distillation strategy combining mutual distillation and self-fusion distillation to enhance the basic classification ability of the detection network.

[0136] Subsequently, a policy network was trained to generate a routing strategy suitable for the multi-branch residual network. This strategy significantly reduces the network's inference cost while maintaining high detection accuracy. The present invention also jointly fine-tunes the policy network with the multi-branch residual network to increase the diversity of the routing strategy, thereby further improving the detection accuracy and inference speed of the model.

[0137] Experimental results on CIFAR and ImageNet datasets show that the hybrid distillation proposed in this paper can significantly improve the performance of the detection network without the guidance of a large teacher model. In addition, visualization experiments further confirm the sparse characteristics of the network channels when processing different images, which verifies the effectiveness and adaptability of the dynamic routing method. While maintaining the same inference cost, the proposed method surpasses mainstream benchmark methods such as early exit, random depth, and pruning in terms of accuracy. With proper customization, our method has the potential to support most existing target classification applications and provide intelligent support for mobile platforms, thereby promoting the development of intelligent mobile systems.

[0138] References

[0139] [1]Han Y,Huang G,Song S,et al.Dynamic neural networks:A survey[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2021,44(11):7436-7456.

[0140] [2]Hinton G,Vinyals O,Dean J.Distilling the Knowledge in a NeuralNetwork[J].Stat,2015,1050:9.

[0141] [3]Li W,Wang J,Ren T,et al.Learning accurate,speedy,lightweight CNNsvia instance-specific multi-

[0142] teacher knowledge distillation for distracted driver postureidentification[J].IEEE Transactions on Intelligent Transportation Systems,2022,23(10):17922-17935.

[0143] [4]An S,Liao Q,Lu Z,et al.Efficient semantic segmentation via self-attention and self-distillation[J].IEEE Transactions on IntelligentTransportation Systems,2022,23(9):15256-15266.

[0144] [5]Jin X,Peng B,Wu Y,et al.Knowledge distillation via routeconstrained optimization[C] / / Proceedings of the IEEE / CVF InternationalConference on Computer Vision.2019:1345-1354.

[0145] [6]Cho J H,Hariharan B.On the efficacy of knowledge distillation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2019:4794-4802.

[0146] [7]Xia W,Yin H,Dai X,et al.Fully dynamic inference with deep neuralnetworks[J].IEEE Transactions on Emerging Topics in Computing,2021,10(2):962-972.

[0147] [8]Graves A.Adaptive Computation Time for Recurrent Neural Networks[J].2016.

[0148] [9]Huang G,Sun Y,Liu Z,et al.Deep networks with stochastic depth[C] / / Computer Vision–ECCV 2016:14th European Conference,Amsterdam,The Netherlands,October 11–14,2016,Proceedings,Part IV 14.

[0149] Springer International Publishing,2016:646-661.

[0150]

[10] Krizhevsky A,Sutskever I,Hinton G E.Imagenet classification withdeep convolutional neural networks[J].

[0151] Advances in Neural Information Processing Systems,2012,25.

[0152]

[11] Simonyan K,Zisserman A.Very deep convolutional networks forlarge-scale image recognition[C] / / 3rd International Conference on LearningRepresentations(ICLR 2015).Computational and Biological Learning Society,2015.

[0153]

[12] Yim J,Joo D,Bae J,et al.A gift from knowledge distillation:Fastoptimization,network minimization and transfer learning[C] / / Proceedings ofthe IEEE Conference on Computer Vision and Pattern Recognition.2017:4133-4141.

[0155]

[13] Krizhevsky A,Hinton G.Learning multiple layers of features fromtiny images[J].2009.

[0156]

[14] Deng J,Dong W,Socher R,et al.Imagenet:A large-scale hierarchicalimage database[C] / / 2009 IEEE Conference on Computer Vision and PatternRecognition.Ieee,2009:248-255.

[0157]

[15] He K,Zhang X,Ren S,et al.Deep residual learning for imagerecognition[C] / / Proceedings of the IEEE Conference on Computer Vision andPattern Recognition.2016:770-778.

[0158]

[16] Figurnov M,Collins M D,Zhu Y,et al.Spatially adaptive computationtime for residual networks[C] / / Proceedings of the IEEE Conference on ComputerVision and Pattern Recognition.2017:1039-1048.

[0160]

[17] Huang G,Sun Y,Liu Z,et al.Deep networks with stochastic depth[C] / / Computer Vision–ECCV 2016:14th European Conference,Amsterdam,TheNetherlands,October 11–14,2016,Proceedings,Part IV 14.

[0161] Springer International Publishing,2016:646-661.

Claims

1. An autonomous driving target detection method based on self-distillation assisted dynamic neural network, which uses the images collected by the autonomous driving mobile platform as the input of the target detection model of the mobile platform to achieve target recognition and classification; Its characteristics are The architecture of the object detection model is based on a self-distillation-assisted dynamic neural network, which consists of a policy network, a detection network and a hybrid distillation network; The detection network is composed of a series of multi-branch residual blocks. It takes the collected image as input and outputs the target classification result. The policy network is strengthened through back propagation. The policy network generates a routing vector based on the input to determine the routing selection of the detection network. The hybrid distillation network contains two subnetworks; the two subnetworks are differentiated from the detection network in the initial stage of training and their parameters are randomly initialized; the subnetworks that have undergone hybrid distillation are fused into a fusion network; through knowledge distillation, the knowledge of the fusion network is compressed and migrated back to the detection network, completing a closed-loop optimization.

2. The automatic driving target detection method according to claim 1, characterized in that The mobile platform includes a vehicle or a drone.

3. The automatic driving target detection method according to claim 1, characterized in that During the training of a dynamic neural network: a. In the initial stage, the dynamic neural network only contains a policy network and a detection network; b. In the training preparation stage, the detection network differentiates into two sub-networks with the same structure as itself; c. In the training phase based on mixed distillation, the process includes: 1) Mutual distillation: The two sub-networks use the mutual distillation method to learn from each other and obtain an enhanced sub-network; 2) Self-fusion distillation: Use autoencoders to fuse two enhanced sub-networks into a fusion network; 3) Endogenous knowledge transfer: The fusion network compresses knowledge through self-fusion distillation and transfers it back to the detection network.

4. The automatic driving target detection method according to claim 3, characterized in that In the inter-distillation of step 1), The two sub-networks in the mutual distillation are denoted as θ1 and θ2; let z 1,c and z 2,c Respectively represent the predicted probabilities of θ1 and θ2 in category c; The prediction result of θ1 is expressed as Where τ represents the distillation temperature, c' represents the category index, and C represents the set of categories; The regular cross entropy formula of the predicted result and the true label of category c is where y c represents the true label of category c; z c Represents the predicted probability of the subnetwork on category c; Then, the regular cross entropy between the soft label generated by θ1, that is, the prediction result and the true label of category c is In order to improve the generalization performance of θ1, another equivalent sub-network θ2 is introduced to provide training experience for θ1; The same method as θ1 is used to obtain the prediction result σ2(z 2,c ), and use formula (2) to get the regularized cross entropy between the soft label generated by θ2 and the real label of category c σ1(z 1,c ) and σ2(z 2,c ) is quantified by KL divergence; from σ1(z 1,c ) to σ2(z 2,c ) is calculated as Similarly: The total loss functions of θ1 and θ2 are expressed as and Then, θ1 and θ2 not only grasp the correct true labels, but also learn the probability estimates of the parallel network.

5. The automatic driving target detection method according to claim 3, characterized in that In the self-fusion distillation of step 2), the feature maps of the two sub-networks θ1 and θ2 are connected in series to form a fusion network; the classification knowledge of the fusion network Feature Fusion is refined and instilled into the detection network; specifically: The feature map F of the Lth convolutional layer of θ1 and θ2 1L and F 2L are connected, and the result after series connection is recorded as F e ; The autoencoder ω1 in the fusion network converts F e Encoded into a meaningful compact feature map F f ; F f The size of is the same as the feature map of the Lth layer of the detection network; F e The output of the fusion network is expressed as F f After being passed to the ensemble classifier, the true label y and predicted probability z are used e The ensemble classifier is trained under supervision; the soft output of the ensemble classifier is represented as z f ; The training goal of the fusion classifier is to minimize the fusion loss L fusion =L CE (y,z f )+τ 2 D KL (With e ||from f ) (6) Self-fusion distillation consists of two stages: the stage of compressing the model knowledge after fusion and the stage of transferring the high-value knowledge in the model knowledge; The detection network finally outputs the feature map F 0L , calculate F after the inverse decoder ω2 0L With F e The cross entropy loss is used to quantify the effectiveness of the feature maps learned by the detection network; Regarding the feature map output by the detection network, the optimization goal is to minimize F 0L With F e The cross entropy loss where ||·||2 is the norm and r(·) is used to align F 0L and F e The channel size; Based on formula (7), the final goal of the detection network is expressed as minimizing in, represents the overall loss of the detection network, z0 is the soft output of the detection network, and λ f is the weight factor of the fusion feature matching degree.

6. The automatic driving target detection method according to claim 1, Its characteristics are that for the detection network: The 1×1 convolution kernel in the residual block is used to reduce the dimensionality of the input, and the multi-branch structure is used to extract data features. Branch 1 and branch 2 use convolution kernels of sizes 3×3 and 5×5 respectively. The feature matrices obtained from each branch are summed to obtain a sum matrix, and then the 1×1 convolution kernel is used to transform the dimension of the sum matrix. Using the shortcut connection in the residual network, the sum matrix and the original features after downsampling are superimposed to obtain the output of the multi-branch residual block; the output y of the previous residual block is i When input to the i-th multi-branch residual block, the output of the residual block is y i+1 =F i (y i )+y i , as the input of the next residual block; The adjustment of the detection network architecture is determined by routing selection. That is, under the guidance of the strategy vector, the multi-branch residual network realizes dynamic routing through shortcut connections in the residual structure; when the residual block receives the "skip" instruction, the convolution kernel in the residual block will not participate in the reasoning, that is, y i+1 =y i .

7. The automatic driving target detection method according to claim 1, Its characteristic is that the policy network is a policy network based on curriculum learning CL; for the policy network: During training, CL selects the last h residual blocks from all K residual blocks for training; in the first round, h is 1. As h increases, it gradually associates and optimizes the switches of more residual blocks until all blocks are covered when h equals K. When the policy network associates h residual blocks, it keeps the first Kh blocks open and only learns the switch strategy of the last h blocks. In the early stage of training, start learning from easy samples; In the middle and late stages of training, the training samples and knowledge are gradually advanced to complex ones. According to the complexity, CL assigns different weights to the training samples. In the initial training phase, simple samples are assigned the highest weights, and the weights of difficult samples are gradually increased; for training samples, each optimization of CL uses different weights for weighting; Assume that the original distribution of the training samples is P(z), and the weight given to each sample in the λth optimization is 0≤W λ (z)≤1, where 0≤λ≤1 and W1(z)=1; at the λth optimization, the sample distribution is expressed as where ∫Q λ (z)dz=1; when λ=1, Q1(z)=P(z); Inequality (10) is used to ensure that Q λ The information entropy of (z) increases monotonically, Inequality (11) is used to ensure that W λ (z) monotonically non-decreasing, 8. The automatic driving target detection method according to claim 1, characterized in that A weakly supervised joint training method is used to improve the synergy between dynamic routing and detection networks, balancing detection accuracy and inference speed; The steps of joint training include: Knowledge transfer: In the initial stage of training, the detection network differentiates into two sub-networks with the same structure, and uses a hybrid distillation network to enhance the diversity of detection network parameters; Strategy generation: The CL-based policy network extracts the features of the input instance and outputs a policy vector as the routing vector of the detection network. The CL training is jointly participated by the policy network and the detection network. Link jump: The strategy vector guides the multi-branch residual network to achieve dynamic jumps through shortcut connections; Dynamic reasoning: residual blocks in the detection network are selectively skipped to achieve dynamic reasoning; Collaborative training: The detection network adjusts its own parameters based on the classification results and uses the results as the input of the reward function in CL; the reward value output by the reward function is back-propagated to the policy network, prompting it to learn a more appropriate route.

9. The automatic driving target detection method according to claim 8, characterized in that for joint training: For an input image, the policy network outputs all residual network routes at once, that is, predicting all actions of the detection network. This is a single-step Markov decision process given the input state. Order k ∈[0,1] is the kth element in the strategy vector s, and its value represents the probability that the residual block k is turned on. Given an image x and a pre-trained detection network consisting of K multi-branch residual blocks, the strategy for selecting the residual block is defined as a K-dimensional Bernoulli distribution, i.e. The policy network here is described as a function f(x; W) about the image x and the weight W; the image x is activated by the activation function σ(x) = 1 / (1+e -x ) The output result after processing is represented as s=f(x;W)(13) The lightweight Resnet-8 model is used to build the policy network; the policy network generates an action vector u based on s to determine the block involved in reasoning; the 0-1 variable u k Corresponding to the kth element in u, u k =1 and u k =0 represents turning on and off the kth residual block respectively; The benefit brought by the action vector u is quantified by a customized reward function, which guides the policy network to find a high-precision and low-cost route. The reward function is defined as in, Represents the proportion of enabled residual blocks in the entire detection network. When the result is predicted correctly, the shorter the route length, the more positive rewards are given to encourage the policy network to skip more residual blocks. γ is used to penalize predictions with low similarity to the teacher network, balancing inference speed and detection accuracy; Selecting more residual blocks is beneficial to improving detection accuracy, but it will reduce the reasoning speed. The optimal policy network parameters are obtained by maximizing the reward expectation. The reward expectation is expressed as Then, the routing generated by formula (13) determines which blocks in the detection network perform forward propagation to predict the classification results; the policy network calculates the reward value based on whether the prediction is correct and the number of blocks used; The training strategy samples are collected from the k-dimensional Bernoulli distribution, for u k ∈{0,1}, the policy gradient is calculated as In the mini-batch samples, Monte Carlo sampling is used to obtain the expected gradient of Equation (16).

Citation Information

Patent Citations

  • Automatic driving multi-target detection method based on instance adaptive dynamic neural network

    CN117237893A

  • Mobile platform multi-target classification method based on multi-teacher auxiliary instance adaptive DNN

    CN118097228A