Multi-object Detection Method, Apparatus, Device, Medium and Product
By dynamically pruning the target branch network of the multi-object detection model, the problems of insufficient accuracy in detection target standards and limited processor resources in the prior art are solved, and efficient and accurate multi-object detection is achieved.
Patent Information
- Application Number
- CN202210523846.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-13
AI Technical Summary
The existing object detection algorithms have the problem of insufficient accuracy when detecting targets of different scales, and it is difficult to achieve high-accuracy detection in scenarios with limited processor resources.
By dynamically pruning multiple target branch networks in the pre-trained multi-objective detection model, the pruned multi-objective detection model is obtained. The model includes a pruned backbone network and multiple pruned target branch networks, which can maximize the compression of model volume while maintaining detection accuracy.
In the scenario where processor resources are limited, multi-objective detection is maintained with high accuracy, and the model size and calculation amount are greatly reduced.
Smart Images

Figure CN114972950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and in particular, to a multi-object detection method, device, equipment, medium and product. Background Art
[0002] In recent years, object detection algorithms have made great breakthroughs and have been applied in scenarios such as autonomous driving and smart refrigerators. However, existing object detection algorithms still have significant limitations in some special application scenarios. Specifically, current object detection algorithms cannot well detect different objects of different scales, and the conventional convolution process treats different semantics of feature representations of different scales equally. That is to say, current object detection algorithms tend to ignore the differences between different objects of different scales, which is also one of the reasons for the slightly lower accuracy of multi-object detection algorithms.
[0003] In addition, due to considerations of processor cost and performance constraints in scenarios such as autonomous driving and smart refrigerators, the computational complexity of the multi-object detection model often has strict limitations. How to obtain relatively accurate results for the detection model with limited computing power has become one of the difficult problems. Summary of the Invention
[0004] The present invention provides a multi-object detection method, device, equipment, medium and product to solve the above problems.
[0005] The present invention provides a multi-object detection method, including:
[0006] Obtain a picture to be detected;
[0007] Input the picture to be detected into the pruned multi-object detection model to obtain the detection results of multiple objects respectively;
[0008] Wherein, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model;
[0009] The pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network;
[0010] Correspondingly, the pruned multi-object detection model is obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively.
[0011] According to a multi-object detection method provided by the present invention, the pruned multi-object detection model includes a pruned backbone network and multiple pruned object branch networks connected to the pruned backbone network;
[0012] Correspondingly, the multiple pruned target branch networks are obtained by dynamically pruning multiple target branch networks in a pre-trained multi-object detection model respectively;
[0013] The pruned backbone network is obtained by dynamically pruning the backbone network in a pre-trained multi-object detection model.
[0014] According to a multi-object detection method provided by the present invention, the pruned multi-object detection model is obtained by dynamically pruning multiple target branch networks in a pre-trained multi-object detection model respectively, including:
[0015] S1. Select a branch pruning rate from a preset set of branch pruning rates as the to-be-analyzed branch pruning rate, select a target branch network from the multiple target branch networks as the to-be-analyzed target branch network, and prune the to-be-analyzed target branch network using the to-be-analyzed branch pruning rate to obtain a pruned target branch network;
[0016] S2. Perform pruning sensitivity analysis on the pruned target branch network to obtain the performance of the pruned branch network;
[0017] S3. Repeat S1 to S2 until all branch pruning rates in the preset set of branch pruning rates are exhausted, to obtain multiple performances of the pruned branch networks;
[0018] S4. Determine the optimal performance of the pruned branch network from the multiple performances of the pruned branch networks, and use the to-be-analyzed branch pruning rate corresponding to the optimal performance of the pruned branch network as the optimal pruning rate of the to-be-analyzed target branch network;
[0019] S5. Repeat S1 to S4 until all target branch networks are exhausted, to obtain the optimal pruning rate corresponding to each target branch network, and prune the target branch networks respectively based on the corresponding optimal pruning rates to obtain the pruned multi-object detection model.
[0020] According to a multi-object detection method provided by the present invention, the preset set of branch pruning rates is obtained as follows:
[0021] According to the value range of the preset branch pruning rate and the value step of the branch pruning rate, all branch pruning rates that meet the value step of the branch pruning rate are exhaustively obtained within the value range of the preset branch pruning rate, so as to obtain the set of branch pruning rates.
[0022] According to a multi-object detection method provided by the present invention, the multiple pruned target branch networks include a pruned face detection branch network, a pruned cigarette detection branch network, and a pruned mobile phone detection branch network;
[0023] The detection results of the multiple targets include face detection results, cigarette detection results, and mobile phone detection results;
[0024] Correspondingly, inputting the picture to be detected into the pruned multi-target detection model to respectively obtain the detection results of multiple targets includes:
[0025] The pruned backbone network extracts features from the picture to be detected to obtain multi-scale feature maps;
[0026] The pruned face detection branch network predicts face detection results based on the multi-scale feature maps;
[0027] The pruned cigarette detection branch network predicts cigarette detection results based on the multi-scale feature maps;
[0028] The pruned mobile phone detection branch network predicts mobile phone detection results based on the multi-scale feature maps.
[0029] According to a multi-target detection method provided by the present invention, the pruned backbone network includes multiple convolutional layers;
[0030] Correspondingly, the pruned backbone network extracts features from the picture to be detected to obtain multi-scale feature maps, including
[0031] Each convolutional layer extracts features from the picture to be detected, thereby obtaining multiple feature maps with different scales;
[0032] According to the scales of the face prediction boxes in the face detection results, the cigarette prediction boxes in the cigarette detection results, and the mobile phone prediction boxes in the mobile phone detection results, respectively determine the face detection feature map, the cigarette detection feature map, and the mobile phone detection feature map from the multiple feature maps with different scales;
[0033] The pruned face detection branch network predicts face detection results based on the face detection feature map;
[0034] The pruned cigarette detection branch network predicts cigarette detection results based on the cigarette detection feature map;
[0035] The pruned mobile phone detection branch network predicts mobile phone detection results based on the mobile phone detection feature map.
[0036] The present invention also provides a multi-target detection device, including:
[0037] A picture acquisition module, configured to acquire a picture to be detected;
[0038] A detection module, configured to input the picture to be detected into the pruned multi-target detection model to respectively obtain the detection results of multiple targets;
[0039] Among them, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model;
[0040] The pre-trained multi-object detection model includes a backbone network and a plurality of target branch networks connected to the backbone network;
[0041] Correspondingly, the pruned multi-object detection model is obtained by dynamically pruning the plurality of target branch networks in the pre-trained multi-object detection model respectively.
[0042] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements any one of the above multi-object detection methods.
[0043] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above multi-object detection methods.
[0044] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above multi-object detection methods.
[0045] The multi-object detection method, device, equipment, medium, and product provided by the present invention obtain a pruned multi-object detection model by dynamically pruning the plurality of target branch networks in the pre-trained multi-object detection model respectively, so that each target branch network can be compressed to the maximum extent, and the detection accuracy corresponding to each target is ensured to be almost not lost. The pruned multi-object detection model can be applied to scenarios where processor cost and performance need to be considered. In addition, since the backbone network and the pruned target branch network are separated, and the multi-scale feature maps output in the backbone network are input into different pruned target branch networks, the differences between different targets at different scales can be concerned, and the detection accuracy of both large targets and small targets is relatively high. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic flowchart of the multi-object detection method provided by the embodiment of the present invention;
[0048] Figure 2 It is a schematic structural diagram of the pruned multi-object detection model provided by the embodiments of the present invention;
[0049] Figure 3 It is a schematic structural diagram of the multi-object detection device provided by the embodiments of the present invention;
[0050] Figure 4 It is a schematic physical structure diagram of an electronic device provided by the embodiments of the present invention. Specific embodiments
[0051] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0052] As one of the methods for model compression, model pruning can reduce the size and computational amount of the multi-object detection model, and at the same time, there is almost no loss in accuracy. However, the existing pruning schemes corresponding to the multi-object detection model set the same pruning ratio for different objects. Although it can make the overall compression performance of the multi-object detection model reach the best, it does not start from each object and fails to make the compression performance corresponding to each object reach the best state.
[0053] To solve the above problems, the embodiments of the present invention provide a multi-object detection method, which is specifically as follows.
[0054] Figure 1 It is a schematic flow diagram of the multi-object detection method provided by the embodiments of the present invention; as Figure 1 shown, a multi-object detection method includes the following steps:
[0055] S101, obtain the picture to be detected.
[0056] In this embodiment, the picture to be detected is an in-cabin picture in the application scenario of autonomous driving. In other embodiments of the present invention, the picture to be detected can also be a picture of the inside of a refrigerator in the application scenario of an intelligent refrigerator, or a picture formed by a road monitoring image in the application scenario of urban monitoring. The present invention does not make any limitations in this regard.
[0057] In addition, the picture to be detected can be obtained by on-site shooting, or can be obtained by decomposing a video stream, or can be a test picture obtained from a database in various multi-object detection application scenarios. The present invention does not make any limitations on the acquisition method of the picture to be detected.
[0058] S102. Input the image to be detected into the pruned multi-object detection model to obtain the detection results of multiple objects respectively.
[0059] Among them, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model; the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network.
[0060] Correspondingly, the pruned multi-object detection model is obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively.
[0061] In this step, the pruned multi-object detection model is obtained by dynamically pruning the pre-trained multi-object detection model, and the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network. Each object branch network is used to detect different objects. For example, when the image to be detected is an in-cabin image in an autonomous driving application scenario, the object branch network can be a mobile phone detection branch network, a key detection branch network, a tissue paper detection branch network, a glasses detection branch network, a face detection branch network, etc., for detecting different objects.
[0062] By performing targeted dynamic pruning on different object branch networks, specifically, based on multiple pruning rates, pruning sensitivity analysis is performed on each object branch network to determine the optimal pruning rate corresponding to each object branch network, and pruning is completed according to the optimal pruning rate corresponding to each object branch network, so as to ensure that the loss of object detection performance is small or even non-existent, and each object branch network can be compressed to the maximum extent.
[0063] After dynamically pruning different object branch networks, the pruned object branch networks are obtained. Therefore, the pruned multi-object detection model consists of a backbone network and pruned object branch networks. Among them, the backbone network extracts features from the image to be detected to obtain multi-scale feature maps and inputs them into different pruned object branch networks. Each pruned object branch network detects the corresponding object based on the multi-scale feature maps. Taking the above-mentioned image to be detected as an in-cabin image in an autonomous driving application scenario as an example, the detection results corresponding to each pruned object branch network are to locate various types of objects on the image to be detected through target boxes of different colors, and display the type name of the object and the probability belonging to that type.
[0064] In addition, the above-mentioned pre-trained multi-task object detection model is trained based on the labeled bounding boxes, corresponding object type labels, and training datasets.
[0065] The multi-object detection method provided by the embodiment of the present invention obtains a pruned multi-object detection model by dynamically pruning multiple target branch networks in a pre-trained multi-object detection model, so that each target branch network can be compressed to the maximum extent, and it is ensured that the detection accuracy corresponding to each target is hardly lost. The pruned multi-object detection model can be applied to scenarios that need to consider processor cost and performance. In addition, since the backbone network and the pruned target branch network are separated, and the multi-scale feature maps output in the backbone network are input into different pruned target branch networks, the differences between different targets at different scales can be concerned, and whether it is the detection of large targets or small targets, there is a high accuracy rate.
[0066] Further, the pruned multi-object detection model includes a pruned backbone network and a plurality of pruned target branch networks connected to the pruned backbone network.
[0067] Correspondingly, the plurality of pruned target branch networks are obtained by dynamically pruning multiple target branch networks in a pre-trained multi-object detection model respectively; the pruned backbone network is obtained by dynamically pruning the backbone network in a pre-trained multi-object detection model.
[0068] Based on the above embodiment, this embodiment also dynamically prunes the backbone network in a pre-trained multi-object detection model to obtain a pruned backbone network. At this time, the pruned multi-object detection model includes a pruned backbone network and a plurality of pruned target branch networks.
[0069] The dynamic pruning process of the backbone network is similar to the dynamic detection process of the target branch network. The pruning sensitivity analysis of the backbone network is carried out based on multiple pruning rates, so as to determine the optimal pruning rate corresponding to the backbone network, and the backbone network is pruned based on the optimal pruning rate to obtain the pruned backbone network. Specifically, after the pruning of the target branch networks is completed, the backbone network is dynamically pruned, that is, a pruning rate is sequentially selected from a preset pruning rate set to prune the backbone network, and then the test set is used to evaluate the accuracy rate and average precision of each target detection task after pruning, and the average value of the accuracy rate and average precision of each target detection task is used as the performance of the pruned backbone network. After pruning is completed using all the pruning rates in the preset pruning rate set and the corresponding performance is obtained, the best performance is selected from them, and the corresponding pruning rate is used as the optimal pruning rate of the backbone network.
[0070] In this embodiment, the multiple object detection tasks are respectively cigarette detection, mobile phone detection, and face detection. Dynamic pruning is performed on the corresponding branch networks and the backbone network, so as to determine that the optimal pruning rate of the backbone network is 0.5, and the pruned backbone network is obtained based on this optimal pruning rate.
[0071] The multi-object detection method provided by the embodiment of the present invention compresses the pre-trained multi-object detection model by dynamically pruning the backbone network, thereby reducing the model volume.
[0072] Further, the pruned multi-object detection model is obtained by respectively performing dynamic pruning on multiple target branch networks in the pre-trained multi-object detection model, and includes:
[0073] S1. Select a branch pruning rate from the preset branch pruning rate set as the to-be-analyzed branch pruning rate, select a target branch network from the multiple target branch networks as the to-be-analyzed target branch network, and use the to-be-analyzed branch pruning rate to prune the to-be-analyzed target branch network to obtain the pruned target branch network.
[0074] In this step, assuming that the branch pruning rate set is [0.1, 0.15, 0.2, 0.25, …, 0.8, 0.85, 0.9, 0.95], then in the order of each branch pruning rate in the set, it is used as the to-be-analyzed pruning rate in turn; and one is determined from the multiple target branch networks as the to-be-analyzed target branch network, and the pruned target branch network is obtained by using the to-be-analyzed pruning rate to prune the to-be-analyzed target branch network.
[0075] S2. Perform pruning sensitivity analysis on the pruned target branch network to obtain the performance of the pruned branch network.
[0076] In this step, by evaluating the performance of the pruned target branch network, the influence of the to-be-analyzed pruning rate on the detection accuracy corresponding to the target branch network is determined, and this process is the pruning sensitivity analysis process.
[0077] In addition, in this embodiment, the above-mentioned performance of the pruned branch network refers to the accuracy (ACC) of the target being detected and the average precision (AP). In other embodiments of the present invention, it may also be other model performance indicators, such as precision (precise), recall (recall), etc. The present invention does not make a limitation on this.
[0078] S3. Repeat the above S1 to S2 until all branch pruning rates in the preset branch pruning rate set are exhausted, and obtain the performances of multiple pruned branch networks.
[0079] In this step, by selecting different branch pruning rates, the performance of the pruned branch network corresponding to each branch pruning rate is analyzed and obtained.
[0080] S4. Determine the optimal performance of the pruned branch network from the performances of the multiple pruned branch networks, and use the branch pruning rate corresponding to the optimal performance of the pruned branch network as the optimal pruning rate of the target branch network to be analyzed.
[0081] In this step, by comparing the performances of the pruned branch networks corresponding to each branch pruning rate, the best performance of the pruned branch network is selected, and then the corresponding branch pruning rate is used as the optimal pruning rate of the target branch network to be analyzed.
[0082] S5. Repeat S1 to S4 until all target branch networks are exhausted, obtain the optimal pruning rate corresponding to each target branch network, and perform pruning on the target branch network based on the corresponding optimal pruning rate respectively to obtain the pruned multi-target detection model.
[0083] In this step, repeat S1 to S4 to perform the above analysis on each target branch network, so as to determine the optimal pruning rate corresponding to each target branch network. After obtaining the optimal pruning rates of different target branch networks, perform pruning on each target branch network to obtain the pruned target branch network, thereby constituting the pruned multi-target detection model.
[0084] In this embodiment, faces, mobile phones, and cigarettes are selected as detection targets. After analyzing through S1 - S5 above, the optimal pruning rate of the face detection branch network is 0.7, the optimal pruning rate of the mobile phone detection branch network is 0.6, and the optimal pruning rate of the cigarette detection branch network is 0.5. According to the above optimal pruning rates, the face detection branch network, the mobile phone detection branch network, and the cigarette detection branch network are respectively pruned to obtain the corresponding pruned face detection branch network, the pruned mobile phone detection branch network, and the pruned cigarette detection branch network.
[0085] The multi-target detection method provided by the embodiment of the present invention can maximize the compression of each target branch network, and ensure that the detection accuracy corresponding to each target is almost not lost. The pruned multi-target detection model can be applied to scenarios where processor cost and performance need to be considered.
[0086] Further, the preset set of branch pruning rates is obtained as follows:
[0087] According to the preset value range of the branch pruning rate and the value step of the branch pruning rate, all branch pruning rates that meet the value step of the branch pruning rate are exhausted within the preset value range of the branch pruning rate, so as to obtain the set of branch pruning rates.
[0088] Specifically, assume that the value range of the branch pruning rate is set to [0.1, 0.95], and the step size of the branch pruning rate value is 0.05. Then, the set of pruning rates to be analyzed [0.1, 0.15, 0.2, 0.25, …, 0.8, 0.85, 0.9, 0.95] is obtained by exhaustive means.
[0089] The multi-object detection method provided by the embodiments of the present invention determines the most suitable branch pruning rate for each target branch network by setting a set of branch pruning rates and performing pruning sensitivity analysis, further ensuring that the model is compressed to the maximum extent with almost no loss of accuracy.
[0090] Further, the multiple pruned target branch networks include a pruned face detection branch network, a pruned cigarette detection branch network, and a pruned mobile phone detection branch network; the detection results of the multiple targets include face detection results, cigarette detection results, and mobile phone detection results.
[0091] Correspondingly, inputting the picture to be detected into the pruned multi-object detection model, the detection results of multiple targets are obtained respectively, including: the pruned backbone network extracts multi-scale feature maps from the picture to be detected; the pruned face detection branch network predicts face detection results based on the multi-scale feature maps; the pruned cigarette detection branch network predicts cigarette detection results based on the multi-scale feature maps; the pruned mobile phone detection branch network predicts mobile phone detection results based on the multi-scale feature maps.
[0092] In this embodiment, taking three targets of face, cigarette, and mobile phone as examples to introduce the specific implementation process of multi-object detection.
[0093] After the picture to be detected is input into the pruned multi-object detection model, the pruned backbone network extracts features from the picture to be detected, thereby obtaining multiple feature maps with different scales. According to the size relationship between the targets, the feature maps with different scales are respectively used as the inputs of each pruned target branch network. Specifically, in terms of the size relationship, the size of the face > the size of the mobile phone > the size of the cigarette, and the scale of the feature map obtained after the convolution calculation of each convolutional layer in the pruned backbone network for the picture to be detected is continuously decreasing. Therefore, in order to improve the detection performance of small targets, the feature map output by the earlier convolutional layer is used as the input of the pruned cigarette detection branch network, the feature map output by the slightly later convolutional layer is used as the input of the pruned mobile phone detection branch network, and the feature map output by the last convolutional layer is used as the input of the pruned face detection branch network. The pruned face detection branch network, the pruned mobile phone detection branch network, and the pruned cigarette detection branch network respectively predict face detection results, mobile phone detection results, and cigarette detection results based on the corresponding feature maps.
[0094] The multi-object detection method provided by the embodiment of the present invention can not only ensure the detection accuracy of large objects, but also ensure the detection accuracy of small objects. Moreover, due to the dynamic pruning of each object branch network, a multi-object detection model with a smaller volume can be obtained.
[0095] Further, the pruned backbone network includes a plurality of convolutional layers.
[0096] Correspondingly, the pruned backbone network extracts features from the image to be detected to obtain multi-scale feature maps, including: each convolutional layer extracts features from the image to be detected, thereby obtaining a plurality of feature maps with different scales; according to the scales of the face prediction boxes in the face detection results, the cigarette prediction boxes in the cigarette detection results, and the mobile phone prediction boxes in the mobile phone detection results, the face detection feature map, the cigarette detection feature map, and the mobile phone detection feature map are respectively determined from the feature maps with different scales; the pruned face detection branch network predicts the face detection results based on the face detection feature map; the pruned cigarette detection branch network predicts the cigarette detection results based on the cigarette detection feature map; the pruned mobile phone detection branch network predicts the mobile phone detection results based on the mobile phone detection feature map.
[0097] Schematically, as Figure 2 shown, the feature maps output by the last three convolutional layers in the pruned backbone network are respectively used as the cigarette detection feature map, the mobile phone detection feature map, and the face detection feature map, and then are respectively input into the corresponding pruned cigarette detection branch network, the pruned mobile phone detection branch network, and the pruned face detection branch network. The convolutional layers in the pruned cigarette detection branch network, the pruned mobile phone detection branch network, and the pruned face detection branch network respectively perform further feature extraction on the cigarette detection feature map, the mobile phone detection feature map, and the face detection feature map, and then detect the corresponding objects. For example, if there is a cigarette in the image to be detected, there will be a cigarette prediction box at the position of the cigarette in the image to be detected, and the target type "cigarette" and the probability that the target belongs to "cigarette" are marked around the cigarette prediction box; if there is no cigarette in the image to be detected, it will not be displayed.
[0098] The multi-object detection method provided by the embodiment of the present invention can not only ensure the detection accuracy of large objects, but also ensure the detection accuracy of small objects. Moreover, due to the dynamic pruning of each object branch network, a multi-object detection model with a smaller volume can be obtained.
[0099] Next, the multi-object detection device provided by the present invention will be described. The multi-object detection device described below can be correspondingly referred to the multi-object detection method described above.
[0100] Figure 3 The following is a schematic structural diagram of the multi-object detection device provided by an embodiment of the present invention. As Figure 3 shown, a multi-object detection device includes:
[0101] An image acquisition module 301, configured to acquire an image to be detected.
[0102] In this embodiment, the image to be detected is an in-cabin image in an autonomous driving application scenario. In other embodiments of the present invention, the image to be detected may also be an internal image of a refrigerator in a smart refrigerator application scenario, or may be an image formed by a road monitoring video in a city monitoring application scenario. The present invention does not make any limitations in this regard.
[0103] In addition, the image to be detected can be obtained by on-site shooting, or can be decomposed from a video stream, or can be a test image obtained from a database in various multi-object detection application scenarios. The present invention does not make any limitations on the acquisition method of the image to be detected.
[0104] A detection module 302, configured to input the image to be detected into the pruned multi-object detection model to obtain detection results of multiple objects respectively.
[0105] Among them, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model; the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network; correspondingly, the pruned multi-object detection model is obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively.
[0106] In this module, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model, and the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network. Each object branch network is used to detect different objects. For example, when the image to be detected is an in-cabin image in an autonomous driving application scenario, the object branch network can be a mobile phone detection branch network, a key detection branch network, a tissue paper detection branch network, a glasses detection branch network, a face detection branch network, etc., for detecting different objects.
[0107] By performing targeted dynamic pruning on different object branch networks, specifically, based on multiple pruning rates, pruning sensitivity analysis is performed on each object branch network to determine the optimal pruning rate corresponding to each object branch network, and pruning is completed according to the optimal pruning rate corresponding to each object branch network, so as to ensure that the performance loss of object detection is small or even non-existent, and each object branch network can be compressed to the maximum extent.
[0108] After dynamically pruning different target branch networks, the pruned target branch networks are obtained. Therefore, the multi-target detection model after pruning is composed of a backbone network and the pruned target branch networks. Among them, the backbone network extracts features from the image to be detected to obtain multi-scale feature maps, and inputs them into different pruned target branch networks. Each pruned target branch network detects the corresponding target based on the multi-scale feature maps. Taking the image to be detected as an in-cabin image in an autonomous driving application scenario as an example, the detection results corresponding to each pruned target branch network are to locate each type of target on the image to be detected through target boxes of different colors, and display the type name of the target and the probability of belonging to this type.
[0109] In addition, the above-mentioned pre-trained multi-task target detection model is trained based on the labeled bounding boxes, the corresponding target type labels, and the training data set.
[0110] The multi-target detection device provided by the embodiments of the present invention dynamically prunes multiple target branch networks in the pre-trained multi-target detection model respectively to obtain the multi-target detection model after pruning, so that each target branch network can be compressed to the maximum extent, and it is ensured that the detection accuracy corresponding to each target is hardly lost. The multi-target detection model after pruning can be applied to scenarios that need to consider issues such as processor cost and performance. In addition, since the backbone network and the pruned target branch networks are independent of each other, and the multi-scale feature maps output by the backbone network are input into different pruned target branch networks, the differences between different targets at different scales can be noticed, and the detection accuracy is relatively high whether it is for large targets or small targets.
[0111] Figure 4 It is a schematic physical structure diagram of an electronic device provided by the embodiments of the present invention, as Figure 4As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a multi-object detection method, and the multi-object detection method includes: obtaining a picture to be detected; inputting the picture to be detected into the pruned multi-object detection model to obtain the detection results of multiple objects respectively; wherein, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model; the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network; correspondingly, the pruned multi-object detection model is obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively.
[0112] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0113] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a multi-object detection method provided by each of the above methods. The multi-object detection method includes: obtaining a picture to be detected; inputting the picture to be detected into a pruned multi-object detection model to respectively obtain detection results of multiple objects; wherein, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model; the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network; correspondingly, the pruned multi-object detection model is obtained by respectively dynamically pruning multiple object branch networks in the pre-trained multi-object detection model.
[0114] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the multi-object detection method provided by each of the above methods. The multi-object detection method includes: obtaining a picture to be detected; inputting the picture to be detected into a pruned multi-object detection model to respectively obtain detection results of multiple objects; wherein, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model; the pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network; correspondingly, the pruned multi-object detection model is obtained by respectively dynamically pruning multiple object branch networks in the pre-trained multi-object detection model.
[0115] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-object detection method, characterized in that, Including: Obtain the picture to be detected; Input the picture to be detected into the pruned multi-object detection model to obtain the detection results of multiple objects respectively; Among them, the pruned multi-object detection model is obtained by dynamically pruning the pre-trained multi-object detection model; The pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network; Correspondingly, the pruned multi-object detection model is obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively.
2. The multi-object detection method according to claim 1, characterized in that, The pruned multi-object detection model includes a pruned backbone network and multiple pruned object branch networks connected to the pruned backbone network; Correspondingly, the multiple pruned object branch networks are obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively; The pruned backbone network is obtained by dynamically pruning the backbone network in the pre-trained multi-object detection model.
3. The multi-object detection method according to claim 1, characterized in that, The pruned multi-object detection model is obtained by dynamically pruning multiple object branch networks in the pre-trained multi-object detection model respectively, including: S1. Select a branch pruning rate from the preset branch pruning rate set as the branch pruning rate to be analyzed, select an object branch network from the multiple object branch networks as the object branch network to be analyzed, and use the branch pruning rate to be analyzed to prune the object branch network to be analyzed to obtain a pruned object branch network; S2. Conduct pruning sensitivity analysis on the pruned object branch network to obtain the performance of the pruned branch network; S3. Repeat S1 to S2 until all branch pruning rates in the preset branch pruning rate set are exhausted to obtain the performance of multiple pruned branch networks; S4. Determine the optimal performance of the pruned branch network from the performance of the multiple pruned branch networks, and use the branch pruning rate corresponding to the optimal performance of the pruned branch network as the optimal pruning rate of the object branch network to be analyzed; S5. Repeat S1 to S4 until all object branch networks are exhausted to obtain the optimal pruning rate corresponding to each object branch network, and prune the object branch network based on the corresponding optimal pruning rate respectively to obtain the pruned multi-object detection model.
4. The multi-object detection method according to claim 3, characterized in that, The preset branch pruning rate set is obtained as follows: According to the preset value range and value step of the branch pruning rate, all branch pruning rates that meet the value step of the branch pruning rate are enumerated within the preset value range of the branch pruning rate, so as to obtain the branch pruning rate set.
5. The multi-object detection method according to any one of claims 2-4, characterized in that, The multiple pruned object branch networks include a pruned face detection branch network, a pruned cigarette detection branch network, and a pruned mobile phone detection branch network; The detection results of the multiple objects include face detection results, cigarette detection results, and mobile phone detection results; Correspondingly, the step of inputting the picture to be detected into the pruned multi-object detection model to obtain the detection results of multiple objects respectively includes: The pruned backbone network extracts features from the image to be detected to obtain multi-scale feature maps; The pruned face detection branch network predicts face detection results based on the multi-scale feature maps; The pruned cigarette detection branch network predicts cigarette detection results based on the multi-scale feature maps; The pruned mobile phone detection branch network predicts mobile phone detection results based on the multi-scale feature maps.
6. The multi-object detection method according to claim 5, characterized in that, The pruned backbone network includes multiple convolutional layers; Correspondingly, the pruned backbone network extracts features from the image to be detected to obtain multi-scale feature maps, including: Each convolutional layer extracts features from the image to be detected, thereby obtaining multiple feature maps with different scales; According to the scales of the face prediction boxes in the face detection results, the cigarette prediction boxes in the cigarette detection results, and the mobile phone prediction boxes in the mobile phone detection results, face detection feature maps, cigarette detection feature maps, and mobile phone detection feature maps are respectively determined from the feature maps with different scales; The pruned face detection branch network predicts face detection results based on the face detection feature maps; The pruned cigarette detection branch network predicts cigarette detection results based on the cigarette detection feature maps; The pruned mobile phone detection branch network predicts mobile phone detection results based on the mobile phone detection feature maps.
7. A multi-object detection device, characterized in that, Including: An image acquisition module for acquiring an image to be detected; A detection module for inputting the image to be detected into the pruned multi-object detection model to respectively obtain detection results of multiple objects; Wherein, the pruned multi-object detection model is obtained by dynamically pruning a pre-trained multi-object detection model; The pre-trained multi-object detection model includes a backbone network and multiple object branch networks connected to the backbone network; Correspondingly, the pruned multi-object detection model is obtained by respectively dynamically pruning multiple object branch networks in the pre-trained multi-object detection model.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the multi-object detection method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by the processor, it implements the multi-object detection method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, wherein, When the computer program is executed by the processor, it implements the multi-object detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for pruning neural network
CN111382839A
Multi-target detection method and device and storage device
CN112001247A