Model compression and environment perception method and device, equipment, medium and product
By applying pruning strategies in the three-dimensional object detection model and compressing the processing task processing model, the problem of large calculation volume and low inference efficiency is solved, and real-time perception capabilities in scenarios such as autonomous driving are achieved.
Patent Information
- Application Number
- CN202510094421.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing three-dimensional object detection model has complex structure, large parameters and high calculations, making it difficult to meet the real-time reasoning needs of scenarios such as autonomous driving.
By obtaining the task scenario of the target environment-aware task, determining the execution performance conditions, and compressing the task processing model based on the pruning strategy. The pruning strategy includes determining the target pruning strategy, determining the channel to be pruned and zeroing its channel parameters to obtain the target model.
When the execution performance conditions are met, the calculation amount of the task processing model is greatly reduced, the inference efficiency is improved, the computing resource consumption is reduced, and it is applied in real-time inference scenarios.
Smart Images

Figure CN120012858A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection technology, and specifically to a model compression and environment perception method and device, equipment, medium, and product. Background Art
[0002] Three-dimensional object detection technology plays a key role in the fields of autonomous driving, intelligent transportation systems, etc. Taking the field of autonomous driving as an example, with the rapid development of autonomous driving technology, the vehicle's perception of the target environment has become a key factor in achieving safe driving. As an indispensable part of autonomous driving, three-dimensional object detection greatly affects the vehicle's recognition, detection and obstacle avoidance decisions for surrounding pedestrians, obstacles and roads. However, the task processing model for three-dimensional object detection is often complex in structure, with a large number of parameters and high computational complexity, making it difficult to meet the vehicle's real-time reasoning scenario requirements. Summary of the invention
[0003] The present application provides a model compression and environment perception method and device, equipment, medium, and product.
[0004] In order to achieve the above purpose, the technical solution adopted in this application is as follows:
[0005] A model compression method, the method comprising: obtaining a task scenario of a target environment perception task and a task processing model to be compressed; the task processing model comprises at least one network layer, and the network layer comprises at least one channel; based on the task scenario, determining the execution performance condition of the target environment perception task; based on the execution performance condition, compressing the task processing model through a pruning strategy to obtain a target model; the compression processing is used to set channel parameters corresponding to the channels to be pruned in the task processing model to zero, the channels to be pruned are determined based on the pruning strategy, and the target model is used to execute the target environment perception task and meet the execution performance condition.
[0006] According to the above technical means, first, the task scenario of the target environment perception task is used to determine the execution performance conditions so that the execution performance conditions match the task scenario. Then, the task processing model is compressed through the pruning strategy to obtain the target model. In this way, while meeting the execution performance conditions, the amount of calculation of the task processing model is greatly reduced, the reasoning efficiency of the task processing model is improved, and the consumption of computing resources of the task processing model is reduced, so that the target model can be applied to task scenarios that require real-time reasoning. In addition, by setting the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the compression processing method for the task processing model can reduce the damage to the model structure of the task processing model, thereby retaining the interactions and dependencies between the channels in the task processing model, and reducing the loss of computational accuracy of the target model.
[0007] Furthermore, based on the execution performance conditions, the task processing model is compressed through a pruning strategy to obtain a target model, including: determining a target pruning strategy based on the execution performance conditions; determining channels to be pruned in the task processing model based on the target pruning strategy; and setting channel parameters corresponding to the channels to be pruned in the task processing model to zero to obtain a target model.
[0008] According to the above technical means, the target pruning strategy can be determined by executing the performance conditions, so that the channel to be pruned can be determined based on the target pruning strategy to obtain the target model, so that the obtained target model meets the execution performance conditions, thereby applying the target model to the target environment perception task.
[0009] Further, the execution performance conditions include target accuracy conditions, target speed conditions and / or target computing volume conditions, and the target pruning strategy includes a target pruning rate; based on the execution performance conditions, the target pruning strategy is determined, including at least one of the following: based on the target accuracy condition, the inference accuracy of the task processing model is determined; based on the inference accuracy, the target pruning rate is determined; the inference accuracy includes the detection accuracy corresponding to at least one detection head in the task processing model; based on the target speed condition, the inference speed of the task processing model is determined; based on the inference speed, the target pruning rate is determined; the inference speed includes the detection speed corresponding to at least one detection head in the task processing model; based on the target computing volume condition, the inference computing volume of the task processing model is determined; based on the inference computing volume, the target pruning rate is determined; the inference speed includes the detection computing volume corresponding to at least one detection head in the task processing model.
[0010] According to the above technical means, a solution is provided for determining the target pruning strategy through execution performance conditions, so that the target pruning rate that meets the execution performance conditions can be determined, and the task processing model to be compressed is compressed through the target pruning rate, so that the operation of the generated compressed task processing model meets the execution performance conditions, so that the compressed task processing model can be applied to the actual target environment perception task.
[0011] Furthermore, based on the execution performance conditions, the task processing model is compressed through a pruning strategy to obtain a target model, including: determining at least one target pruning strategy; for each target pruning strategy, determining the channels to be pruned in the task processing model based on the target pruning strategy, and setting the channel parameters corresponding to the channels to be pruned in the task processing model to zero to obtain a candidate model; based on the execution performance conditions, determining the target model from each candidate model.
[0012] According to the above-mentioned technical means, at least one candidate model is generated through at least one target pruning strategy, and the performance of at least one candidate model is evaluated. In this way, through the performance evaluation indicators of each candidate model, the impact of compression processing on the task processing model can be comprehensively analyzed, so that the model that meets the execution performance conditions is determined as the target model. This method can efficiently screen out target models suitable for task scenarios in target environment perception tasks.
[0013] Further, based on the execution performance condition, a target model is determined from each candidate model, including: determining a performance evaluation index for each candidate model respectively; and selecting a target model whose performance evaluation index meets the execution performance condition from each candidate model.
[0014] According to the above-mentioned technical means, the candidate model corresponding to the performance evaluation indicator that meets the execution performance conditions is determined as the target model, so that the model that is most suitable for executing the target environment perception task can be selected, and the target model can be applied to the target environment perception task, which can reduce the loss of accuracy of the perception results of the target environment, and thus facilitate the subsequent determination of an execution strategy that is more suitable for the target environment based on the perception results of the target environment after the loss of accuracy is reduced.
[0015] Furthermore, the task processing model includes a multi-task model, which includes a feature extraction module shared by multiple subtasks and a detection head module corresponding to each subtask, and the target environment perception task includes at least one subtask; based on the target pruning strategy, the channels to be pruned in the task processing model are determined, including: determining the weight value of each subtask corresponding to the multi-task model; based on the weight value of each subtask, determining the pruning rate of the detection head module corresponding to each subtask; based on the pruning rate of each detection head module, determining the channels to be pruned in the network layer of the detection head module.
[0016] According to the above-mentioned technical means, by applying the weights of different subtasks in different task scenarios from the task processing model, different pruning rates can be determined based on the weights of different subtasks, so that the pruning rate corresponding to the subtask with a higher weight is lower, and more feature channels can be retained for the subtask with a higher weight, so that the output result of the subtask with a higher weight has a higher accuracy, while the output result of the subtask with a lower weight has a lower accuracy, so that the task processing results of the task processing model are more suitable for the corresponding task scenarios.
[0017] Furthermore, based on the target pruning strategy, channels to be pruned in the task processing model are determined, including: determining the contribution rate of each network layer in the task processing model; determining the pruning rate of each network layer based on the contribution rate of each network layer; and determining channels to be pruned in each network layer based on the pruning rate of each network layer.
[0018] According to the above-mentioned technical means, by using the contribution rates of different network layers in different task scenarios from the task processing model, different pruning rates can be determined based on the contribution rates of different network layers, so that the pruning rate corresponding to the network layer with a higher contribution rate is lower, and more feature channels can be reserved for the network layer with a higher contribution rate, so that the output result of the network layer with a higher contribution rate has a higher accuracy to maintain the overall performance of the task processing model, while the output result of the network layer with a lower contribution rate has a lower accuracy, thereby effectively improving the performance degradation problem caused by the global unified pruning rate.
[0019] Furthermore, based on the target pruning strategy, channels to be pruned in the task processing model are determined, including: obtaining a contribution value of each channel of each network layer in the task processing model; and determining channels to be pruned in each network layer based on a global pruning rate and the contribution value of each channel.
[0020] According to the above technical means, by analyzing the contribution values of different channels in different task scenarios from the task processing model, channels with lower contribution values can be determined as channels to be pruned based on the global pruning rate. In this way, the task processing model can be compressed from a global perspective. While maintaining the high calculation accuracy of the task processing model, the amount of model parameters and the amount of calculation of the task processing model can be greatly reduced, thereby improving the calculation speed of the task processing model. In this way, the task processing model can be applied in application scenarios with higher real-time requirements.
[0021] Furthermore, based on the execution performance condition, the task processing model is compressed through a pruning strategy to obtain a target model, including: based on the initial pruning rate of each network layer in the task processing model, determining the channels to be pruned in each network layer, setting the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and obtaining a candidate model; determining the performance evaluation index of the candidate model; when the performance evaluation index of the candidate model meets the execution performance condition, compressing the candidate model at least once through a pruning strategy until the performance evaluation index of the candidate model after this compression meets the target stop condition, and determining the candidate model obtained by the last compression process as the target model; the target stop condition includes at least one of the following: the performance evaluation index of the candidate model after this compression does not meet the execution performance condition; the change state of the performance evaluation index of the candidate model after this compression compared with the performance evaluation index of the candidate model obtained by the last compression process does not meet the target change condition.
[0022] According to the above technical means, the pruning rate of this compression process can be determined by the initial pruning rate and the pruning rate adjustment step size, so that the candidate model obtained by the previous compression process is compressed based on the pruning rate to obtain the candidate model after this compression. This gradual compression method by gradually increasing the pruning rate can solve the problem of a sudden drop in the performance of the target model caused by a one-time large-scale compression.
[0023] Furthermore, the performance evaluation indicators of the candidate model include accuracy indicators, and the target change condition includes an accuracy change threshold. When the difference between the accuracy indicator of the candidate model after this compression and the accuracy indicator obtained by the previous compression process is greater than the accuracy change threshold, it is determined that the change state of the performance evaluation indicator of the candidate model after this compression compared to the performance evaluation indicator of the candidate model obtained by the previous compression process does not meet the target change condition.
[0024] According to the above-mentioned technical means, through the relationship between the difference between the accuracy index of the candidate model after this compression and the accuracy index obtained by the last compression process, and the accuracy change threshold, it can be determined whether to end the compression of the candidate model. If it is determined to be ended, the candidate model obtained by the last compression will be determined as the target model. The performance evaluation index of the target model obtained in this way is relatively better.
[0025] Furthermore, obtaining the task processing model to be compressed includes: performing perspective conversion on image data collected by at least one collection device to determine feature data corresponding to the image data under the bird's-eye view perspective; inputting the feature data into the initial task processing model, training the initial task processing model, and obtaining the task processing model to be compressed.
[0026] According to the above technical means, first, by processing the image data collected by at least one acquisition device, feature data corresponding to the image data at a bird's-eye view perspective is obtained, wherein the data image is converted to a bird's-eye view perspective, and the local information of the target environment at different perspectives can be integrated into a unified view, thereby eliminating the influence of perspective differences. Then, based on the feature data unified in the bird's-eye view, it is input into the initial task processing model, so that the initial task processing model can perform model training without the need to integrate the image data information, thereby improving the convergence speed and training efficiency of the task processing model.
[0027] A method for environmental perception, the method comprising: acquiring image data of a target environment; using a target model to perform a target environmental perception task based on the image data to obtain an environmental perception result; the target model is determined based on the execution performance condition of the target environmental perception task and after compressing a task processing model through a pruning strategy; the compression processing is used to set channel parameters corresponding to channels to be pruned in the task processing model to zero, and the channels to be pruned are determined based on the pruning strategy; the execution performance condition is determined based on the task scenario of the target environmental perception task.
[0028] According to the above-mentioned technical means, first, the execution performance conditions are determined through the task scenario of the target environmental perception task, and then the task processing model is compressed through the pruning strategy to obtain the target model, so that the target model is more closely matched with the task scenario of the target environmental perception task. In this way, applying the target model to the same task scenario of the target environmental perception task can not only reduce the waiting time for obtaining the environmental perception results, but also improve the calculation accuracy of the environmental perception results.
[0029] A model compression device, comprising:
[0030] A first acquisition module is used to acquire a task scenario of a target environment perception task and a task processing model to be compressed; the task processing model includes at least one network layer, and the network layer includes at least one channel;
[0031] A determination module, used to determine the execution performance conditions of the target environment perception task based on the task scenario;
[0032] The first obtaining module is used to compress the task processing model through a pruning strategy based on the execution performance conditions to obtain a target model; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the channels to be pruned are determined based on the pruning strategy, and the target model is used to execute the target environment perception task and meet the execution performance conditions.
[0033] An environment sensing device, comprising:
[0034] A second acquisition module is used to acquire image data of the target environment;
[0035] The second obtaining module is used to use the target model to perform the target environment perception task based on the image data to obtain the environment perception result; the target model is determined based on the execution performance conditions of the target environment perception task, and the task processing model is compressed through the pruning strategy; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and the channels to be pruned are determined based on the pruning strategy; the execution performance conditions are determined based on the task scenario of the target environment perception task.
[0036] A computer device comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, part or all of the steps in the above method are implemented.
[0037] A computer-readable storage medium stores a computer program, which implements part or all of the steps in the above method when executed by a processor.
[0038] A computer program product comprises a computer program or instructions, which, when executed by a processor, perform some or all of the steps in the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of the implementation process of a model compression method proposed in this application;
[0040] Figure 2 A schematic diagram of the implementation process of an environment perception method proposed in this application;
[0041] Figure 3A A schematic diagram of the implementation process of a multi-task model compression task based on a bird's-eye view proposed in this application;
[0042] Figure 3B Schematic diagram of the implementation of a multi-task model compression task based on a bird's-eye view proposed in this application Figure 1 ;
[0043] Figure 3C Schematic diagram of the implementation of a multi-task model compression task based on a bird's-eye view proposed in this application Figure 2 ;
[0044] Figure 4 A schematic diagram of the structure of a model compression device proposed in this application;
[0045] Figure 5 A schematic diagram of the structure of an environment sensing device proposed in this application;
[0046] Figure 6 A hardware entity schematic diagram of a computer device proposed in this application. DETAILED DESCRIPTION
[0047] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.
[0048] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0049] The present application embodiment proposes a model compression method, such as Figure 1 As shown, the model compression method includes the following steps S101 to S103, wherein:
[0050] Step S101: Acquire a task scenario of a target environment perception task and a task processing model to be compressed; the task processing model includes at least one network layer, and the network layer includes at least one channel.
[0051] Here, the target environment perception task refers to the task of obtaining surrounding environment information through sensors and other devices, and processing and analyzing this information to identify and monitor surrounding dynamic and static obstacles, such as vehicles, pedestrians, buildings, etc. For example, in the field of autonomous driving, a comprehensive view of the traffic environment can be constructed through the target environment perception task, so that these obstacles can be identified and monitored, and the passable areas can be distinguished, providing a reliable basis for planning the autonomous driving path.
[0052] The task scenario refers to the sum of factors such as environment, conditions, interactions and behaviors involved when performing the target environment perception task.
[0053] The task processing model to be compressed refers to the model that needs to be compressed due to its large size and high computational complexity when performing the target environment perception task. Among them, the task processing model includes at least one network layer. The network layer is the basic building block of the task processing model. Each network layer performs a specific calculation or data processing task. By stacking multiple layers, a complex task processing model can be constructed to process various tasks, such as image recognition, speech recognition, natural language processing, etc. The network layer includes at least one channel. The channel is a key concept that describes the depth of an image or feature map. It determines the number of data components contained in each pixel or feature point and affects the performance and computational complexity of the network layer. By reasonably setting and adjusting the number of channels, an efficient and accurate task processing model can be constructed.
[0054] In some embodiments, when the task processing model is implemented as a convolutional neural network (CNN) model, the network layer refers to a convolutional layer, and the channel refers to a convolutional channel.
[0055] In some implementations, the target environment perception task may be issued by a user through the human-computer interaction interface of the system, or may be identified and perceived by the system based on the current operating state and the environment in which it is located.
[0056] In some implementations, the task scenario of the target environment perception task may be acquired through task information of the target environment perception task.
[0057] In some implementations, the task processing model to be processed may be obtained by training an initial task processing model with image data of a target environment.
[0058] Step S102: Determine the execution performance conditions of the target environment perception task based on the task scenario.
[0059] Here, the execution performance conditions refer to the performance standards and requirements that the task processing model needs to meet when performing the target environment perception task.
[0060] In some embodiments, the focus objects in different mission scenarios are different. For example, for the vehicle's autonomous driving scenario, the focus is on the vehicle's driving road. Therefore, the standards for execution performance conditions on the driving road and the execution performance conditions on the non-driving road are different.
[0061] Step S103: Based on the execution performance conditions, the task processing model is compressed through the pruning strategy to obtain a target model; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the channels to be pruned are determined based on the pruning strategy, and the target model is used to execute the target environment perception task and meet the execution performance conditions.
[0062] Here, pruning strategy refers to removing some unnecessary or redundant elements from complex models or algorithms in order to simplify the model structure, reduce the amount of calculation and improve the generalization ability. Its main purpose is to prevent overfitting by reducing the complexity of the model while maintaining or improving the predictive performance of the model.
[0063] In some implementations, the channels to be pruned in the task processing model can be determined by the pruning strategy, so that the channel parameters of the channels to be pruned can be set to zero, and the channels to be pruned can be removed, so that in the subsequent model application process, the zeroed channels are not calculated. This method can prevent the structure of the task processing model from being destroyed and improve the calculation accuracy of the target model.
[0064] In some implementations, after the pruning strategy compresses the task processing model, a target model is obtained;
[0065] In other implementations, after the pruning strategy compresses the task processing model, a first candidate model is obtained, so that the first candidate model can be fine-tuned and trained, and a performance evaluation index is performed on the fine-tuned first candidate model, and a target model is obtained based on the performance evaluation index. Among them, the fine-tuning training can include optimizing training parameters, data enhancement, precision recovery and other means.
[0066] In this embodiment, first, the task scenario of the target environment perception task is used to determine the execution performance condition so that the execution performance condition matches the task scenario. Then, the task processing model is compressed through the pruning strategy to obtain the target model. In this way, when the execution performance condition is met, the calculation amount of the task processing model is greatly reduced, the reasoning efficiency of the task processing model is improved, and the consumption of computing resources of the task processing model is reduced, so that the target model can be applied to task scenarios that require real-time reasoning. In addition, by setting the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the compression processing method of the task processing model can reduce the damage to the model structure of the task processing model, thereby retaining the interaction and dependency between the channels in the task processing model, and reducing the loss of calculation accuracy of the target model.
[0067] In some embodiments, based on the execution performance condition, the task processing model is compressed by a pruning strategy to obtain a target model, which may include the following steps S1031 and S1032, wherein:
[0068] Step S1031: determining a target pruning strategy based on the execution performance condition; determining a channel to be pruned in the task processing model based on the target pruning strategy;
[0069] Here, the target pruning strategy refers to a pruning strategy that will satisfy the execution performance condition.
[0070] In some implementations, the execution performance condition includes a target accuracy condition, and an accuracy threshold indicated in the target accuracy condition may be obtained, thereby determining a channel to be pruned in the task processing model based on the accuracy threshold.
[0071] In some implementations, the execution performance condition includes a target speed condition, and a speed threshold indicated in the target speed condition may be obtained, thereby determining a channel to be pruned in the task processing model based on the speed threshold.
[0072] In some implementations, the execution performance condition includes a target computational load condition, and a computational load threshold indicated in the target computational load condition may be obtained, thereby determining a channel to be pruned in the task processing model based on the computational load threshold.
[0073] Step S1032: Set the channel parameters corresponding to the channels to be pruned in the task processing model to zero to obtain the target model.
[0074] In this embodiment, the target pruning strategy can be determined by executing the performance conditions, so that the channel to be pruned can be determined based on the target pruning strategy to obtain the target model, so that the obtained target model meets the execution performance conditions, and the target model is applied to the target environment perception task.
[0075] In some embodiments, the execution performance condition includes a target accuracy condition, a target speed condition and / or a target computation amount condition, and the target pruning strategy includes a target pruning rate;
[0076] Determining the target pruning strategy based on the execution performance condition may include at least one of the following steps S10311 to S10313, wherein:
[0077] Step S10311: determining the inference accuracy of the task processing model based on the target accuracy condition; determining the target pruning rate based on the inference accuracy; the inference accuracy includes the detection accuracy corresponding to at least one detection head in the task processing model;
[0078] Here, the target accuracy condition may represent the accuracy change condition of the task processing model, or may represent the accuracy condition of the compressed task processing model.
[0079] In some implementations, the inference accuracy may be the detection accuracy of the task processing model and / or the detection accuracy corresponding to the detection head of each task in the task processing model.
[0080] In some embodiments, when the target accuracy condition represents the accuracy change condition of the task processing model, the accuracy change condition may include an accuracy change threshold, wherein the smaller the accuracy change threshold, the lower the corresponding target pruning rate should be.
[0081] In some embodiments, when the target accuracy condition represents the accuracy condition of the compressed task processing model, the accuracy condition may include an accuracy threshold, wherein the smaller the accuracy threshold, the higher the corresponding target pruning rate should be.
[0082] Step S10312: determining the inference speed of the task processing model based on the target speed condition; determining the target pruning rate based on the inference speed; the inference speed includes the detection speed corresponding to at least one detection head in the task processing model;
[0083] Here, the target speed condition may represent a speed change condition of the task processing model, or may represent a speed condition of the compressed task processing model.
[0084] In some implementations, the inference speed may be the execution speed of the task processing model and / or the execution speed corresponding to the detection head of each task in the task processing model.
[0085] In some embodiments, when the target speed condition represents a speed change condition of the task processing model, the speed change condition may include a speed change threshold, wherein the smaller the speed change threshold, the lower the corresponding target pruning rate should be.
[0086] In some implementations, when the target speed condition represents the speed condition of the compressed task processing model, the speed condition may include a speed threshold, wherein the smaller the speed threshold, the lower the corresponding target pruning rate should be.
[0087] Step S10313: Based on the target computational load condition, determine the inference computational load of the task processing model; based on the inference computational load, determine the target pruning rate; the inference speed includes the detection computational load corresponding to at least one detection head in the task processing model.
[0088] Here, the target computational load condition may represent a computational load variation condition of the task processing model, or may represent a computational load condition of the compressed task processing model.
[0089] In some implementations, the inference computation load may be the computation load of the task processing model and / or the computation load corresponding to the detection head of each task in the task processing model.
[0090] In some embodiments, when the target computing volume condition represents the computing volume change condition of the task processing model, the computing volume change condition may include a computing volume change threshold, wherein the smaller the computing volume change threshold, the lower the corresponding target pruning rate should be.
[0091] In some embodiments, when the target computing volume condition represents the computing volume condition of the compressed task processing model, the computing volume condition may include a computing volume threshold, wherein the smaller the computing volume threshold, the higher the corresponding target pruning rate should be.
[0092] In this embodiment, a solution is provided for determining a target pruning strategy by executing performance conditions, so that a target pruning rate that meets the execution performance conditions can be determined, and the task processing model to be compressed is compressed by the target pruning rate, so that the operation of the generated compressed task processing model meets the execution performance conditions, so that the compressed task processing model can be applied to the actual target environment perception task.
[0093] In some embodiments, based on the execution performance condition, the task processing model is compressed by a pruning strategy to obtain a target model, which may include the following steps S1033 to S1035, wherein:
[0094] Step S1033: Determine at least one target pruning strategy;
[0095] In some implementations, the target pruning strategy may be at least one of a task-based pruning strategy, a network layer-based pruning strategy, and a global-based pruning strategy.
[0096] In some implementations, the target pruning strategy may be a combination of a global pruning strategy and a local pruning strategy, wherein the local pruning strategy includes at least one of a task-based pruning strategy and a network layer-based pruning strategy.
[0097] In some embodiments, the target pruning strategy may be a pruning rate of each network layer in a preset task processing model, and at least one target pruning strategy may be at least one pruning rate of each network layer in the preset task processing model, wherein each pruning rate corresponds to a target pruning strategy.
[0098] Step S1034: for each target pruning strategy, based on the target pruning strategy, determine the channel to be pruned in the task processing model, and set the channel parameters corresponding to the channel to be pruned in the task processing model to zero, so as to obtain a candidate model;
[0099] In some embodiments, when the target pruning strategy may be a pruning rate of a network layer in a preset task processing model, the step is implemented as follows: obtaining multiple pruning rates of each network layer in the task processing model determined by the pruning strategy; for each pruning rate of each network layer in the task processing model, based on each pruning rate, determining a channel to be pruned in each network layer, setting a channel parameter corresponding to the channel to be pruned in the task processing model to zero, and obtaining a candidate model;
[0100] In some embodiments, when the target pruning strategy can be a combination of a global pruning strategy and a local pruning strategy, in which the local pruning strategy is illustrated as a task-based pruning strategy, the step is implemented as follows: obtaining the contribution value of each channel of each network layer in the task processing model; based on the global pruning rate and the contribution value of each channel, determining the channels to be pruned in each network layer, and setting the channel parameters corresponding to the channels to be pruned in the task processing model to zero, to obtain a second candidate model; determining the weight values of each subtask corresponding to the second candidate model; based on the weight value of each subtask, determining the pruning rate of the detection head module corresponding to each subtask; based on the pruning rate of each detection head module, determining the channels to be pruned in the network layer of the detection head module and setting the channel parameters corresponding to the channels to be pruned in the task processing model to zero, to obtain a candidate model.
[0101] Step S1035: Determine a target model from among the candidate models based on the execution performance condition.
[0102] In some implementations, a performance evaluation is performed on each candidate model obtained in the above step S1034 to obtain a performance evaluation index for each candidate model, and a target model is determined from each candidate model based on the performance evaluation index and performance execution condition of each candidate model.
[0103] In this embodiment, at least one candidate model is generated through at least one target pruning strategy, and the performance of at least one candidate model is evaluated. In this way, the impact of compression processing on the task processing model can be comprehensively analyzed through the performance evaluation indicators of each candidate model, so that the model that meets the execution performance conditions is determined as the target model. This method can efficiently screen out target models suitable for task scenarios in target environment perception tasks.
[0104] In some embodiments, the above step S1035, determining the target model from the candidate models based on the execution performance condition, may include the following steps S10351 and S10352, wherein:
[0105] Step S10351: Determine the performance evaluation index of each candidate model respectively;
[0106] In some implementations, at least one performance evaluation indicator among the inference accuracy, inference speed, and inference computation amount of each candidate model may be determined.
[0107] In some embodiments, when determining the inference accuracy of each candidate model, the first inference accuracy of the task processing model to be compressed can be obtained, and the second inference accuracy of each candidate model can be obtained, wherein the inference accuracy can be determined from at least one of the dimensions of the ratio of the number of samples predicted correctly by the model to the total number of samples, and the difference between the model prediction result and the actual result.
[0108] In some embodiments, when determining the inference speed of each candidate model, a first inference speed of the task processing model to be compressed on different hardware platforms can be obtained, and a second inference speed of each candidate model on different hardware platforms can be obtained, wherein the inference speed can be determined from at least one dimension of the number of images or frames that the model can process per second, and the processing time required for the model to process a single image or a single image frame.
[0109] In some embodiments, when determining the amount of computation of each candidate model, the first inference computation of the task processing model to be compressed can be obtained, and the second inference computation of each candidate model can be obtained, wherein the inference computation can be determined from at least one dimension of the model's parameter quantity, computational complexity, and memory occupancy and storage requirements.
[0110] Step S10352: From each candidate model, select a target model whose performance evaluation index meets the execution performance condition.
[0111] In some embodiments, the execution performance condition may include at least one of a target accuracy condition, a target speed condition, and a target computing amount condition, wherein the execution performance condition may be a performance evaluation indicator change characterizing the task processing model before and after compression, or may be a performance evaluation indicator characterizing the compressed task processing model.
[0112] In some embodiments, when the execution performance condition characterizes the change in performance evaluation indicators before and after compression of the task processing model, the execution performance condition includes a target speed condition as an example for explanation. The change value of the reasoning speed can be determined through the first reasoning speed and the second reasoning speed. When the change value of the reasoning speed meets the target accuracy condition, the candidate model corresponding to the second reasoning speed is determined as the target model.
[0113] In some embodiments, when the execution performance condition represents the performance evaluation index of the compressed task processing model, the execution performance condition includes the target accuracy condition as an example for explanation. When the second reasoning speed satisfies the target accuracy condition, the candidate model corresponding to the second reasoning speed is determined as the target model.
[0114] It should be noted that in actual implementation, there may be multiple candidate models that meet the performance evaluation indicators. In this case, the performance evaluation indicators of each candidate model can be compared, and the optimal performance evaluation indicator can be selected from them. The candidate model corresponding to the optimal performance evaluation indicator can be determined as the target model.
[0115] In this embodiment, the candidate model corresponding to the performance evaluation indicator that meets the execution performance conditions is determined as the target model, so that the model that is most suitable for executing the target environment perception task can be selected, and the target model can be applied to the target environment perception task, which can reduce the loss of accuracy of the perception results of the target environment, and thus facilitate the subsequent determination of an execution strategy that is more suitable for the target environment based on the perception results of the target environment after the loss of accuracy is reduced.
[0116] In some embodiments, the task processing model includes a multi-task model, the multi-task model includes a feature extraction module shared by multiple subtasks and a detection head module corresponding to each subtask, and the target environment perception task includes at least one subtask;
[0117] Determining the channel to be pruned in the task processing model based on the target pruning strategy may include steps S110 to S112:
[0118] Step S110: Determine the weight value of each subtask corresponding to the multi-task model;
[0119] In some implementations, the reasoning performance index of each subtask in the multi-task model can be obtained first; then, based on the reasoning performance index of the detection head of each subtask in the target environment perception task, the weight value of each subtask is determined. Taking the autonomous driving scenario as an example, the vehicle detection and pedestrian detection tasks have higher accuracy requirements, and the corresponding weight values should be larger, while the executable area segmentation and road sign recognition have lower accuracy requirements, and the corresponding weight values should be smaller.
[0120] In some implementations, the task scenario of the target environment perception task may be analyzed to obtain the task importance of each subtask in the task scenario; and then based on the task importance of each subtask, a weight value of each subtask may be determined.
[0121] In some implementations, the pruning rate allocation may be based on a fixed task weight. During implementation, a fixed task weight value may be set for each task in the task processing model, corresponding to the pruning rate p t The calculation formula is the inverse of the task weight. The task weight of each task can be an empirical value set based on human experience. For example, in autonomous driving, the target detection task is more important than the semantic segmentation task, so w can be set. detection =0.7, w segmentation =0.3.
[0122] In some implementations, the pruning rate may be allocated by dynamically adjusting task weights according to task difficulty or task importance. During implementation, the task weights may be set based on task performance indicators or based on the ratio of task loss values.
[0123] The task weight is set based on the performance index of the task. When implemented, the task weight can be adjusted based on the performance of the task on the performance evaluation index. The calculation method of the task weight is shown in formula (1):
[0124]
[0125] Here, w t represents the task weight of task t, perf t represents the performance evaluation index of task t, where the performance evaluation index may include but is not limited to reasoning speed and / or reasoning accuracy, etc.
[0126] The task weight is set based on the ratio of the task loss value. In implementation, during the training of the task processing model, the task loss value of each task can be detected, and the task weight of the task can be determined according to the size of the task loss value of each task. The larger the task loss value, the higher the task weight. The calculation method of the task weight is shown in formula (2):
[0127]
[0128] Here, wt represents the task weight of task t, losst represents the task loss value of task t, and sum(loss tasks ) represents the sum of the task loss values of all tasks in the task processing model.
[0129] In some implementations, the pruning rate allocation can be based on gradient sensitivity, by measuring the contribution of each task to the parameter update of the task processing model (gradient sensitivity). The greater the gradient sensitivity of the task, the higher the task weight. t and the importance of each network layerI {t,j}Combined with each other, determine the task weight of each task, where I {t,j} represents the gradient sensitivity of task t to the jth layer of the network layer. The gradient sensitivity of the network layer can be determined based on the gradient norm or Fisher information. Then, based on the task weight w t , determine the importance of the j-th network layer; finally, based on the importance of the j-th network layer, determine the pruning rate of the network layer.
[0130] In some embodiments, pruning rate allocation can be performed based on task uncertainty. During implementation, first, the task uncertainty of each task is calculated, where the task uncertainty can be estimated by observing the volatility of the loss function or using a Bayesian method, or by using methods such as heteroscedastic uncertainty or model uncertainty. Then, based on the task uncertainty of each task, the task weight of the task is determined, and the task weight is inversely proportional to its task uncertainty, that is, a task with a lower task uncertainty has a higher task weight, while a task with a higher task uncertainty has a lower task weight. Finally, the pruning rate is set according to the task weight of each task.
[0131] In some implementations, pruning rate allocation can be performed based on a learning method. During implementation, a weight network is designed to predict the pruning rate. The architecture of the weight network can be designed according to actual conditions. For example, a multi-layer perceptron, a convolutional neural network, or a graph neural network can be used. The weight network of a multi-layer perceptron architecture is used for illustration. The calculation method of the pruning rate is shown in formula (3):
[0132] p t =MLP(w t ,x) (3).
[0133] Here, w t represents the task weight of task t, P t represents the pruning rate of task t, and x represents the training data, where the training data can be the data used to train the task processing model.
[0134] Step S111: based on the weight value of each subtask, respectively determine the pruning rate of the detection head module corresponding to each subtask;
[0135] Here, the detection head module is usually located at the last stage of the task processing model, receives feature maps from the feature extraction network (such as the backbone network Backbone and the neck network Neck), and performs detection and classification based on these feature maps. The detection head module includes multiple network layers.
[0136] In some embodiments, different weight values correspond to different pruning rates, wherein the larger the weight value, the smaller the pruning rate should be. In this way, for the detection head module corresponding to the subtask with a larger weight value, a smaller pruning rate can be set to retain more feature channels for the subtask; for the subtask with a smaller weight value, a larger pruning rate can be set to perform more channel pruning, which can reduce the computational burden of the task processing model.
[0137] Step S112: Based on the pruning rate of each detection head module, determine the channels to be pruned in the network layer of the detection head module.
[0138] In some embodiments, the contribution values of the channels in the network layer of the detection head module can be determined first; then, based on the contribution values of all the channels in the detection head module, the channels to be pruned in the detection head module are determined from low to high importance.
[0139] In this embodiment, by applying the weights of different subtasks in different task scenarios from the task processing model, different pruning rates can be determined based on the weights of different subtasks, so that the pruning rate corresponding to the subtask with a higher weight is lower, and more feature channels can be retained for the subtask with a higher weight, so that the output result of the subtask with a higher weight has a higher accuracy, while the output result of the task with a lower weight has a lower accuracy, so that the task processing result of the task processing model is more suitable for the task scenario corresponding to the task.
[0140] In some embodiments, determining the channel to be pruned in the task processing model based on the target pruning strategy may include the following steps S113 to S115, wherein:
[0141] Step S113: determining the contribution rate of each network layer in the task processing model;
[0142] In some embodiments, the contribution rate of each network layer can be evaluated based on information gain. During implementation, the output in each network layer can be regarded as a feature, and then a gradient-based method or an activation value-based method can be used to evaluate the importance of the output to the model output. In this way, the contribution rate of each network layer can be determined based on the importance of the output of each network layer.
[0143] In some implementations, each network layer can be removed or replaced to observe changes in the accuracy of the task processing model on a validation set or a test set, thereby determining the contribution rate of each network layer based on the precise changes in each network layer.
[0144] In some implementations, the weights of the network layers may be checked to determine the characteristic patterns captured by each network layer, thereby determining the contribution rate of each network layer in combination with the task scenario of the target environment perception task.
[0145] In some implementations, the task scenario of the target environment perception task can be analyzed to obtain the importance of the output of each network layer in the task scenario to the output in the task scenario; then, based on the importance of each network layer, the contribution rate of each network layer is determined.
[0146] Step S114: based on the contribution rate of each network layer, determining the pruning rate of each network layer respectively;
[0147] In some embodiments, different contribution rates correspond to different pruning rates, wherein the greater the contribution rate, the smaller the pruning rate should be, so that for convolutional layers with larger contributions, more channels can be retained by setting a smaller pruning rate to maintain the overall performance of the task processing model; for convolutional layers with smaller contributions, more channels can be pruned by setting a larger pruning rate, thereby effectively improving the performance degradation problem caused by the global unified pruning rate.
[0148] Step S115: Based on the pruning rate of each network layer, determine the channels to be pruned in each network layer.
[0149] In some implementations, the contribution value of the channels in each network layer may be determined first; then, based on the contribution values of all the channels in each network layer, the channels to be pruned in each network layer are determined from low to high importance.
[0150] In this embodiment, by calculating the contribution rates of different network layers in different task scenarios from the task processing model, different pruning rates can be determined based on the contribution rates of different network layers, so that the pruning rate corresponding to the network layer with a higher contribution rate is lower, and more feature channels can be reserved for the network layer with a higher contribution rate, so that the output result of the network layer with a higher contribution rate has a higher accuracy to maintain the overall performance of the task processing model, while the output result of the network layer with a lower contribution rate has a lower accuracy, thereby effectively improving the performance degradation problem caused by the globally unified pruning rate.
[0151] In some embodiments, based on the target pruning strategy, determining the channel to be pruned in the task processing model may include the following steps S116 and S117, wherein:
[0152] Step S116: Obtain the contribution value of each channel of each network layer in the task processing model;
[0153] In some embodiments, the contribution value of each channel in the network layer can be determined by the parameters in the two-dimensional batch normalization (BatchNorm2d) layer in the convolutional layer in the task processing model; wherein the parameters can be selected as scaling parameters, and the scaling parameters can reflect the activation strength of the channel.
[0154] In some implementations, the contribution value of each channel in the network layer can be determined by the weight amplitude in the convolutional layer in the task processing model; wherein, the larger the weight amplitude, the larger the contribution value of the channel.
[0155] In some implementations, the contribution value of each channel may be determined by calculating the norm of each channel; wherein the larger the norm, the larger the contribution value of the channel.
[0156] Step S117: Based on the global pruning rate and the contribution value of each channel, determine the channels to be pruned in each network layer.
[0157] Here, the global pruning rate refers to the pruning ratio set for the entire network or model during the neural network pruning process.
[0158] In some implementations, the global pruning rate may be preset.
[0159] In some implementations, the global pruning rate may be obtained by weighting the local pruning rates determined in step S114 and / or step S111.
[0160] In some implementations, the global pruning rate may be first determined based on the execution performance conditions of the task scenarios in the multi-task model.
[0161] In some implementations, the contribution value of each channel may be sorted, and the channel to be pruned may be determined from low to high contribution value.
[0162] In this embodiment, by applying different contribution values of different channels in different task scenarios from the task processing model, channels with lower contribution values can be determined as channels to be pruned based on the global pruning rate. In this way, the task processing model can be compressed from a global perspective. While maintaining a high calculation accuracy of the task processing model, the amount of model parameters and the amount of calculation of the task processing model can be greatly reduced, thereby improving the calculation speed of the task processing model. In this way, the task processing model can be applied in application scenarios with higher real-time requirements.
[0163] In some embodiments, based on the execution performance condition, the task processing model is compressed by a pruning strategy to obtain a target model, which may include the following steps S1036 to S1038, wherein:
[0164] Step S1036: based on the initial pruning rate of each network layer in the task processing model, determine the channels to be pruned in each network layer, set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and obtain a candidate model;
[0165] Here, the initial pruning rate refers to the initial value of the pruning rate of each network layer, wherein the initial pruning rate may be preset, for example, the initial pruning rate may be 0.
[0166] In some implementations, the initial pruning rate for each network layer may be the same.
[0167] In some implementations, the initial pruning rate of each network layer may be different.
[0168] Step S1037: Determine the performance evaluation index of the candidate model;
[0169] Step S1038: When the performance evaluation index of the candidate model meets the execution performance condition, the candidate model is compressed at least once by using the pruning strategy until the performance evaluation index of the candidate model after the current compression meets the target stop condition, and the candidate model obtained by the previous compression process is determined as the target model;
[0170] The target stop condition includes at least one of the following: the performance evaluation index of the candidate model after this compression does not meet the execution performance condition; the change state of the performance evaluation index of the candidate model after this compression compared with the performance evaluation index of the candidate model obtained by the previous compression process does not meet the target change condition.
[0171] In some embodiments, when the performance evaluation index of the candidate model does not meet the execution performance conditions, the pruning rate of each network layer is updated based on the initial pruning rate of each network layer and the pruning rate adjustment step of each network layer; the candidate model is compressed based on the updated pruning rate of each network layer to determine a third candidate model; and it is determined whether the third candidate model meets the execution performance conditions.
[0172] In this embodiment, the pruning rate for this compression process can be determined by the initial pruning rate and the pruning rate adjustment step size. By compressing the candidate model obtained in the previous compression process based on the pruning rate, the candidate model after this compression can be obtained. This gradual compression method by gradually increasing the pruning rate can solve the problem of a sudden drop in the performance of the target model caused by a one-time large-scale compression.
[0173] In some embodiments, the performance evaluation indicators of the candidate model include an accuracy indicator, and the target change condition includes an accuracy degradation threshold. When the difference between the accuracy indicator of the candidate model after this compression and the accuracy indicator obtained by the previous compression process is greater than the accuracy degradation threshold, it is determined that the change state of the performance evaluation indicator of the candidate model after this compression compared to the performance evaluation indicator of the candidate model obtained by the previous compression process does not meet the target change condition.
[0174] In this embodiment, the relationship between the difference between the accuracy index of the candidate model after this compression and the accuracy index obtained by the previous compression process, and the accuracy change threshold, can be used to determine whether to end the compression of the candidate model. If it is determined to be ended, the candidate model obtained by the previous compression will be determined as the target model. The performance evaluation index of the target model obtained in this way is relatively better.
[0175] In some embodiments, in step S101, obtaining the task processing model to be compressed may include the following steps S1011 and S1012, wherein:
[0176] Step S1011: performing perspective conversion on image data collected by at least one collection device to determine feature data corresponding to the image data under the perspective of a bird's-eye view;
[0177] Here, the acquisition device refers to a device that can automatically collect and transmit data, such as a video camera or a camera that can capture images. Bird's-Eye-View (BEV) is a technique for observing an object or scene from above. This perspective can provide a comprehensive view of a wide area, allowing people to clearly see the details and layout on the ground.
[0178] In some implementations, the image data collected by at least one collection device may be subjected to perspective conversion through geometric mapping technology, so as to convert the two-dimensional image data into a unified bird's-eye view perspective.
[0179] In some embodiments, the perspective conversion may be performed on image data collected by at least one acquisition device through a deep learning-based method to convert the two-dimensional image data into a unified bird's-eye view perspective, wherein the deep learning-based method may be, for example, algorithms such as BEVFormer and BEVFusion, which extract image features through a deep learning model and perform BEV conversion.
[0180] In some embodiments, the image data collected by at least one acquisition device may be converted into a unified bird's-eye view by performing perspective conversion based on traditional computer vision methods, wherein the traditional computer vision-based methods may be, for example, stereo matching, monocular or multi-camera depth estimation algorithms, which extract depth information from the image and project the information into the BEV perspective.
[0181] In some embodiments, after the image data is subjected to perspective conversion, the obtained BEV image may still contain noise, irregular edges or other information irrelevant to the perception task. By performing image enhancement processing on the BEV image, the image data quality can be optimized, making the image more suitable for input to the deep learning model. The image enhancement processing may include denoising, cropping and region selection, and image enhancement technology.
[0182] In some embodiments, after obtaining the optimized BEV image data, the optimized BEV image data is subjected to data standardization processing, which can effectively eliminate the image differences between different acquisition devices and ensure that the model can stably receive consistent input data during training and reasoning. Among them, the data standardization processing can include pixel value normalization, brightness and color consistency adjustment, and image size standardization. The BEV image data after data standardization processing is the feature data corresponding to the image data under the bird's-eye view perspective.
[0183] Step S1012: input the feature data into the initial task processing model, train the initial task processing model, and obtain the task processing model to be compressed.
[0184] In some embodiments, the initial task processing model may be an initial task processing model generated by using a convolutional neural network architecture, MixVarGENet as a basic backbone network, and a multi-task learning method (MTL).
[0185] In some embodiments, taking the vehicle autonomous driving scenario as an example, the initial task processing model can be used for three-dimensional target detection such as vehicle detection, pedestrian detection, drivable area segmentation, obstacle occupancy detection and lane line recognition.
[0186] In some implementations, training the initial task processing model may include a feature extraction process, a multi-task learning process, and a prediction output process.
[0187] In some embodiments, when extracting features from feature data, different levels of features are extracted from the input feature data through layer-by-layer convolution operations, wherein MixVarGENet, as a backbone network, is designed to handle large-scale, multi-task visual tasks, and can extract low-level geometric features (such as edges, corners) and high-level complex object features (such as the outlines and shapes of vehicles and pedestrians) of images layer by layer. The feature extraction process may include input data processing, convolutional layer feature extraction, and feature fusion.
[0188] In some embodiments, multi-task learning can be performed on the feature information in the extracted feature data. During implementation, a multi-task learning architecture can be used to set up an independent detection head for each specific task based on the shared convolutional feature extraction layer, so that multiple tasks can be processed simultaneously without repeating the feature extraction operation for each task. By sharing convolutional features in this way, not only can computing resources be effectively saved, but also the feature utilization rate of the model can be improved, so that each task can support each other and improve the overall perception accuracy. Among them, multi-task learning can include shared feature layer processing, independent detection head design, and task learning and optimization.
[0189] In some embodiments, based on feature extraction and multi-task learning, the final prediction results are generated by each independent detection head. Taking the vehicle automatic driving scenario as an example, each detection head outputs specific perception information for different tasks, including the location of vehicles and pedestrians, the segmentation of drivable areas, the occupancy status of obstacles, and the geometry of lane lines. Among them, for the vehicle and pedestrian detection heads, the location information of the target object is extracted through the convolution layer and the regression layer, and the three-dimensional bounding box of the object and the category information of the target object (vehicle or pedestrian) can be output. In addition, the detection head can also predict the confidence of each target object so as to filter out uncertain detection results during the reasoning process. For the drivable area segmentation detection head, it is implemented based on a convolutional fully connected network (Fully Convolutional Network, FCN), and the output is a binary mask map to identify the drivable area and non-drivable area on the road. The obstacle occupancy detection head extracts the spatial features of the object to identify static obstacles (such as guardrails, roadblocks) and dynamic obstacles (such as other vehicles or pedestrians) on the road. The detection head outputs the occupancy status of the obstacle, including the location, size, and occupancy ratio of the obstacle on the road. The lane line recognition detection head is based on the feature extraction of the convolutional network and can accurately identify the classification information of the lane lines on the road. The detection head outputs the segmentation information of the lane lines.
[0190] In this embodiment, first, by processing image data collected by at least one acquisition device, feature data corresponding to the image data at a bird's-eye view perspective is obtained, wherein the data image is converted to a bird's-eye view perspective, so that local information of the target environment at different perspectives can be integrated into a unified view, eliminating the impact of perspective differences, and then, based on the feature data unified in the bird's-eye view, it is input into the initial task processing model, so that the initial task processing model can perform model training without the need to integrate information of the image data, thereby improving the convergence speed and training efficiency of the task processing model.
[0191] The present application embodiment proposes an environment perception method, such as Figure 2 As shown, the environment perception method includes the following steps S201 and S202, wherein:
[0192] Step S201: Acquire image data of the target environment;
[0193] Here, the image data of the target environment is collected based on at least one collection device at the target environment.
[0194] Step S202: Utilize the target model and perform the target environment perception task based on the image data to obtain the environment perception result; the target model is determined based on the execution performance condition of the target environment perception task and after compressing the task processing model through the pruning strategy; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and the channels to be pruned are determined based on the pruning strategy; the execution performance condition is determined based on the task scenario of the target environment perception task.
[0195] In some embodiments, based on the image data, feature data corresponding to the image data from a bird's-eye view perspective is determined; the feature data is input into a target model, and an environmental perception result is obtained through the target model.
[0196] In this embodiment, first, the execution performance conditions are determined through the task scenario of the target environment perception task, and then the task processing model is compressed through a pruning strategy to obtain a target model, so that the target model is more closely matched with the task scenario of the target environment perception task. In this way, applying the target model to the same task scenario of the target environment perception task can not only reduce the waiting time for obtaining the environmental perception results, but also improve the calculation accuracy of the environmental perception results.
[0197] The following describes the application of the embodiments of the present application in actual scenarios.
[0198] With the rapid development of autonomous driving technology, the ability of vehicles to perceive the surrounding environment has become a key factor in achieving safe driving. As an indispensable part of autonomous driving, three-dimensional target detection greatly affects the vehicle's recognition, detection and obstacle avoidance decisions of surrounding pedestrians, obstacles and roads. In order to achieve 360-degree all-round perception of the surrounding environment, autonomous driving vehicles currently generally use sensor fusion solutions, relying on sensor data collected by sensors such as lidar, cameras, and millimeter-wave radars to complete three-dimensional target detection. Among them, although lidar can provide high-precision three-dimensional point cloud data, the use of lidar will increase the production cost of the vehicle, and the working principle and data processing of lidar sensors are relatively complex, and there are certain limitations in practical applications. Therefore, vision-based autonomous driving perception solutions are increasingly valued. The advantages of visual sensors (such as cameras) are low cost, light weight, flexible installation, easy deployment, and the ability to obtain rich semantic information. Therefore, pure visual perception solutions have gradually become the focus of autonomous driving research.
[0199] In the scenario of autonomous driving, BEV perspective perception technology provides an effective way for vehicles to perceive the three-dimensional environment. Through the multi-camera surround view system, the vehicle can obtain a bird's-eye view of the surrounding environment from a top-down perspective. BEV perception not only helps improve the vehicle's detection accuracy of obstacles in the surrounding environment, but also helps the vehicle make reasonable decisions in complex traffic scenarios.
[0200] Although the pure visual perception solution based on BEV has broad application prospects, the current 3D object detection technology still faces some challenges. First, the traditional 3D object detection model is often complex in structure, with a large number of parameters and high computational complexity, which makes it difficult to meet the needs of real-time reasoning for autonomous driving, especially in scenarios with high-resolution images and complex network structures. The delay in model reasoning becomes a bottleneck that affects the real-time decision-making of the vehicle. Secondly, the existing 3D object detection solutions usually handle different detection tasks such as vehicles, pedestrians, and obstacles separately, which makes it impossible to share feature information between tasks and increases the computational overhead.
[0201] To solve these problems, researchers began to explore solutions that combine multi-task learning with model compression technology. On the one hand, multi-task learning reduces the overall computational complexity by sharing convolutional features and allowing multiple tasks to be inferred simultaneously. On the other hand, model pruning technology reduces the amount of computation and storage requirements by cutting redundant parameters in the neural network, thereby increasing the inference speed. Combining these technologies will help promote further improvements in the real-time and reliability of autonomous driving technology.
[0202] However, in the pruning process of multi-task models, how to effectively select pruning strategies while ensuring model accuracy is a key challenge. The general channel pruning method is to directly delete unimportant channels, but this method may destroy the model structure and cause a significant decrease in the prediction effect of the pruned model. To address this problem, an innovative channel pruning strategy is proposed, combined with a model selection method based on a funnel-type screening mechanism to ensure that the pruned model can achieve a good balance between inference speed and detection accuracy.
[0203] The embodiment of the present application discloses a multi-task model compression method and system based on a bird's-eye view of the surround view. Through multi-task learning, it integrates tasks such as vehicle and pedestrian detection, movable area segmentation, obstacle occupancy detection, and lane line recognition. The model is streamlined by combining channel pruning technology, which can greatly reduce the calculation amount of the model while ensuring the accuracy of the model, thereby improving the reasoning efficiency and reducing the consumption of computing resources. The technical solution of the present application has a wide range of applicability, especially for real-time three-dimensional target detection tasks in autonomous driving scenarios.
[0204] First, let's introduce the multi-task model compression system based on bird's-eye view. The system consists of five main modules: preprocessing module, feature extraction and prediction module, pruning and arrangement module, fine-tuning module and analysis module. Each module has different responsibilities in the whole system, forming a complete, multi-task perception and compression closed-loop system.
[0205] 1. Preprocessing module
[0206] The preprocessing module is the first step of the entire surround bird's-eye view multi-task model optimization system. Its core function is to convert the two-dimensional image data collected by multiple cameras into a unified bird's-eye view perspective, and perform necessary cleaning, enhancement and formatting on these data, mainly including perspective conversion and data processing. Perspective conversion refers to the system using perspective conversion technology to uniformly convert images from different cameras into a bird's-eye view perspective. The converted bird's-eye view can provide a global spatial perception perspective for the vehicle and help the system capture the spatial information of objects such as vehicles, pedestrians and road obstacles. Data processing refers to the necessary denoising, cropping and normalization of the image to ensure the effectiveness of subsequent feature extraction.
[0207] This module plays a vital role in the entire system because the quality of preprocessed data directly affects the performance of subsequent feature extraction and multi-task learning models. Through the preprocessing module, the system can ensure the consistency and high quality of the input data, thus providing a solid foundation for subsequent steps.
[0208] In the scenario of autonomous driving, vehicles usually use a surround view system composed of multiple cameras to cover a 360-degree environmental perception range. The two-dimensional image obtained by each camera has a unique perspective and different spatial information. Therefore, a key task of the preprocessing module is to integrate these image data from different perspectives into the same coordinate system, that is, a bird's-eye view, so that the vehicle can obtain a global perception of the surrounding environment. This process is achieved through operations such as perspective conversion, image enhancement, and data standardization to improve the reliability and uniformity of the overall data. The specific implementation steps are as follows:
[0209] 1. Perspective conversion: Perspective conversion is the first step of the preprocessing module, which is responsible for converting the two-dimensional images from multiple cameras into a unified bird's-eye view. Since each camera is installed in a different position, the images it captures have different projection angles and field of view. If the original image is used directly for processing, it will not only cause local deviations in perception due to the difference in perspective, but also increase the complexity of subsequent feature extraction. Therefore, the perspective conversion through geometric mapping technology can ensure that the images captured by multiple cameras are uniformly mapped to the same coordinate system, thereby providing a unified global perspective for subsequent perception tasks. The benefit of doing so is that the implementation of perspective conversion enables the system to integrate local information from different perspectives into a unified view, eliminating the impact of perspective differences. This global perspective not only improves the system's perception ability, but also provides consistent input in spatial position for subsequent multi-task models, ensuring that the vehicle can more accurately detect and understand the environment during driving. The specific steps are:
[0210] ① Geometric mapping technology: Through geometric mapping technology, two-dimensional images of different perspectives are spatially transformed, and these images are projected onto the plane where the vehicle is located (usually the ground plane) according to the pre-set camera calibration parameters (including the camera's internal and external parameters). This process generates a bird's-eye view in which each pixel corresponds to a spatial position in the real world, thereby helping the system capture the three-dimensional spatial information of the object.
[0211] ② Construction of global perspective: The image data of each camera will be converted to a unified BEV perspective, which can provide the vehicle with a bird's-eye view of the global environment. The advantage of the BEV perspective is that it can more intuitively display the spatial distribution of objects, especially in road scenes, where the relative position relationship between vehicles, pedestrians and obstacles can be more clearly displayed.
[0212] 2. Image enhancement: After the perspective conversion, the obtained BEV image may still contain noise, irregular edges or other information irrelevant to the perception task, which will affect the accuracy of subsequent feature extraction. Therefore, the purpose of the image enhancement operation is to further optimize the data quality through denoising, cropping and normalization, so that the image is more suitable for input into the deep learning model. The benefit of doing so is that the image enhancement operation can effectively improve the quality of image data, eliminate noise and irrelevant information, thereby ensuring that the data input into the model is purer. Through cropping and normalization processing, the computational burden can also be effectively reduced and the overall efficiency of the system can be improved. In addition, the use of image enhancement technology enables the system to have better adaptability under various lighting conditions, providing high-quality input data for subsequent multi-task models. The specific steps are:
[0213] ① Denoising: In the converted BEV image, there may be interference information due to camera noise or unsatisfactory data acquisition conditions. Denoising techniques, such as Gaussian filtering and median filtering, can effectively eliminate these noises and enhance the clarity of the image. Denoising can significantly reduce the impact of irrelevant pixels on feature extraction and improve the accuracy of subsequent models.
[0214] ② Cropping and region selection: In order to reduce unnecessary computation and focus on important perception areas, the system crops the BEV image and selects the region of interest. By focusing on an area within a certain range around the vehicle, the system can reduce the size of the image and retain the most critical environmental information. This method can not only improve computational efficiency, but also effectively reduce the computational burden of the model.
[0215] ③ Image enhancement technology: The image quality is further optimized through brightness adjustment, contrast enhancement and other technologies, so that it can maintain stable input quality in various environments (such as low light, strong light, etc.). This process helps to improve the robustness of the system in various complex environments.
[0216] 3. Image data standardization: Image data from different cameras may be inconsistent in terms of color range, brightness, etc. Direct input into the deep learning model may affect the learning effect of the model. In order to ensure that the model can be trained stably and obtain consistent feature representation, the preprocessing module standardizes the input image data. Standardization can not only eliminate the differences between image data, but also improve the convergence speed and training efficiency of the model. The benefit of doing so is that the data standardization operation can effectively eliminate the image differences between different cameras, ensuring that the model can stably receive consistent input data during training and inference. Through standardization, the system not only improves the convergence speed of the model, but also reduces the errors between different data sources, providing a reliable input basis for subsequent feature extraction and multi-task learning. The specific steps are:
[0217] ① Pixel value normalization: In order to make the images collected by different cameras have a unified input range, the system normalizes the pixel values of the images to a certain range (for example, between 0 and 1 or -1 and 1). This standardization process can effectively prevent the negative impact of differences between different input data on model training.
[0218] ②Brightness and color consistency adjustment: The system adjusts the brightness and color of each image to make the images from each camera more visually consistent. This adjustment ensures that even data collected under different weather or lighting conditions can maintain consistency at the input stage.
[0219] ③ Image size standardization: Since the resolution of different cameras may be different, the system needs to standardize the image size to ensure that the images collected by each camera have the same dimensions before being input into the model. By adjusting the image resolution, the system can make each image match the expected input size of the model when it is input.
[0220] 2. Feature Extraction and Prediction Module
[0221] The feature extraction and prediction module is one of the core parts of the autonomous driving system. Its main task is to extract rich multi-level features from the input bird's-eye view through a convolutional neural network, and use these features to predict multiple tasks. This module uses the advanced MixVarGENet as the basic backbone network, and gradually generates a deep feature representation of the target object through layer-by-layer convolution operations. On this basis, the detection head completes the prediction of multiple tasks such as vehicle detection, pedestrian detection, drivable area detection, obstacle occupancy detection, and lane line recognition. This multi-task learning method can not only effectively share the underlying convolution features between multiple perception tasks, but also greatly reduce the amount of calculation for each task and reduce calculation redundancy, thereby significantly improving reasoning efficiency.
[0222] Through feature sharing and task allocation, this module can provide reliable prediction results for multiple tasks and reduce resource waste while improving feature utilization. In addition, MixVarGENet performs well in capturing low-level geometric features and high-level semantic information, enabling the system to efficiently detect complex objects and identify their spatial locations. The detailed implementation steps of this module include feature extraction of convolutional neural networks, design of multi-task learning structures, and prediction outputs of each task. The specific implementation steps are as follows:
[0223] 1. Convolutional neural network feature extraction: Feature extraction is the basic task of this module. It extracts features of different levels from the input BEV image through layer-by-layer convolution operations. MixVarGENet, as the backbone network, is designed to handle large-scale, multi-task visual tasks. It can extract low-level geometric features (such as edges and corners) and high-level complex object features (such as the outlines and shapes of vehicles and pedestrians) of the image layer by layer. The specific operation steps are as follows:
[0224] ① Input data processing: The high-quality bird's-eye view images output by the preprocessing module are input into the convolutional neural network. Since the BEV images have been processed by perspective conversion, cropping and enhancement, their structural information and geometric shapes are very clear and suitable as input for the neural network.
[0225] ② Convolutional layer feature extraction: MixVarGENet contains multiple convolutional layers, each of which extracts features at different levels through convolution operations and nonlinear activation functions. Low-level convolutional layers can extract geometric features such as edges and textures in images, while high-level convolutional layers can capture more complex object shapes and structures, such as vehicles, pedestrians, obstacles, etc.
[0226] ③ Feature fusion: After extracting features at different levels, MixVarGENet fuses these features to combine the underlying geometric information with the high-level semantic information, thereby providing a more comprehensive feature representation for subsequent multi-task prediction. Through this fusion, the system can identify various types of objects and scene information in complex road environments.
[0227] 2. Multi-task learning: In order to improve the computational efficiency of the system, the feature extraction and prediction module adopts a multi-task learning architecture. Based on the shared convolutional feature extraction layer, the system sets up an independent detection head for each specific task, so that the system can handle multiple tasks at the same time without having to repeat the feature extraction operation for each task. By sharing convolutional features in this way, the system can not only effectively save computing resources, but also improve the feature utilization of the model, so that each task can support each other and improve the overall perception accuracy. The specific operation steps are:
[0228] ① Shared feature layer: The features extracted by the convolutional neural network are passed layer by layer, and all tasks share these extracted underlying features. This feature sharing strategy enables different tasks to extract the required information from the same features, avoiding the redundancy problem of repeated feature calculations between different tasks.
[0229] ② Independent detection head design: In order to adapt to the prediction requirements of different tasks, the system designs an independent detection head for each task based on the shared feature layer. The detection head is responsible for extracting information related to the task from the shared features and generating the final prediction results. The detection head of each task consists of a specific convolutional layer, a fully connected layer, or other specific layers, which can be optimized according to the task requirements.
[0230] ③Task optimization and learning: In multi-task learning, each task performs backpropagation on the shared feature layer and updates the weights of the convolutional layer. In this way, different tasks can promote each other. For example, the vehicle detection task can help improve the accuracy of pedestrian detection, and vice versa. This collaborative optimization process enables the system to achieve good performance on multiple tasks.
[0231] 3. Prediction output: Based on feature extraction and multi-task learning, the system generates the final prediction results through each independent detection head. Each detection head outputs specific perception information for different tasks, including the location of vehicles and pedestrians, the segmentation of the drivable area, the occupancy status of obstacles, and the geometry of lane lines. The specific operation steps are as follows:
[0232] ① Vehicle and pedestrian detection: The vehicle and pedestrian detection head extracts the location information of the target object through the convolution layer and regression layer. The detection head can output the 3D bounding box of the object and the category information of the target object (vehicle or pedestrian). In addition, the detection head can also predict the confidence of each target object so as to filter out uncertain detection results during the reasoning process.
[0233] ② Driving area segmentation: The driving area segmentation task is implemented through a convolution-based fully connected network, and the output is a binary mask map to identify the drivable and non-drivable areas on the road. This task can provide strong support for vehicle path planning and ensure that the vehicle can identify the road area where it is safe to drive.
[0234] ③Obstacle occupancy detection: The obstacle detection head extracts the spatial features of the object and identifies static obstacles (such as guardrails, roadblocks) and dynamic obstacles (such as other vehicles or pedestrians) on the road. The detection head outputs the occupancy status of the obstacle, including the location, size and occupancy ratio of the obstacle on the road.
[0235] ④ Lane line recognition: The lane line recognition detection head is based on the feature extraction of the convolutional network and can accurately identify the classification information of the lane lines on the road. The detection head outputs the segmentation information of the lane lines.
[0236] 3. Pruning Arrangement Module
[0237] After the model training is completed, the pruning and permutation module prunes unimportant channels by analyzing the importance of each channel in the convolutional neural network (i.e., the contribution value of each channel mentioned above). Specifically, the pruning and permutation module calculates the importance of each convolution channel according to the global pruning rate (i.e., the global pruning rate mentioned above) and sorts them. For channels with lower importance, the module will not delete them directly, but set the parameters of the channels to zero. This processing method can reduce the damage to the model structure. At the same time, for different network parts, such as the backbone network and the detection head part, the pruning and permutation module can set different pruning rates to meet the computing requirements of different tasks. Through this differentiated pruning strategy, the system can significantly reduce the number of model parameters and calculations while maintaining high detection accuracy. In addition to the global pruning strategy, the pruning and permutation module also supports the importance evaluation of local layers and dynamically adjusts the pruning ratio according to the weight contribution of each layer. This adaptive pruning method ensures the flexibility and efficiency of the model and is suitable for the optimization needs of deep convolutional networks.
[0238] In the model pruning part, this application proposes several pruning strategies, aiming to effectively reduce the computational complexity of the model and improve the reasoning speed through different methods. The following is a detailed pruning plan and specific implementation steps:
[0239] 1. Global pruning strategy: The core idea of the global pruning strategy is to calculate the importance of each channel in each convolutional layer based on a global unified perspective, and prune the model according to the preset pruning rate. The specific steps are as follows:
[0240] ① Calculate channel importance: In the trained model, use the parameters of the BatchNorm2d layer or the weight amplitude to calculate the importance of each convolution channel. Usually, the scaling parameters in the BatchNorm2d layer reflect the activation strength of the channel. The larger the weight amplitude, the more important the channel.
[0241] ② Sorting and screening: Sort all channels according to the calculated channel importance. According to the preset global pruning rate, determine the number of channels to be pruned, and select channels with lower importance for pruning.
[0242] ③ Parameter zeroing: In order to avoid damaging the model structure caused by directly deleting channels, unimportant channels are not directly removed during pruning, but their parameters are set to zero. This method ensures that the model structure is not damaged, and some performance can still be restored during subsequent fine-tuning.
[0243] ④ Optimization of reasoning speed: By setting unimportant channel parameters to zero, although the channel is not completely removed, the model does not calculate the zeroed channel during the actual reasoning process, thereby greatly reducing the reasoning complexity and improving the reasoning speed.
[0244] 2. Local pruning strategy: The local pruning strategy provides a more flexible pruning method. According to the importance of different network parts, the pruning rate of each layer is dynamically adjusted to minimize the computational overhead while maintaining the detection accuracy of the model. The specific plan is as follows:
[0245] ① Specific layer pruning: Different parts of the model have different effects on accuracy. For example, the extraction of low-level features (referring to the basic elements and attributes of the image) in the backbone network has a greater impact on the overall accuracy, so a lower pruning rate should be set for these layers. In the detection head, since the main task of these layers is to output prediction results, the pruning ratio can be larger to reduce computational overhead.
[0246] ② Dynamic pruning adjustment: By analyzing the contribution rate of the convolution parameters of each layer, the system can adaptively adjust the pruning ratio of each layer. For example, for convolution layers with greater contributions, more channels are retained to maintain the overall performance of the model; for convolution layers with smaller contributions, more channels are pruned to avoid performance degradation caused by a global unified pruning rate.
[0247] 3. Funnel-type pruning and screening mechanism: To ensure that the pruned model meets the requirements in terms of reasoning speed and accuracy, this application proposes a funnel-type screening mechanism, which generates multiple models with different pruning rates and screens out models that meet the requirements layer by layer, ensuring that the final deployed model meets both reasoning efficiency and high detection accuracy.
[0248] ① Multi-model generation: The system generates multiple pruned models at one time according to different pruning rates. Models with different pruning rates have different effects on inference speed and accuracy, so it is necessary to find a balance through screening.
[0249] ②Inference speed evaluation: Perform inference speed test on the generated model (i.e. the above-mentioned inference speed) to select the model that meets the latency requirements. Since the main purpose of pruning is to improve inference efficiency, inference speed is the primary criterion for model selection.
[0250] ③Precision fine-tuning: For the models selected by the inference speed evaluation, further fine-tuning training is performed to restore the precision loss caused by pruning. Through fine-tuning, the detection performance of the model can be significantly improved to achieve a higher detection accuracy.
[0251] ④Final model selection: After inference speed evaluation and accuracy fine-tuning, the system selects the model with the best performance for deployment based on the final inference speed and detection accuracy.
[0252] 4. Pruning based on task importance: In a multi-task model, each task has different degrees of dependence on different parts of the model. For example, vehicle detection and pedestrian detection have high accuracy requirements, so more feature channels can be retained, while for simpler tasks such as road area segmentation, a higher proportion of pruning can be performed. Based on this, this application proposes a strategy for pruning based on task importance:
[0253] ① Task classification: Vehicle detection and pedestrian detection tasks have high requirements on accuracy. Therefore, more feature channels should be retained in the network layers related to these tasks. For relatively simple tasks, such as executable area segmentation and road sign recognition, a higher proportion of pruning can be adopted to reduce the computational burden.
[0254] ② Task weight setting: During multi-task learning, the system can set different weights for different tasks (i.e., the above-mentioned subtasks) to ensure that more resources are reserved for important tasks during pruning, while more aggressive optimization is performed on secondary tasks.
[0255] (5) Progressive pruning strategy: The progressive pruning strategy avoids the problem of sudden drop in model performance caused by a one-time large-scale pruning by gradually increasing the pruning ratio. The main steps of this strategy are as follows:
[0256] ① Gradually increase the pruning ratio: During the training process, the pruning ratio is gradually increased from a smaller value, and the system gradually reduces the number of model parameters through multiple stages of pruning operations.
[0257] ②Performance monitoring and adjustment: At each pruning stage, the system monitors the performance of the model. If it is found that pruning at a certain stage has a greater impact on performance, pruning can be stopped or rolled back to the previous state in a timely manner. This strategy makes the accuracy of the model drop more gradually and facilitates the adjustment of the pruning process.
[0258] 4. Fine-tuning module
[0259] In the process of optimizing deep learning models, pruning is a commonly used technical means that can significantly reduce the number of model parameters and computational complexity, thereby improving reasoning speed and efficiency. However, pruning operations will inevitably lead to a decrease in the detection accuracy of the model, especially in the case of large-scale pruning, the model may lose some key feature representation capabilities. In order to restore the performance of the pruned model, the fine-tuning module helps the pruned model to re-adapt to its streamlined structure through further training and optimization, and ensures that its prediction ability is not significantly affected. During the fine-tuning process, the system dynamically adjusts the training parameters according to the amount of pruning, and retrains the pruned model to restore its prediction ability. Through fine-tuning, the pruned model can achieve an ideal balance between reasoning speed and detection accuracy.
[0260] The role of the fine-tuning module is not only to simply restore the accuracy of the pruned model, but also to further improve the generalization ability of the model by optimizing training parameters and introducing data enhancement technology, so that it can maintain efficient performance in various complex scenarios. This module dynamically adjusts training parameters such as learning rate and batch size according to the number of channels after pruning, changes in model structure, etc., to ensure that the pruned network can be retrained under optimal conditions. In addition, the fine-tuning module introduces a variety of data enhancement technologies to expand the diversity of training data and improve the robustness and generalization ability of the model, thereby ensuring that the model not only performs well on known data sets, but also maintains high performance in actual application scenarios. The specific implementation steps are as follows:
[0261] 1. Training adjustment after pruning: The essence of pruning is to remove redundant or unimportant channels in the network to reduce the number of model parameters and computational burden. However, the pruning process usually destroys the model's weight distribution and feature extraction capabilities. Therefore, the primary task of the fine-tuning module is to adjust the training parameters according to the number of channels and changes in the model structure after pruning to help the model gradually restore its prediction performance. The specific steps are:
[0262] ① Learning rate adjustment: During the fine-tuning process after pruning, the setting of the learning rate (which controls the step size of the model at each weight update) is crucial. Since the pruning operation reduces the number of channels in the network, the remaining channels need to take on more feature representation tasks, and the system needs to adjust the learning rate to adapt to the new weight update requirements. Generally speaking, a moderate reduction in the learning rate can allow the model to restore accuracy more stably under the new weight distribution, avoiding excessive disturbances to the model caused by drastic parameter updates.
[0263] ② Batch size adjustment: After pruning, the number of model parameters is reduced, and the computing resources occupied are also reduced accordingly. The system can increase the batch size of training to process more data samples in the same training cycle. This can not only speed up the training process, but also improve the generalization ability of the model, because a larger batch size can provide the model with more sample gradient information, which helps to update the weights more stably.
[0264] ③ Optimization of loss function: After pruning, the network structure changes, and the optimization target of the loss function (the difference or error between the quantified model prediction value and the true value) also needs to be adjusted according to the characteristics of the pruned model. The system needs to dynamically adjust the weight coefficients in the loss function according to different model tasks (such as object detection, semantic segmentation, etc.) to ensure that the model's prediction performance can be gradually restored after pruning.
[0265] 2. Additional data enhancement: Data enhancement is a common technique for expanding the diversity of training datasets. It can simulate different scenes and perspective changes by introducing different image transformation methods (such as rotation, translation, scaling, etc.), thereby improving the generalization ability of the model. In the fine-tuning process after pruning, the introduction of data enhancement technology is particularly important, because the pruned model has changed in both weight and feature extraction capabilities. Data enhancement can effectively prevent the model from overfitting to the weight distribution before pruning.
[0266] 3. Gradually restore accuracy: The pruning operation cuts off some channels of the model, resulting in a decrease in the detection accuracy of the model in the initial stage. The core task of the fine-tuning module is to gradually restore the accuracy of the model through further training, so that it can maintain high detection performance while significantly improving the inference speed. To achieve this goal, the system dynamically monitors the loss changes during the training process and gradually optimizes the weight distribution of the model to ensure that the model can gradually recover to the accuracy level before pruning. The specific steps are as follows:
[0267] ① Dynamic monitoring of training loss: The system continuously monitors the model's training loss and validation set performance during the training process. When the loss tends to be stable, the system can adjust the training parameters appropriately according to the situation, such as further reducing the learning rate, adjusting the gradient descent algorithm, etc., to ensure that the model can gradually restore accuracy.
[0268] ② Local fine-tuning strategy: For the layers that are most affected in the model after pruning, the system can adopt a local fine-tuning strategy to make more detailed weight adjustments for specific layers. By performing more frequent updates locally, the system can accelerate the weight recovery of key layers, thereby improving the overall accuracy of the model.
[0269] ③Strategy of gradually improving accuracy: As training progresses, the accuracy of the model will gradually recover. The system prioritizes the recovery of detection heads for specific tasks based on the priorities of different tasks (such as vehicle detection, pedestrian detection, etc.), ensuring that the detection accuracy of key tasks reaches the expected level first. Through this strategy of gradually improving accuracy, the system can ensure that the prediction results of each task remain reliable while significantly improving the reasoning efficiency of the pruned model.
[0270] 5. Evaluation Module
[0271] The evaluation module is responsible for a comprehensive evaluation of the pruned model. First, the evaluation module evaluates the compression effect of the model, including indicators such as the reduction of model parameters, the reduction of computational complexity, and the improvement of inference speed. Secondly, the system also evaluates the target detection performance after pruning to ensure that the pruning operation does not significantly affect the detection accuracy of the model. Through the comprehensive evaluation of these two aspects, the system finally selects the pruned model with the best performance.
[0272] The core task of the evaluation module is to ensure that the compression effect, inference speed, detection accuracy and task performance of the pruned model meet the expected goals by performing multi-dimensional analysis and evaluation on the pruned model. Specifically, first, the evaluation module needs to examine the compression effect of the pruned model, including the reduction in the number of model parameters, the reduction in computational complexity and the improvement in inference speed. Secondly, the system will also evaluate the detection performance of the pruned model to ensure that the pruning operation does not significantly affect the detection accuracy of the model, the retention of detection accuracy, and the task performance trade-off under the multi-task learning framework. Through these evaluation indicators, the system can comprehensively analyze the impact of pruning operations on the model and select the optimal pruned model for actual deployment. The following are several important evaluation indicators:
[0273] 1. Model compression effect: Model compression effect is one of the core goals of pruning technology. By comparing the number of parameters and computational complexity of the model before and after pruning, evaluate whether the pruning operation significantly reduces the number of parameters and computing resource consumption of the model. For complex perception tasks in autonomous driving scenarios, reducing the computational complexity of the model can directly reduce the computational overhead during reasoning, thereby improving the response speed of the system. The specific steps are as follows:
[0274] ① Parameter comparison: The evaluation module first calculates the parameters of the model before and after pruning. By comparing, we can intuitively understand the compression effect of pruning on the model. If the number of model parameters after pruning is significantly reduced, it means that pruning has successfully removed unimportant channels and redundant weights.
[0275] ② Computational complexity comparison: In addition to the number of parameters, the evaluation module also needs to evaluate the computational complexity of the model before and after pruning (usually measured in floating-point operations FLOPs). The pruning operation should effectively reduce the computational operations that the model needs to perform during the inference process, thereby reducing the consumption of computing resources.
[0276] ③Memory usage and storage requirements: The evaluation module also needs to measure the memory usage and storage requirements of the pruned model during inference. Reducing the number of model parameters and computational complexity can not only increase the inference speed, but also reduce the demand for hardware memory, making the model more suitable for deployment on devices with limited resources.
[0277] 2. Improved reasoning speed: Improving the reasoning speed of the model is one of the main goals of pruning technology, especially in autonomous driving scenarios, where the model needs to detect and respond to vehicles, pedestrians, road obstacles, etc. in real time in a very short time. Therefore, the evaluation module needs to test the reasoning speed of the pruned model to ensure that it can significantly improve the reasoning efficiency on different hardware platforms. The specific steps are as follows:
[0278] ① Inference speed test: The evaluation module performs multiple inference speed tests on the pruned model and counts the inference latency of the model on different hardware platforms (such as GPU, CPU, and embedded devices). The latency measurement indicators include the processing time per frame and the number of frames processed per second.
[0279] ② Hardware adaptability test: In order to ensure that the pruned model can run efficiently on different hardware platforms, the system needs to test the model in multiple hardware environments. For example, an autonomous driving system may be deployed on a GPU with high computing power, but it may also be deployed on an embedded device or an in-vehicle computing platform. The evaluation module needs to ensure that the pruned model can meet the requirements of real-time processing under different hardware conditions.
[0280] ③Inference efficiency comparison: By comparing the inference speed of the model before pruning, the system can intuitively measure the degree to which the pruning operation improves the inference efficiency. The evaluation module should ensure that the inference speed of the pruned model is significantly improved and can meet the real-time requirements in the autonomous driving scenario.
[0281] (3) Detection accuracy retention: Although pruning can reduce the number of model parameters and computational complexity, it may also lead to a decrease in detection accuracy, especially when pruning is performed drastically, the model may lose some feature extraction capabilities. Therefore, the evaluation module needs to focus on the detection accuracy retention of the pruned model on each task to ensure that the model does not significantly affect the prediction accuracy of the task while improving the inference speed. The specific steps are as follows:
[0282] ① Standard detection index evaluation: The system uses a series of standard detection indicators (such as precision, recall, average precision, etc.) to evaluate the accuracy of the model before and after pruning in tasks such as vehicle detection, pedestrian detection, and road area segmentation. The pruned model should remain stable in accuracy to avoid a significant drop.
[0283] ② Evaluation of fine-tuning effect after pruning: After the pruning operation, the system restores the detection accuracy of the model through fine-tuning training. The evaluation module needs to test the fine-tuned model to ensure that its accuracy is restored to the level before pruning. By comparing the model performance before pruning, after pruning, and after fine-tuning, you can intuitively understand the actual effect of fine-tuning on accuracy recovery.
[0284] ③ Comparison of detection task accuracy: The evaluation module also needs to compare the accuracy of different tasks one by one to ensure that the pruned model can maintain high-precision performance in tasks such as vehicle detection, pedestrian detection, and obstacle recognition. For multi-task models, the retention of detection accuracy is crucial because the prediction results of each task will directly affect the decision-making of the autonomous driving system.
[0285] 4. Task performance trade-off: In multi-task learning, different tasks have different tolerances for model pruning. For example, vehicle detection and pedestrian detection tasks usually require higher model accuracy, while tasks such as executable area segmentation may have a higher tolerance for pruning. Therefore, the evaluation module needs to analyze the detection results of each task to ensure that the performance of each task is reasonably balanced during the pruning process, focus on optimizing key tasks, and avoid over-optimization of low-priority tasks that affect the overall performance of the model. The specific steps are as follows:
[0286] ① Task priority setting: The system sets the task priority according to the importance of different tasks. For example, vehicle and pedestrian detection are key tasks in the autonomous driving system and have a higher priority. After pruning, the system should give priority to retaining the detection accuracy of these tasks.
[0287] ②Critical task accuracy evaluation: For tasks with higher priority (such as vehicle and pedestrian detection), the system needs to focus on the impact of pruning on these tasks. The evaluation module needs to ensure that the detection accuracy of the model on these key tasks will not be significantly affected after pruning, ensuring the decision-making accuracy of the autonomous driving system.
[0288] ③ Optimization of low-priority tasks: For low-priority tasks (such as executable area segmentation or obstacle detection), the evaluation module can allow more optimization during the pruning process, thereby further reducing the computational complexity and improving the reasoning speed. This task-performance trade-off strategy ensures that while optimizing the overall performance of the model, the detection accuracy of key tasks is focused on.
[0289] Next, we introduce the multi-task model compression task based on the bird's-eye view, such as Figure 3A As shown, the multi-task model compression task based on the bird's-eye view includes the following steps S301 to S305, wherein:
[0290] S301: converting the two-dimensional image data (i.e., the above-mentioned image data) collected by multiple cameras into a unified bird's-eye view, and then encoding it into a BEV image as an image input;
[0291] Here, it is performed through the preprocessing module.
[0292] like Figure 3B As shown, by collecting images 11 by multiple cameras installed on the vehicle body 1, multiple two-dimensional images can be obtained, and the multiple two-dimensional images are input into the preprocessing module 2. The preprocessing module 2 can convert the two-dimensional images into a unified bird's-eye view.
[0293] S302: extracting features from the input data (i.e., the feature data) and performing prediction in the convolutional neural network model (i.e., the initial task processing model);
[0294] Here, the output of the preprocessing module is input into the feature extraction and prediction module, which performs feature extraction and prediction on the input data to generate a model to be pruned.
[0295] like Figure 3B As shown, the image data from the bird's-eye view is input into the feature extraction and prediction module 3, and the DNN 31 network model is trained to obtain a trained DNN network model.
[0296] S303: Multi-model generation, generating multiple pruned models at one time according to different pruning rates, using a channel pruning scheme, and arranging the models according to the importance of each channel after the model training is completed;
[0297] Here, the model to be pruned output by the feature extraction and prediction module is input into the pruning and arrangement module, and the pruning and arrangement module generates multiple compressed models.
[0298] like Figure 3B As shown, the trained DNN network model is input into the pruning and arranging module 4, and a plurality of pruned models can be obtained through the processing of the pruning and arranging module 4.
[0299] In some embodiments, the channel pruning scheme may be a plurality of pruning rates preset for each network layer in the model to be compressed, so that the model to be pruned is pruned at a plurality of pruning rates respectively to obtain a plurality of pruned models. In some embodiments, the channels in each network layer may be sorted according to their importance, and the channels may be removed in order from low to high importance until the corresponding pruning rate is met to obtain a pruned model.
[0300] For example, the above pruning process is explained using a pruning strategy based on task importance, such as Figure 3C As shown, during implementation, five pruning rate setting schemes 41 are selected to set the pruning rate for each task in the trained DNN network model 42, and then, the trained DNN network model 42 is pruned based on the pruning rate of each task set by each of the above pruning rate setting schemes 41 to obtain the pruned model 43. It can be seen that after the implementation of all pruning rate setting schemes, five pruned models can be obtained, namely, model 431, model 432, model 433, model 434, and model 435. Among them, the five pruning rate setting schemes may include a pruning rate setting 411 based on fixed weights, a pruning rate setting 412 based on dynamic weights, a pruning rate setting 413 based on gradient sensitivity, a pruning rate setting 414 based on task uncertainty, and a pruning rate setting 415 based on a learning method.
[0301] In other embodiments, the models to be pruned may be pruned using multiple pruning strategies in the pruning arrangement module to obtain multiple models. For example, the models to be pruned may be pruned using a local pruning strategy, a pruning strategy based on task importance, a progressive pruning strategy, and a global pruning strategy to obtain multiple models.
[0302] In other implementations, the model to be pruned may be pruned using a global pruning strategy in a pruning arrangement module to obtain a first model, and then the first model may be pruned using multiple pruning strategies in a pruning arrangement module to obtain multiple models. For example, the model to be pruned may be pruned using a local pruning strategy, a pruning strategy based on task importance, and a progressive pruning strategy to obtain multiple models.
[0303] It should be noted that the pruning strategy based on task importance corresponds to the description of steps S110 to S112 above; the local pruning strategy corresponds to the description of steps S113 to S115 above; the global pruning strategy corresponds to the description of steps S116 and S117 above; and the progressive pruning strategy corresponds to the description of steps S1036 to S1038 above.
[0304] S304: a funnel-type pruning screening mechanism generates multiple models with different pruning rates, and screens models that meet the requirements layer by layer;
[0305] Here, first, by evaluating the inference speed of the multiple pruned models generated in the above step S203, when implementing, the inference speeds of the multiple pruned models are obtained respectively, and compared with the preset inference speed threshold, the pruned models that meet the inference speed threshold are determined as candidate models to generate a first candidate set; then, each model in the first candidate set is fine-tuned and trained to obtain a second candidate set; finally, the inference speed and accuracy of each model in the second candidate set are evaluated to determine the final model to be pruned.
[0306] like Figure 3B As shown, the pruning and arrangement module 4 performs performance evaluation on multiple pruned models to screen out models that meet performance requirements.
[0307] S305: fine-tune the selected pruned model, and consider the amount of fine-tuning according to the amount of pruning;
[0308] Here, the model to be pruned is fine-tuned through the fine-tuning module to obtain a fine-tuned model under different training parameters.
[0309] like Figure 3B As shown, the screened model that meets the requirements is input into the fine-tuning module 5, and fine-tuning training is performed through the fine-tuning module 5 to obtain a fine-tuned trained model.
[0310] S306: Comprehensively analyze the model from two aspects: model compression and target detection to obtain the final model (i.e., the target model mentioned above).
[0311] Here, multiple fine-tuned models are evaluated through the evaluation module, and the final model is obtained from two dimensions: model compression effect and target detection.
[0312] like Figure 3B As shown, the fine-tuned trained model is input into the evaluation module 6, which evaluates the fine-tuned trained model from two dimensions: model compression effect and target detection, and determines the optimal model obtained by the evaluation as the final model.
[0313] In summary, the embodiment of the present application proposes a new pruning strategy that combines global pruning and local pruning. First, the system calculates the importance of the convolution layer parameters globally and uniformly, and sorts all channels according to their importance. Then, the system determines the number of channels that need to be pruned based on the preset global pruning rate, and sets these unimportant channel parameters to zero instead of directly deleting them, thereby retaining the flexibility of the model.
[0314] In addition, the system can set different pruning rates for different network parts. For example, in the backbone network, the system can adopt a lower pruning rate to ensure the integrity of the underlying features; while in the detection head, the system can adopt a higher pruning rate to speed up the reasoning process.
[0315] In order to further improve the efficiency of model selection after pruning, this application introduces a funnel-type screening mechanism. Specifically, the system generates multiple models with different pruning rates at one time and evaluates the inference speed of these models. Then, the system screens out models whose delays meet the requirements based on the preset inference delay standard. Next, these models are fine-tuned for accuracy to restore the accuracy loss caused by pruning. Through this funnel-type screening mechanism, the system can efficiently select models with both performance and efficiency.
[0316] Based on the above embodiments, the present application provides a model compression device, such as Figure 4 As shown, the model compression device 400 includes:
[0317] The first acquisition module 401 is used to acquire a task scenario of a target environment perception task and a task processing model to be compressed; the task processing model includes at least one network layer, and the network layer includes at least one channel;
[0318] A determination module 402 is used to determine the execution performance conditions of the target environment perception task based on the task scenario;
[0319] The first obtaining module 403 is used to compress the task processing model through a pruning strategy based on the execution performance conditions to obtain a target model; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the channels to be pruned are determined based on the pruning strategy, and the target model is used to execute the target environment perception task and meet the execution performance conditions.
[0320] In some embodiments, the first obtaining module includes: a first determination unit, used to determine a target pruning strategy based on execution performance conditions; based on the target pruning strategy, determining the channels to be pruned in the task processing model; the first obtaining unit, used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, to obtain the target model.
[0321] In some embodiments, the execution performance conditions include target accuracy conditions, target speed conditions and / or target computing volume conditions, and the target pruning strategy includes a target pruning rate; the first determination unit includes at least one of the following: a first determination subunit, used to determine the inference accuracy of the task processing model based on the target accuracy condition; determine the target pruning rate based on the inference accuracy; the inference accuracy includes the detection accuracy corresponding to at least one detection head in the task processing model; a second determination subunit, used to determine the inference speed of the task processing model based on the target speed condition; determine the target pruning rate based on the inference speed; the inference speed includes the detection speed corresponding to at least one detection head in the task processing model; a third determination subunit, used to determine the inference computing volume of the task processing model based on the target computing volume condition; determine the target pruning rate based on the inference computing volume; the inference speed includes the detection computing volume corresponding to at least one detection head in the task processing model.
[0322] In some embodiments, the first obtaining module includes: a second determination unit, used to determine at least one target pruning strategy; the second obtaining unit, used to determine, for each target pruning strategy, the channels to be pruned in the task processing model based on the target pruning strategy, and set the channel parameters corresponding to the channels to be pruned in the task processing model to zero to obtain a candidate model; and a third determination unit, used to determine the target model from each candidate model based on an execution performance condition.
[0323] In some embodiments, the third determination unit includes: a fourth determination subunit, used to determine the performance evaluation index of each candidate model respectively; and a selection subunit, used to select a target model whose performance evaluation index meets the execution performance condition from each candidate model.
[0324] In some embodiments, the task processing model includes a multi-task model, the multi-task model includes a feature extraction module shared by multiple subtasks and a detection head module corresponding to each subtask, and the target environment perception task includes at least one subtask; the first determination unit or the second determination unit includes: a fifth determination subunit, used to determine the weight value of each subtask corresponding to the multi-task model; a sixth determination subunit, used to determine the pruning rate of the detection head module corresponding to each subtask based on the weight value of each subtask; and a seventh determination subunit, used to determine the channel to be pruned in the network layer of the detection head module based on the pruning rate of each detection head module.
[0325] In some embodiments, the first determination unit or the second determination unit includes: an eighth determination subunit, used to determine the contribution rate of each network layer in the task processing model; a ninth determination subunit, used to determine the pruning rate of each network layer based on the contribution rate of each network layer; and a tenth determination subunit, used to determine the channels to be pruned in each network layer based on the pruning rate of each network layer.
[0326] In some embodiments, the first determination unit or the second determination unit includes: an acquisition subunit, used to obtain the contribution value of each channel of each network layer in the task processing model; an eleventh determination subunit, used to determine the channel to be pruned in each network layer based on the global pruning rate and the contribution value of each channel.
[0327] In some embodiments, the first obtaining module includes: a fourth determination unit, which is used to determine the channels to be pruned in each network layer based on the initial pruning rate of each network layer in the task processing model, and set the channel parameters corresponding to the channels to be pruned in the task processing model to zero to obtain a candidate model; a fifth determination unit, which is used to determine the performance evaluation index of the candidate model; a sixth determination unit, which is used to compress the candidate model at least once through the pruning strategy when the performance evaluation index of the candidate model meets the execution performance condition, until the performance evaluation index of the candidate model after this compression meets the target stop condition, and determine the candidate model obtained by the last compression process as the target model; the target stop condition includes at least one of the following: the performance evaluation index of the candidate model after this compression does not meet the execution performance condition; the change state of the performance evaluation index of the candidate model after this compression compared with the performance evaluation index of the candidate model obtained by the last compression process does not meet the target change condition.
[0328] In some embodiments, the performance evaluation indicators of the candidate model include an accuracy indicator, and the target change condition includes an accuracy change threshold. When the difference between the accuracy indicator of the candidate model after this compression and the accuracy indicator obtained by the previous compression process is greater than the accuracy change threshold, it is determined that the change state of the performance evaluation indicator of the candidate model after this compression compared to the performance evaluation indicator of the candidate model obtained by the previous compression process does not meet the target change condition.
[0329] In some embodiments, the first acquisition module includes: a seventh determination unit, used to perform perspective conversion on image data collected by at least one acquisition device, and determine feature data corresponding to the image data under the bird's-eye view perspective; a third obtaining unit, used to input the feature data into an initial task processing model, train the initial task processing model, and obtain the task processing model to be compressed.
[0330] The present application embodiment provides an environment sensing device, such as Figure 5 As shown, the environment perception device 500 includes:
[0331] The second acquisition module 501 is used to acquire image data of the target environment;
[0332] The second obtaining module 502 is used to use the target model to perform the target environment perception task based on the image data to obtain the environment perception result; the target model is determined based on the execution performance conditions of the target environment perception task, and the task processing model is compressed through the pruning strategy; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and the channels to be pruned are determined based on the pruning strategy; the execution performance conditions are determined based on the task scenario of the target environment perception task.
[0333] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiment of the present application can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0334] An embodiment of the present application also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements some or all of the steps in the above method when executing the program.
[0335] The embodiment of the present application also proposes a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.
[0336] The embodiment of the present application also proposes a computer program, including a computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0337] The present application also proposes a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0338] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.
[0339] It should be noted that the embodiment of the present application provides a hardware entity of a computer device, such as Figure 6 As shown, the hardware entity of the computer device 600 includes: a processor 601 generally controls the overall operation of the computer device 600. A communication interface 602 can enable the computer device to communicate with other terminals or servers through a network. A memory 603 is configured to store instructions and applications executable by the processor 601, and can also cache data to be processed or processed by the processor 601 and each module in the computer device 600 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be performed between the processor 601, the communication interface 602, and the memory 603 through a bus 604.
[0340] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.
[0341] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0342] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0343] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the embodiments of the present application may be all integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0344] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0345] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0346] The above embodiments are only preferred embodiments for fully illustrating the present application, and the protection scope of the present application is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present application is within the protection scope of the present application.
Claims
1. A model compression method, characterized in that: The method comprises: Acquire a task scenario of a target environment perception task and a task processing model to be compressed; the task processing model includes at least one network layer, and the network layer includes at least one channel; Based on the task scenario, determining the execution performance conditions of the target environment perception task; Based on the execution performance conditions, the task processing model is compressed through a pruning strategy to obtain a target model; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the channels to be pruned are determined based on the pruning strategy, and the target model is used to execute the target environment perception task and meet the execution performance conditions.
2. The method according to claim 1, characterized in that: Based on the execution performance condition, the task processing model is compressed by a pruning strategy to obtain a target model, including: Based on the execution performance condition, determining a target pruning strategy; based on the target pruning strategy, determining a channel to be pruned in the task processing model; The channel parameters corresponding to the channel to be pruned in the task processing model are set to zero to obtain a target model.
3. The method according to claim 2, characterized in that: The execution performance condition includes a target accuracy condition, a target speed condition and / or a target computation amount condition, and the target pruning strategy includes a target pruning rate; The determining of the target pruning strategy based on the execution performance condition includes at least one of the following: Based on the target accuracy condition, determining the inference accuracy of the task processing model; based on the inference accuracy, determining the target pruning rate; the inference accuracy includes the detection accuracy corresponding to at least one detection head in the task processing model; Based on the target speed condition, determining the reasoning speed of the task processing model; Determining the target pruning rate based on the inference speed; the inference speed includes a detection speed corresponding to at least one detection head in the task processing model; Based on the target computational load condition, the inference computational load of the task processing model is determined; based on the inference computational load, the target pruning rate is determined; the inference speed includes the detection computational load corresponding to at least one detection head in the task processing model.
4. The method according to claim 1, characterized in that: Based on the execution performance condition, the task processing model is compressed by a pruning strategy to obtain a target model, including: Determine at least one target pruning strategy; For each of the target pruning strategies, based on the target pruning strategy, determine the channel to be pruned in the task processing model, and set the channel parameters corresponding to the channel to be pruned in the task processing model to zero, so as to obtain a candidate model; Based on the execution performance condition, a target model is determined from each of the candidate models.
5. The method according to claim 4, characterized in that: The step of determining a target model from each of the candidate models based on the execution performance condition includes: Determining the performance evaluation index of each candidate model respectively; From each of the candidate models, a target model whose performance evaluation index satisfies the execution performance condition is selected.
6. The method according to claim 2 or 4, characterized in that: The task processing model includes a multi-task model, the multi-task model includes a feature extraction module shared by multiple subtasks and a detection head module corresponding to each of the subtasks, and the target environment perception task includes at least one of the subtasks; The determining, based on the target pruning strategy, the channel to be pruned in the task processing model includes: Determining a weight value of each subtask corresponding to the multi-task model; Based on the weight value of each of the subtasks, respectively determine the pruning rate of the detection head module corresponding to each of the subtasks; Based on the pruning rate of each of the detection head modules, channels to be pruned in the network layer of the detection head module are determined.
7. The method according to claim 2 or 4, characterized in that: The determining, based on the target pruning strategy, the channel to be pruned in the task processing model includes: Determining a contribution rate of each network layer in the task processing model; Based on the contribution rate of each of the network layers, respectively determine the pruning rate of each of the network layers; Based on the pruning rate of each of the network layers, channels to be pruned in each of the network layers are determined.
8. The method according to claim 2 or 4, characterized in that: The determining, based on the target pruning strategy, the channel to be pruned in the task processing model includes: Obtaining a contribution value of each channel of each of the network layers in the task processing model; Based on the global pruning rate and the contribution value of each of the channels, a channel to be pruned in each of the network layers is determined.
9. The method according to any one of claims 1 to 5, characterized in that: Based on the execution performance condition, the task processing model is compressed by a pruning strategy to obtain a target model, including: Based on the initial pruning rate of each of the network layers in the task processing model, determine the channels to be pruned in each of the network layers, set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and obtain a candidate model; Determining a performance evaluation indicator of the candidate model; When the performance evaluation index of the candidate model meets the execution performance condition, the candidate model is compressed at least once by using a pruning strategy until the performance evaluation index of the candidate model after the current compression meets the target stop condition, and the candidate model obtained by the previous compression process is determined as the target model; The target stop condition includes at least one of the following: the performance evaluation index of the candidate model after this compression does not meet the execution performance condition; the change state of the performance evaluation index of the candidate model after this compression compared with the performance evaluation index of the candidate model obtained by the previous compression process does not meet the target change condition.
10. The method according to claim 9, characterized in that: The performance evaluation indicators of the candidate model include accuracy indicators, and the target change condition includes an accuracy change threshold. When the difference between the accuracy indicator of the candidate model after this compression and the accuracy indicator obtained by the previous compression process is greater than the accuracy change threshold, it is determined that the change state of the performance evaluation indicator of the candidate model after this compression compared to the performance evaluation indicator of the candidate model obtained by the previous compression process does not meet the target change condition.
11. The method according to any one of claims 1 to 5, characterized in that: Get the task processing model to be compressed, including: Performing perspective conversion on image data collected by at least one collection device to determine feature data corresponding to the image data under the perspective of a bird's-eye view; The feature data is input into an initial task processing model, and the initial task processing model is trained to obtain a task processing model to be compressed.
12. A method for environmental perception, characterized in that: The method comprises: Acquire image data of the target environment; Using the target model, based on the image data, a target environment perception task is performed to obtain an environment perception result; the target model is determined based on the execution performance conditions of the target environment perception task, after compressing the task processing model through a pruning strategy; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and the channels to be pruned are determined based on the pruning strategy; the execution performance conditions are determined based on the task scenario of the target environment perception task.
13. A model compression device, characterized in that: The device comprises: A first acquisition module is used to acquire a task scenario of a target environment perception task and a task processing model to be compressed; the task processing model includes at least one network layer, and the network layer includes at least one channel; A determination module, used to determine the execution performance conditions of the target environment perception task based on the task scenario; The first obtaining module is used to compress the task processing model through a pruning strategy based on the execution performance conditions to obtain a target model; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, the channels to be pruned are determined based on the pruning strategy, and the target model is used to execute the target environment perception task and meet the execution performance conditions.
14. An environment sensing device, characterized in that: The device comprises: A second acquisition module is used to acquire image data of the target environment; The second obtaining module is used to use the target model to perform the target environment perception task based on the image data to obtain the environment perception result; the target model is determined based on the execution performance conditions of the target environment perception task, and the task processing model is compressed through a pruning strategy; the compression processing is used to set the channel parameters corresponding to the channels to be pruned in the task processing model to zero, and the channels to be pruned are determined based on the pruning strategy; the execution performance conditions are determined based on the task scenario of the target environment perception task.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 12 are implemented.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
17. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 12 are implemented.
Citation Information
Cited By
Flight control end intelligent algorithm deployment method based on adaptive pruning
CN120725086A
A method for deploying intelligent algorithms for flight control based on adaptive pruning
CN120725086B
Communication segmentation learning system and method for adaptive channel compression, and medium
CN120896677A