Multi-task processing method and device, equipment and medium

By employing a perception task model in the driver assistance system and utilizing the shared BEV features between the encoding module and the output head, the problem of wasted computing resources is solved, and the utilization of computing resources is maximized.

CN121859232APending Publication Date: 2026-04-14ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In driver assistance systems, the large number of separately designed perception task models leads to a waste of vehicle-side computing resources and low utilization of computing resources.

Method used

A perception task model is adopted, which includes a BEV feature extraction network and a task prediction head. Through an encoding module and multiple output heads, the BEV features are shared for encoding processing to output a variety of perception results, supporting multi-task processing.

Benefits of technology

This maximizes the utilization of vehicle-side computing resources, reduces the need to deploy models separately for each perception task, and improves the utilization rate of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859232A_ABST
    Figure CN121859232A_ABST
Patent Text Reader

Abstract

The invention provides a multi-task processing method and device, equipment and a medium, and the method comprises the steps: obtaining sensor data, and inputting the sensor data into a BEV aerial view feature extraction network of a perception task model to obtain BEV features; wherein the sensing task model comprises a BEV feature extraction network and a task prediction head, the BEV feature extraction network is used for extracting BEV features of the sensor data, the task prediction head comprises an encoding module and a plurality of output heads, the encoding module is used for encoding the BEV features to obtain encoding features, and the output heads are used for outputting the encoding features to the sensing task model. Each output head is used for processing the coding features to output respective sensing results; and inputting the BEV features into the task prediction head to output a plurality of perception results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of driver assistance technology, and in particular to multitasking methods, devices, equipment and media. Background Technology

[0002] In advanced driver assistance systems (ADAS), the BEV (Bird's Eye View) perception task typically involves converting sensor data into BEV features from the BEV's perspective using a perception task model. Based on these features, perception results are generated to provide decision-making information for decision-making, planning, and control in ADAS. For example, the decision-making stage can use the perception results to determine the vehicle's status, such as needing to slow down if there's a pedestrian ahead; the planning stage can use the perception results to generate a path, such as avoiding obstacles; and the control stage can use the perception results to adjust the vehicle's state, such as maintaining a safe distance.

[0003] The relevant solutions generally involve designing a separate perception task model for each perception task and deploying all the perception task models on the vehicle side.

[0004] However, the more perception task models there are, the more computing resources are required, which leads to a waste of computing resources on the vehicle side by designing separate perception task models. Summary of the Invention

[0005] To overcome the problems existing in related technologies, this specification provides multitasking methods, apparatus, devices and media.

[0006] According to a first aspect of the embodiments of this specification, a multitasking method is provided, the method comprising: Sensor data is acquired and input into the BEV (Bird's Eye View) feature extraction network of the perception task model to obtain BEV features. The perception task model includes a BEV feature extraction network and a task prediction head. The BEV feature extraction network is used to extract BEV features from the sensor data. The task prediction head includes an encoding module and multiple output heads. The encoding module is used to encode the BEV features to obtain encoded features. Each output head is used to process the encoded features to output its own perception result. The BEV features are input into the task prediction head to output various perception results.

[0007] According to a second aspect of the embodiments of this specification, a multitasking apparatus is provided, the apparatus comprising: A BEV feature extraction unit is used to acquire sensor data and input the sensor data into the BEV bird's-eye view feature extraction network of the perception task model to obtain BEV features. The perception task model includes a BEV feature extraction network and a task prediction head. The BEV feature extraction network is used to extract BEV features from the sensor data. The task prediction head includes an encoding module and multiple output heads. The encoding module is used to encode the BEV features to obtain encoded features. Each output head is used to process the encoded features to output its respective perception result. The perception result output unit is used to input the BEV features into the task prediction head to output various perception results.

[0008] According to a third aspect of the embodiments of this specification, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described in the first aspect.

[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method as described in the first aspect.

[0010] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: In the embodiments described in this specification, the task prediction head of the improved perception task model includes an encoding module and multiple output heads. The encoding module can be used to encode the shared BEV features of various perception tasks output by the BEV feature extraction network to obtain encoded features. Each output head can process the encoded features to output its own perception result. When performing a perception task, sensor data can be input into this perception task model to output various perception results.

[0011] As can be seen, this solution can output the perception results required by multiple perception tasks based on a single perception task model, without the need to deploy a separate perception task model for each perception task, thus maximizing the utilization of computing resources on the vehicle side.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.

[0014] Figure 1This is a flowchart illustrating a training method for a perception task model according to an exemplary embodiment of this specification.

[0015] Figure 2 This is a schematic diagram illustrating a perception task model according to an exemplary embodiment of this specification.

[0016] Figure 3 This is a schematic diagram illustrating a task prediction head according to an exemplary embodiment of this specification.

[0017] Figure 4 This is a schematic diagram illustrating another task prediction head according to an exemplary embodiment of this specification.

[0018] Figure 5 This is a schematic diagram of a BEV feature extraction network illustrated in this specification according to an exemplary embodiment.

[0019] Figure 6 This is a schematic diagram illustrating another perception task model according to an exemplary embodiment of this specification.

[0020] Figure 7 This is a flowchart illustrating a multitasking method according to an exemplary embodiment of this specification.

[0021] Figure 8 This specification illustrates an architecture diagram of a deployment-aware task model based on an exemplary embodiment.

[0022] Figure 9 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of this specification.

[0023] Figure 10 This is a block diagram illustrating a multitasking device according to an exemplary embodiment of this specification. Detailed Implementation

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0025] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0027] In advanced driver assistance systems (ADAS), the BEV (Bird's Eye View) perception task typically involves converting sensor data into BEV features from the BEV's perspective using a perception task model. Based on these features, perception results are generated to provide decision-making information for decision-making, planning, and control in ADAS. For example, the decision-making stage can use the perception results to determine the vehicle's status, such as needing to slow down if there's a pedestrian ahead; the planning stage can use the perception results to generate a path, such as avoiding obstacles; and the control stage can use the perception results to adjust the vehicle's state, such as maintaining a safe distance.

[0028] The relevant solutions generally involve designing a separate perception task model for each perception task and deploying all the perception task models on the vehicle side.

[0029] However, the more perception task models there are, the more computing resources are required, which leads to a waste of computing resources on the vehicle side by designing separate perception task models.

[0030] To address the aforementioned technical issues, this specification provides a training method for a perception task model, which trains a perception task model that enables the simultaneous execution of multiple perception tasks based on the model, thereby maximizing the utilization of computing resources on the vehicle side.

[0031] The embodiments described in this specification will now be described in detail.

[0032] Figure 1 This is a flowchart illustrating a training method for a perception task model according to an exemplary embodiment of this specification. It includes steps 101-102: Step 101: Acquire sensor sample data and input the sensor sample data into the perception task model; wherein, the perception task model includes a BEV bird's-eye view feature extraction network and a task prediction head, the BEV feature extraction network is used to extract BEV features from the sensor data, the task prediction head includes an encoding module and multiple output heads, the encoding module is used to encode the BEV features to obtain encoded features, and each output head is used to process the encoded features to output its own perception result.

[0033] Step 102: Update the parameters of the perception task model based on the loss between the perception results of each sample output by the perception task model and the corresponding ground truth.

[0034] In this embodiment, sensor data can be image data and point cloud data, etc. Image data can be image data acquired by various vehicle-mounted image acquisition devices, such as data acquired by vehicle-mounted cameras. Point cloud data can be data acquired by various types of radar, such as millimeter-wave radar, lidar, and ultrasonic radar. The perception task model in this solution can input at least one type of sensor data. For example, it can input only image data or point cloud data, or it can input both image data and point cloud data into the perception task model.

[0035] like Figure 2 As shown, the perception task model 2 in this scheme consists of two parts: a BEV feature extraction network 21 and a task prediction head 22. The BEV feature extraction network 21 can draw on existing network architectures, such as BEVformer and LSS (Lift-Splat-Shoot). Alternatively, the network architecture of the BEV feature extraction network 21 designed in this scheme can also be used, which will not be elaborated upon here.

[0036] The task prediction head 22 may include an encoding module and multiple output heads. The encoding module may be an encoder consisting of convolutional layers and fully connected layers, or an existing Transformer may be used as an encoder. It may transform the BEV features output from the BEV feature extraction network 21 into task-related features relevant to the output heads.

[0037] Each output head processes the encoded features output by the encoding module to produce its own perception results. These perception results can be used in downstream tasks such as decision-making, planning, and control for assisted driving. For example, these output heads can output perception results for OD (Object Detection) tasks, OCC (Occupancy Prediction) tasks, and MAP (Mount Assistance) tasks.

[0038] In this embodiment, the improved perception task model's task prediction head includes an encoding module and multiple output heads. The encoding module encodes the shared BEV features output by the BEV feature extraction network for each perception task to obtain encoded features. Each output head processes the encoded features to output its own perception result. When executing a perception task, sensor data can be input into this perception task model to output multiple perception results. Therefore, this solution can output the required perception results for multiple perception tasks based on a single perception task model, eliminating the need to deploy separate perception task models for each task, thus maximizing the utilization of vehicle-side computing resources.

[0039] The following section introduces two preferred network architectures for the task prediction head 22 designed in this scheme: In one embodiment, such as Figure 3 As shown, the task prediction head 22 can contain the same number of independent encoding modules and output heads. Each independent encoding module is used to encode the BEV features output by the BEV feature extraction network 21 to obtain its own encoded features. Each output head is used to process the encoded features output by the specified encoding module to output its own perception result.

[0040] For example, each encoding module uniquely corresponds to one output head. For instance, encoding module 221A may only input the output encoded features to output head 222A; encoding module 221B may only input the output encoded features to output head 222B; and encoding module 221C may only input the output encoded features to output head 222C.

[0041] For example, each encoding module can correspond to multiple output heads and input the encoded features into its corresponding multiple output heads. Each encoding module can correspond to all or some of the output heads, and the multiple output heads corresponding to each encoding module can not be completely identical. For example, encoding module 221A can correspond to output heads 222A, 222B, and 222C, and simultaneously input the output encoded features into output heads 222A, 222B, and 222C. As another example, encoding module 221B can correspond to output heads 222A and 222B, and simultaneously input the output encoded features into output heads 222A and 222B. In this embodiment, the perceptual task model designed in this scheme integrates multi-task outputs. Based on mutual influence during training, different output heads share the encoded features output by the encoding modules corresponding to other output heads, further improving cross-task feature synergy, thereby improving task accuracy.

[0042] In one embodiment, such as Figure 4As shown, the task prediction head 22 may include a shared encoding module 221 and multiple output heads. The encoding module 221 is used to encode the BEV features output by the BEV feature extraction network 21 to obtain encoded features. Each output head is used to process the same encoded feature to output its own perception result.

[0043] In this embodiment, multiple output heads share the same encoding module, and the encoded features only need to be calculated once, which can avoid repeated calculations and thus save computing resources on the vehicle side.

[0044] Next, we will introduce the preferred network architecture of the BEV feature extraction network 21 designed in this scheme: Multimodal data can combine data from different sensors to provide a more comprehensive environmental perception. The perception task model 2 in this scheme can also support the processing of multimodal data. For example... Figure 5 As shown, the BEV feature extraction network 21 may include a fusion module 212 and feature extraction modules corresponding to the various input sensor data. Each feature extraction module extracts BEV features from its corresponding sensor data, and the fusion module 212 fuses all extracted BEV features to obtain a fused BEV feature. For example, sensor data A, sensor data B, and sensor data C are input to feature extraction modules 211A, 211B, and 211C, respectively. Feature extraction module 211A extracts BEV features from sensor data A; feature extraction module 211B extracts BEV features from sensor data B; and feature extraction module 211C extracts BEV features from sensor data C. All extracted BEV features are then fused in the fusion module 212 to obtain the fused BEV feature. The fusion module 212 can dynamically adjust the weight ratio of different sensor data during the fusion process, ensuring system operation even if one input is lost. For example, if sensor data A cannot be obtained or the confidence level of the obtained sensor data A is low, the weight of sensor data A can be adjusted to 0 or its weight can be reduced.

[0045] In one embodiment, if the sensor data includes point cloud data, the feature extraction module corresponding to the point cloud data may include a voxelization module and a sparse convolution module. The voxelization module is used to extract voxel features from the point cloud data, and the sparse convolution module is used to process the voxel features to obtain point cloud BEV features.

[0046] To address the issues of high resource consumption and long processing time in traditional point cloud data processing, this embodiment designs a voxelization module and a sparse convolution module for sequential joint processing to obtain point cloud BEV features. The voxelization module aggregates features of point clouds within the same 3D mesh, and then the sparse convolution module uses sparse convolution operations to quickly calculate the 3D mesh containing point cloud features in the vicinity, thereby rapidly obtaining the point cloud BEV features. The sparse convolution module can perform calculations only on non-empty voxels, solving the problem of invalid computation of dense meshes output by the voxelization module, thus saving computational resources for processing point cloud data and improving computational efficiency.

[0047] like Figure 6 As shown, this specification provides a preferred model architecture for a perception task. Sensor data may include image data and point cloud data. The feature extraction module corresponding to the image data may include an image backbone network, an FPN network (Feature Pyramid Network), and a view transformation module. The feature extraction module corresponding to the point cloud data may include a voxelization module and a sparse convolution module.

[0048] The image backbone network can employ a multi-layer convolutional neural network. Together with the FPN network, it is used to extract features from image data to obtain semantically meaningful 2D features. The view transformation module is used to convert the features of the source view into the feature representation of the target view. In this scheme, it can be used to upscale 2D features to 3D features, which, after pooling, yield the image's BEV features.

[0049] The point cloud BEV features output by the sparse convolution module and the image BEV features output by the view transformation module can be jointly input into the fusion module for fusion processing to obtain fused BEV features. These fused BEV features are then input into the task prediction head 22 for further processing. For an embodiment of the task prediction head 22, please refer to the foregoing. Figure 3 and Figure 4 This will not be elaborated upon here.

[0050] Figure 7 This is a flowchart illustrating a multitasking method according to an exemplary embodiment of this specification. It includes steps 701-702: Step 701: Acquire sensor data and input the sensor data into the BEV feature extraction network of the perception task model to obtain BEV features.

[0051] Step 702: Input the BEV features into the task prediction head to output various perception results.

[0052] In this embodiment, the solution can output the perception results required by multiple perception tasks based on a single perception task model, without the need to deploy a separate perception task model for each type of perception task, thereby maximizing the utilization of computing resources on the vehicle side.

[0053] When deploying a perception task model, it can be deployed as a whole, or it can be split into different parts and deployed separately, with parallel computation performed on each part.

[0054] In one embodiment, when there are multiple types of input sensor data, the BEV feature extraction network can be divided into a backbone and front-end parts corresponding to each of the multiple types of sensor data, which are deployed separately.

[0055] By leveraging the multi-threading capabilities and the parallel computing power of GPUs, each type of sensor data can be input into its respective front-end component for processing to obtain individual front-end output data. All front-end output data can then be input into the main component for further processing to obtain fused BEV features.

[0056] In this embodiment, by separating the front-end part of the BEV feature extraction network that can be processed in parallel and deploying it separately, the advantages of multi-process can be fully utilized to transform the serial inference process of the model into parallelization, improve inference efficiency, and thus respond to perception results faster.

[0057] In one embodiment, the sensor data may include image data and point cloud data. The preprocessing portion corresponding to the image data can be used to preprocess the image data to obtain preprocessed image data. The preprocessing portion corresponding to the point cloud data can be used to extract features from the point cloud data to obtain point cloud BEV features. The backbone portion is used to extract features from the preprocessed image data to obtain image BEV features, and to fuse the image BEV features and the point cloud BEV features to obtain fused BEV features.

[0058] In one embodiment, the point cloud preprocessing section may include a voxelization module and a sparse convolution module. By deploying time-consuming modules such as the voxelization module and the sparse convolution module separately, the advantages of parallelization can be fully utilized to improve processing efficiency.

[0059] like Figure 8As shown, when the BEV feature extraction network is deployed, the intrinsic and extrinsic parameter preprocessing module can be assigned to the pre-processing part 81 corresponding to the image data, and the voxelization module and sparse convolution module can be assigned to the pre-processing part 82 corresponding to the point cloud data for independent deployment. The remaining parts are centrally deployed in the backbone part 83. After obtaining the sensor intrinsic and extrinsic parameters, the relevant parameters can be cached for reuse in each model inference stage. The sensor outputs corresponding to the image data and point cloud data are timestamped to ensure that different modal inputs are in the same time domain. The point cloud data is processed separately in the pre-processing part 82 to obtain the point cloud BEV features. The point cloud BEV features and image data are then input together into the backbone part 83 for processing to obtain the fused BEV features.

[0060] In one embodiment, during actual deployment, the BEV features output by the BEV feature extraction network can be parallelized and input into each independent encoding module for processing, or the encoded features output by a shared encoding module can be parallelized and input into each output head for processing. Parallel processing of each perception task in the task prediction head can improve processing efficiency.

[0061] Corresponding to the embodiments of the foregoing methods, this specification also provides embodiments of the apparatus and the terminal to which it is applied.

[0062] Figure 9 This is a schematic diagram illustrating the structure of an electronic device according to an exemplary embodiment. Figure 9 As shown, at the hardware level, the electronic device 900 includes a processor 902, an internal bus 904, a network interface 906, a memory 908, and a non-volatile memory 910. It may also include other hardware required for various services. One or more embodiments of this specification can be implemented in software, for example, the processor 902 can read the corresponding computer program from the non-volatile memory 910 into the memory 908 and then run it. Of course, besides software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software. That is to say, the execution entity of the following processing flow is not limited to individual logic modules, but can also be hardware or logic devices.

[0063] Figure 10 This is a block diagram illustrating a multitasking device according to an exemplary embodiment of this specification. Figure 10 As shown, this device can be applied to, for example Figure 9 The electronic device 900 shown implements the technical solution of this specification. The device includes: BEV feature extraction unit 1002 is used to acquire sensor data and input the sensor data into the BEV bird's-eye view feature extraction network of the perception task model to obtain BEV features; wherein, the perception task model includes a BEV feature extraction network and a task prediction head, the BEV feature extraction network is used to extract BEV features from the sensor data, the task prediction head includes an encoding module and multiple output heads, the encoding module is used to encode the BEV features to obtain encoded features, and each output head is used to process the encoded features to output its own perception result.

[0064] The perception result output unit 1004 is used to input the BEV features into the task prediction head to output various perception results.

[0065] Optionally, the sensor data can be of multiple types, and the BEV feature extraction network includes a fusion module and feature extraction modules corresponding to each of the multiple sensor data types. Each feature extraction module is used to extract BEV features from the corresponding sensor data, and the fusion module is used to fuse all the extracted BEV features to obtain fused BEV features.

[0066] Optionally, the sensor data includes point cloud data, and the feature extraction module corresponding to the point cloud data includes a voxelization module and a sparse convolution module; the voxelization module is used to extract voxel features from the point cloud data; the sparse convolution module is used to process the voxel features to obtain point cloud BEV features.

[0067] Optionally, the number of sensor data points can be multiple, and the BEV feature extraction network is divided into a backbone and pre-processing parts corresponding to each of the multiple sensor data points, which are deployed separately. Specifically, the BEV feature extraction unit 1102 is used to parallelize the input of each type of sensor data point to its corresponding pre-processing part for processing to obtain individual pre-processing output data, and then input all the pre-processing output data into the backbone for processing to obtain fused BEV features.

[0068] Optionally, the sensor data includes image data and point cloud data. The preprocessing portion corresponding to the image data is used to preprocess the image data to obtain preprocessed image data; the preprocessing portion corresponding to the point cloud data is used to extract features from the point cloud data to obtain point cloud BEV features; the backbone portion is used to extract features from the preprocessed image data to obtain image BEV features, and to fuse the image BEV features and the point cloud BEV features to obtain fused BEV features.

[0069] Optionally, the front-end portion corresponding to the point cloud data includes a voxelization module and a sparse convolution module; wherein, the voxelization module is used to extract voxel features from the point cloud data, and the sparse convolution module is used to process the voxel features to obtain point cloud BEV features.

[0070] Optionally, the task prediction head includes an encoding module and multiple output heads. The encoding module is used to encode the BEV features to obtain encoded features, and each output head is used to process the encoded features to output its respective perception result, including: The task prediction head includes a shared encoding module and multiple output heads. The shared encoding module is used to encode the BEV features to obtain encoded features, and each output head is used to process the encoded features to output its own perception result. Alternatively, the task prediction head includes the same number of independent encoding modules and output heads. Each independent encoding module is used to encode the BEV features to obtain its own encoded features, and each output head is used to process the encoded features output by a specified encoding module to output its own perception result.

[0071] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0072] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0073] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned multitasking methods provided in this application.

[0074] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0075] This specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the aforementioned multitasking methods.

Claims

1. A multi-task processing method, characterized in that, The method includes: Sensor data is acquired and input into the BEV (Bird's Eye View) feature extraction network of the perception task model to obtain BEV features. The perception task model includes a BEV feature extraction network and a task prediction head. The BEV feature extraction network is used to extract BEV features from the sensor data. The task prediction head includes an encoding module and multiple output heads. The encoding module is used to encode the BEV features to obtain encoded features. Each output head is used to process the encoded features to output its own perception result. The BEV features are input into the task prediction head to output various perception results.

2. The method according to claim 1, characterized in that, The sensor data is of various types, and the BEV feature extraction network includes a fusion module and feature extraction modules corresponding to each of the various sensor data. Each feature extraction module is used to extract BEV features from the corresponding sensor data, and the fusion module is used to fuse all the extracted BEV features to obtain fused BEV features.

3. The method according to claim 2, characterized in that, The sensor data includes point cloud data, and the feature extraction module corresponding to the point cloud data includes a voxelization module and a sparse convolution module. The voxelization module is used to extract voxel features from the point cloud data; the sparse convolution module is used to process the voxel features to obtain point cloud BEV features.

4. The method according to claim 1, characterized in that, The sensor data is diverse, and the BEV feature extraction network is divided into a backbone and front-end components corresponding to each of the various sensor data types, which are deployed separately. The step of inputting the sensor data into the BEV bird's-eye view feature extraction network of the perception task model to obtain BEV features includes: Parallelization inputs each type of sensor data into its corresponding front-end part for processing to obtain each front-end output data, and then inputs all the front-end output data into the main part for processing to obtain fused BEV features.

5. The method according to claim 4, characterized in that, The sensor data includes image data and point cloud data. The preprocessing portion corresponding to the image data is used to preprocess the image data to obtain preprocessed image data; The preceding part corresponding to the point cloud data is used to extract features from the point cloud data to obtain point cloud BEV features. The main part is used to extract features from the preprocessed image data to obtain image BEV features, and to fuse the image BEV features and the point cloud BEV features to obtain fused BEV features.

6. The method according to claim 5, characterized in that, The front-end portion corresponding to the point cloud data includes a voxelization module and a sparse convolution module; wherein, the voxelization module is used to extract voxel features from the point cloud data, and the sparse convolution module is used to process the voxel features to obtain point cloud BEV features.

7. The method according to any one of claims 1-6, characterized in that, The task prediction head includes an encoding module and multiple output heads. The encoding module is used to encode the BEV features to obtain encoded features. Each output head is used to process the encoded features to output its respective perception result, including: The task prediction head includes a shared encoding module and multiple output heads. The shared encoding module is used to encode the BEV features to obtain encoded features, and each output head is used to process the encoded features to output its own perception result. or, The task prediction head contains the same number of independent encoding modules and output heads. Each independent encoding module is used to encode the BEV features to obtain its own encoded features. Each output head is used to process the encoded features output by the specified encoding module to output its own perception result.

8. A multitasking processing device, characterized in that, The device includes: A BEV feature extraction unit is used to acquire sensor data and input the sensor data into the BEV bird's-eye view feature extraction network of the perception task model to obtain BEV features. The perception task model includes a BEV feature extraction network and a task prediction head. The BEV feature extraction network is used to extract BEV features from the sensor data. The task prediction head includes an encoding module and multiple output heads. The encoding module is used to encode the BEV features to obtain encoded features. Each output head is used to process the encoded features to output its respective perception result. The perception result output unit is used to input the BEV features into the task prediction head to output various perception results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.