Multi-task processing method and device, electronic equipment and readable storage medium

By designing a multi-task processing model in an autonomous driving system, and using the physical perception range of the task to output the feature map of the corresponding task, the problem of low reliability of task processing results in the prior art is solved, and more efficient multi-task processing and correlation are achieved.

CN120163987APending Publication Date: 2025-06-17HAOMO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311719019.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, the reliability of task processing results such as lane line detection and obstacle detection is low, and single-task optimization is difficult to ensure the correlation between tasks.

Method used

By designing a multitasking processing model, the model includes multiple branches that correspond to multiple tasks one by one. Each branch outputs the first feature map of the corresponding task in the BEV perspective according to the physical perception range of the task, and adapts the feature extraction requirements of different tasks through pixel mapping relationships.

Benefits of technology

It improves the reliability of multitasking, ensures the correlation between different tasks and the accuracy of feature extraction, and improves the effects of lane line detection and obstacle detection in autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163987A_ABST
    Figure CN120163987A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computer models, and provides a multi-task processing method and device, electronic equipment and a readable storage medium. The multi-task processing method comprises the following steps: acquiring a to-be-processed image; the to-be-processed image is input into a multi-task processing model, the multi-task processing model comprises a plurality of branches corresponding to a plurality of tasks in a one-to-one mode, and each branch is used for outputting a first feature map of the corresponding task at the BEV view angle according to the physical sensing range of the corresponding task, different physical sensing ranges correspond to different pixel mapping relationships, and the pixel mapping relationships represent the position and size of each pixel on the first feature map in a corresponding pixel region on the to-be-processed image; and determining a multi-task processing result according to the first feature maps which are output by the multi-task processing model and are in one-to-one correspondence with the plurality of tasks. According to the embodiment of the invention, the reliability of multi-task processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of computer models, and particularly relates to a multi-task processing method, apparatus, electronic device, and readable storage medium. Background Art

[0002] In recent years, with the development of technology, the intelligence and automation of automobiles have been increasingly favored by the academic and industrial communities. Autonomous driving technology is an important development direction in the current and future automotive industries, and visual perception is one of the core technologies for realizing autonomous driving. Lane line detection and obstacle detection, as basic perception tasks, have always attracted much attention in autonomous driving technology. Existing lane line detection and obstacle detection almost all exist in the form of single tasks, and it is difficult to ensure the relevance between different tasks when each task is optimized independently. Some solutions apply images from the same Bird's Eye View (BEV) perspective to different tasks (such as lane line detection and obstacle detection), but it is found in practical applications that the obtained BEV perspective images cannot give reliable task processing results in some tasks. Summary of the Invention

[0003] Embodiments of this application provide a multi-task processing method, apparatus, electronic device, and readable storage medium, which can solve the problem of low reliability of task processing results in related technologies.

[0004] In a first aspect of the embodiments of this application, a multi-task processing method is provided, including: obtaining an image to be processed; inputting the image to be processed into a multi-task processing model, where the multi-task processing model includes multiple branches corresponding to multiple tasks one by one, and each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task, different physical perception ranges correspond to different pixel mapping relationships, and the pixel mapping relationship represents the position and size of the pixel area corresponding to each pixel on the first feature map in the image to be processed; determining a multi-task processing result according to the first feature maps corresponding to the multiple tasks output by the multi-task processing model.

[0005] In some embodiments of the first aspect, the image to be processed includes perception images of multiple perspectives of a vehicle; the outputting a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task includes: determining second feature maps corresponding to the multiple perspectives in the BEV perspective according to the physical perception range of the corresponding task and the perception images of each perspective; determining the first feature map of the corresponding task according to the second feature maps.

[0006] In some embodiments of the first aspect, before outputting the first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task, it includes: for each of the multiple perspectives, processing the perception image into multiple preprocessed images with different resolutions; determining the second feature maps corresponding to the multiple perspectives in the BEV perspective according to the physical perception range of the corresponding task and the perception image of each perspective, including: for each of the multiple perspectives, performing feature extraction on the multiple preprocessed images respectively according to the physical perception range of the corresponding task, to obtain the multiple second feature maps corresponding to the multiple preprocessed images in the BEV perspective.

[0007] In some embodiments of the first aspect, determining the first feature map of the corresponding task according to the second feature map includes: for the same perspective, fusing the multiple second feature maps corresponding to the multiple preprocessed images to obtain a third feature map; fusing the third feature maps corresponding to different perspectives to obtain the first feature map of the corresponding task.

[0008] In some embodiments of the first aspect, performing feature extraction on the multiple preprocessed images respectively according to the physical perception range of the corresponding task, to obtain the multiple second feature maps corresponding to the multiple preprocessed images in the BEV perspective, includes: performing fully connected processing on the multiple preprocessed images respectively to obtain multiple global feature maps corresponding to the multiple preprocessed images; fusing each global feature map with the corresponding preprocessed image to obtain multiple fused feature maps corresponding to the multiple preprocessed images; converting the fused feature maps to the BEV perspective to obtain the multiple second feature maps corresponding to the multiple preprocessed images.

[0009] In some embodiments of the first aspect, the first feature map corresponding to a single task includes sub-feature maps of multiple feature channels; determining the multi-task processing result according to the first feature maps corresponding to the multiple tasks output by the multi-task processing model includes: for each of the multiple tasks, based on the corresponding relationship between each task operation of the task and the feature channels, performing the corresponding task operation respectively based on the sub-feature maps of different feature channels to obtain the task processing result of the task; generating the multi-task processing result based on the task processing results of each task in the multiple tasks.

[0010] In some embodiments of the first aspect, the multiple tasks include an obstacle detection task and a lane line recognition task, the physical perception range of the obstacle detection task is within a first preset distance around the vehicle, and the physical perception range of the lane line recognition task is beyond a second preset distance in front of the vehicle.

[0011] A multitasking processing device provided in the second aspect of the embodiments of the present application includes: an image acquisition unit for acquiring an image to be processed; a model processing unit for inputting the image to be processed into a multitasking processing model, where the multitasking processing model includes a plurality of branches corresponding one-to-one to a plurality of tasks, and each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task. Different physical perception ranges correspond to different pixel mapping relationships, and the pixel mapping relationship represents the position and size of the pixel area corresponding to each pixel on the first feature map on the image to be processed; a result determination unit for determining a multitasking processing result according to the first feature maps corresponding one-to-one to the plurality of tasks output by the multitasking processing model.

[0012] An electronic device provided in the third aspect of the embodiments of the present application includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above multitasking processing method are implemented.

[0013] A computer-readable storage medium provided in the fourth aspect of the embodiments of the present application stores a computer program, and when the computer program is executed by a processor, the steps of the above multitasking processing method are implemented.

[0014] A computer program product provided in the fifth aspect of the embodiments of the present application causes an electronic device to execute the above multitasking processing method when the computer program product runs on the electronic device.

[0015] In the embodiments of the present application, by acquiring an image to be processed and inputting the image to be processed into a multitasking processing model, a multitasking processing result is determined according to the first feature maps corresponding one-to-one to a plurality of tasks output by the multitasking processing model. Among them, the multitasking processing model includes a plurality of branches corresponding one-to-one to a plurality of tasks, and each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task. Different physical perception ranges correspond to different pixel mapping relationships. Since the pixel mapping relationship represents the position and size of the pixel area corresponding to each pixel on the first feature map on the image to be processed, different feature maps can be extracted according to the differences in the physical perception ranges between different tasks to adapt to the feature extraction requirements of multiple tasks, which helps to improve the reliability of multitasking processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of the implementation of a multi-task processing method provided by an embodiment of the present application;

[0018] Figure 2 It is a schematic flowchart of the specific implementation of step S102 provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic structural diagram of a multi-task processing device provided by an embodiment of the present application;

[0020] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific Embodiments

[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0022] For an autonomous driving system, images from the same Bird's Eye View (BEV) perspective are often applied to different tasks. However, in actual applications, it is found that the obtained BEV perspective images cannot provide reliable task processing results in some tasks.

[0023] Based on this, the present application proposes a multi-task processing method that can adapt to the physical perception ranges of different tasks and output corresponding first feature maps according to the needs of different tasks to improve the reliability of multi-task processing.

[0024] To illustrate the technical solutions of the present application, the following will be described through specific embodiments.

[0025] Figure 1 It shows a schematic flowchart of the implementation of a multi-task processing method provided by an embodiment of the present application. This method can be applied to a processor and is applicable to situations where the reliability of multi-task processing needs to be improved.

[0026] Among them, the above-mentioned processor can be integrated in electronic devices such as smart phones and computers, or can be integrated in vehicles, and the present application does not limit this.

[0027] Specifically, the above multi-task processing method may include the following steps S101 to S103.

[0028] Step S101, obtain the image to be processed.

[0029] Among them, the image to be processed is an image for multi-task processing. In a specific scenario of this application, it may refer to a perception image collected by an in-vehicle camera installed on a vehicle. It should be understood that the image to be processed can be obtained in real time, pre-stored in a memory, or input by a user, and this application does not limit this.

[0030] Step S102, input the image to be processed into the multi-task processing model.

[0031] Among them, the multi-task processing model is used to perform feature extraction according to the task requirements of multiple tasks.

[0032] Specifically, the multi-task processing model may include multiple branches corresponding one-to-one to multiple tasks, and each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task.

[0033] Among them, the physical perception range refers to the region of interest of the task in physical space. In an implementation manner of this application, different physical perception ranges correspond to different pixel mapping relationships, and the pixel mapping relationship represents the position and size of the pixel region corresponding to each pixel on the first feature map in the image to be processed. In other words, if the physical perception range of the task is different, the position and size of the pixel region corresponding to the pixel at the same position on the first feature map in the image to be processed are different, and thus the first feature maps of different tasks in the BEV perspective are also different from each other.

[0034] Specifically, for the pixel region of the image content in the image to be processed corresponding to the physical perception range, there may be more pixels corresponding in the first feature map. Or rather, on the first feature map, for the pixels of the image region mapped to the physical perception range, each pixel can map a smaller pixel region within this image region, so there are more pixels for mapping to the image region corresponding to the physical perception range. For the pixels of the image region mapped to outside the physical perception range, each pixel can map a larger pixel region within this image region, so there are fewer pixels for mapping to the image region corresponding to outside the physical perception range.

[0035] In some specific examples, the multiple tasks may include an obstacle detection task and a lane line recognition task.

[0036] Since the purpose of obstacle detection is usually to facilitate vehicle obstacle avoidance, the physical perception range of the obstacle detection task can be within a first preset distance around the vehicle. Correspondingly, the pixel region within the first preset distance around the vehicle in the image to be processed can correspond to more pixels in the first feature map, while the pixel region outside the first preset distance around the vehicle can correspond to fewer pixels in the first feature map, so that the first feature map corresponding to the obstacle detection task can describe the situation of obstacles within the first preset distance around the vehicle more, facilitating the output of a more reliable obstacle detection result and contributing to the current obstacle avoidance of the vehicle.

[0037] Since the purpose of lane line recognition is usually to facilitate vehicle lane change, the physical perception range of the lane line recognition task can be outside a second preset distance in front of the vehicle. Correspondingly, the pixel region outside the second preset distance in front of the vehicle in the image to be processed can correspond to more pixels in the first feature map, while the pixel region within the second preset distance in front of the vehicle and the pixel regions on other sides (rear side, left and right sides, etc.) can correspond to fewer pixels in the first feature map, so that the first feature map corresponding to the lane line recognition task can describe the situation of lane lines outside the second preset distance in front of the vehicle more, facilitating the output of a more reliable lane line recognition result and contributing to the future lane change of the vehicle.

[0038] Step S103: Determine the multi-task processing result according to the first feature maps corresponding to multiple tasks output by the multi-task processing model.

[0039] In the embodiments of the present application, each branch of the multi-task processing model can output a first feature map corresponding to a task. Based on the first feature maps corresponding to multiple tasks one by one, the multi-task processing result can be fused.

[0040] In the embodiments of the present application, by obtaining the image to be processed and inputting the image to be processed into the multi-task processing model to determine the multi-task processing result according to the first feature maps corresponding to multiple tasks output by the multi-task processing model, wherein the multi-task processing model includes multiple branches corresponding to multiple tasks one by one, and each branch is used to output the first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task. Different physical perception ranges correspond to different pixel mapping relationships. Since the pixel mapping relationship represents the position and size of the pixel region corresponding to each pixel on the first feature map in the image to be processed, different feature maps can be extracted according to the differences in the physical perception ranges between different tasks to adapt to the feature extraction requirements of multiple tasks, which helps to improve the reliability of multi-task processing.

[0041] In some embodiments of the present application, the image to be processed may include perception images of multiple perspectives of a vehicle, specifically including a front view image, a rear view image, a side view image, etc. Considering that the information in the vehicle front view image is usually relatively important, the resolution of the front view image may be greater than the resolution of the perception images of other perspectives.

[0042] As an example, a front view image, a rear view image, a left view image, a right view image, a left rear view image, and a right rear view image collected by an in-vehicle camera may be obtained. Among them, the resolution of the front view image is 1024*576, and the resolution of the other five images is 512*288.

[0043] Of course, the perception images of each perspective can also be preprocessed once. This preprocessing is to select a group within the perspective of each in-vehicle camera as the standard perspective, and perform a projection operation to convert the images with different extrinsic parameters of other relative perspectives to the standard perspective. This preprocessing can solve the model jitter problem caused by differences in the extrinsic parameters of different in-vehicle cameras, and can effectively increase the generalization ability of the model.

[0044] For the branches in the above multi-task processing model, as Figure 2 shown, according to the physical perception range of the corresponding task, the first feature map of the corresponding task on the BEV perspective can be output, which may include steps S201 to S202.

[0045] Step S201, according to the physical perception range of the corresponding task and the perception images of each perspective, determine the second feature map corresponding to each of the multiple perspectives on the BEV perspective.

[0046] That is to say, for each perspective, feature extraction can be performed according to the physical perception range of the task to obtain the second feature map of this perspective converted to the BEV perspective.

[0047] In some embodiments of the present application, for each perspective, the processor may process the image to be processed into multiple preprocessed images with different resolutions.

[0048] Specifically, the images of each camera perspective pass through the backbone network of the Residual Neural Network (ResNet) structure to obtain features at 3 scales. The features at 3 scales within each perspective are fused by Feature Pyramid Networks (FPN) to obtain 3 preprocessed images. The preprocessed images contain richer information than the perception images. Among them, the above 3 scales correspond to different resolutions.

[0049] Taking the 6 perspectives described above as examples, the 3 scales corresponding to the front view image can include 72*128 with 8x downsampling, 36*64 with 16x downsampling, and 18*32 with 32x downsampling. The 3 scales corresponding to the perception images of the remaining 5 perspectives can include 36*64 with 8x downsampling, 18*32 with 16x downsampling, and 9*16 with 32x downsampling. In this way, on the one hand, the parameter quantity can be reduced and the multi-task processing speed can be improved. On the other hand, the information of the front view image can be guaranteed to be retained to a greater extent.

[0050] Moreover, for the preprocessed image, channel dimensionality reduction can also be performed to reduce the data volume.

[0051] Next, for each of the multiple perspectives, the processor can perform feature extraction on the multiple preprocessed images respectively according to the physical perception range of the corresponding task, and obtain multiple second feature maps corresponding one by one to the multiple preprocessed images in the BEV perspective.

[0052] Specifically, the processor can perform full connection (FC) processing on the multiple preprocessed images respectively to obtain multiple global feature maps corresponding one by one to the multiple preprocessed images. Then, each global feature map and the corresponding preprocessed image are fused to obtain multiple fused feature maps corresponding one by one to the multiple preprocessed images. Subsequently, the fused feature maps are converted to the BEV perspective to obtain multiple second feature maps corresponding one by one to the multiple preprocessed images.

[0053] Among them, the global feature map contains the global features of the image to be processed. Fusing the global feature map and the corresponding preprocessed image can fuse the global features and local features, making the fused feature map more referential. The existing methods such as BEVFormer and LSS (Lift Splat Shoot) can be used to convert the fused feature map to the BEV perspective, and this application does not limit this.

[0054] After obtaining the second feature maps, the processor can also uniformly process the resolutions of the respective second feature maps to a preset resolution.

[0055] Step S202, determine the first feature map corresponding to the task according to the second feature map.

[0056] Specifically, for the same perspective, the processor can fuse the multiple second feature maps corresponding one by one to the multiple preprocessed images to obtain a third feature map, and then fuse the third feature maps corresponding to different perspectives to obtain the first feature map corresponding to the task.

[0057] Exemplarily, the specific processing process of a single branch is described below.

[0058] For a single preprocessed image, first, the processor can reduce the number of channels of the preprocessed image through a module of covn2d (convolutional layer) + bn (batch normalization layer) + relu (activation layer), and perform channel attention mechanism on the features through the structure of SE-Net (Squeez and Excitation Networks) to further reduce the amount of data. At this time, the result input = [B, C, H, W] can be obtained, where B represents the batch size, which refers to the number of samples used to update the model parameters in one training; C represents the number of feature channels; H and W represent the height and width of the feature map. For subsequent processing convenience, input = [B, C, H, W] can be converted to input_reshape = [B, C, H*W] through feature reshaping (reshape). Then, perform global attention mechanism on the image space of input_reshape through FC operation, and output the global feature map input_reshape_a = [B, C, H*W]. Fuse the global features and local features in the image space through input_reshape += input_reshape_a. Then, convert the fused feature map input_reshape obtained after fusion to output = [B, C, Ho*Wo] through FC operation, where Ho and Wo are the height and width of the BEV space. Finally, perform feature reshaping on output to obtain the second feature map output_reshape = [B, C, Ho, Wo].

[0059] After performing the above processing on the preprocessed images of 3 scales at 6 viewpoints in sequence, 18 second feature maps with the same resolution in the BEV viewpoint can be obtained. At this time, a concat (connection) operation can be performed on the 3 second feature maps of the same viewpoint to obtain a third feature map, and then the third feature maps corresponding to different viewpoints are fused and a sum (addition) operation is performed to obtain the first feature map output by a single branch (that is, the first feature map corresponding to a single task).

[0060] After obtaining the first feature maps corresponding to different tasks, the processor can determine the multi-task processing result.

[0061] Specifically, the first feature map corresponding to a single task can contain sub-feature maps of multiple feature channels.

[0062] The above step S103 can specifically include: for each task among multiple tasks, based on the corresponding relationship between each task operation of the task and the feature channels, perform the corresponding task operation respectively based on the sub-feature maps of different feature channels to obtain the task processing result of the task. Then, based on the task processing results of each task among multiple tasks, generate a multi-task processing result.

[0063] Taking the lane line recognition task as an example, the lane line recognition task may include an image segmentation step, a clustering step, a lane line color recognition step, a lane line function recognition step, a lane line attribute recognition step, and so on. The first feature map may include sub-feature maps of multiple feature channels. Each task operation may correspond to one or more sub-feature maps in the first feature map. Correspondingly, when performing a task operation, the corresponding task operation may be performed based on the sub-feature maps of different feature channels to obtain the task processing result of the task. For example, in the first feature map, channels 1 and 2 correspond to the image segmentation step. According to the sub-feature maps of channels 1 and 2, a segmentation map of the lane line can be output; channel 3 corresponds to the clustering step. According to the sub-feature map of channel 3, a clustering map of the lane line can be output, and so on.

[0064] Combining the results of these task operations, post-processing can be further performed to obtain task processing results such as the fitting equation of all lane lines in the ego-vehicle coordinate system and the start and end point information in the image to be processed.

[0065] In this way, the processing of multiple different task steps of the same task can be completed based on a single first feature map, which can effectively reduce the number of parameters of the model compared to outputting a feature map for each step.

[0066] Of course, other recognition methods are also applicable to this application. Taking the obstacle detection task as an example: The first feature map can be processed in a manner similar to YOLO (You Only Look Once) to detect the center point, length, width, height, attitude angle, etc. of obstacles within the corresponding physical perception range as the task processing result.

[0067] For the task processing results of different tasks, multi-task processing results can be obtained through fusion. If there is a correlation between different tasks, the task processing result of the second task can be obtained by processing based on the task processing result of one task and the first feature map of the other task, and then the task processing results of different tasks can be fused.

[0068] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences.

[0069] As Figure 3 shown in the structural schematic diagram of a multi-task processing device 300 provided by an embodiment of this application, the multi-task processing device 300 is configured on a processor.

[0070] Specifically, the multi-task processing device 300 may include:

[0071] An image acquisition unit 301 for acquiring an image to be processed;

[0072] A model processing unit 302 for inputting the image to be processed into a multi-task processing model, the multi-task processing model including a plurality of branches corresponding one-to-one to a plurality of tasks, each branch being configured to output a first feature map of the corresponding task in the BEV view according to the physical perception range of the corresponding task, different physical perception ranges corresponding to different pixel mapping relationships, the pixel mapping relationship representing the position and size of the pixel region corresponding to each pixel on the first feature map in the image to be processed;

[0073] A result determination unit 303 for determining a multi-task processing result according to the first feature maps corresponding one-to-one to the plurality of tasks output by the multi-task processing model.

[0074] In some embodiments of the present application, the model processing unit 302 may specifically be configured to: determine second feature maps corresponding one-to-one to the plurality of views in the BEV view according to the physical perception range of the corresponding task and the perception images of each view; determine the first feature map of the corresponding task according to the second feature map.

[0075] In some embodiments of the present application, the model processing unit 302 may specifically be configured to: for each view of the plurality of views, process the perception image into a plurality of preprocessed images with different resolutions; for each view of the plurality of views, perform feature extraction on the plurality of preprocessed images respectively according to the physical perception range of the corresponding task to obtain a plurality of the second feature maps corresponding one-to-one to the plurality of preprocessed images in the BEV view.

[0076] In some embodiments of the present application, the model processing unit 302 may specifically be configured to: for the same view, fuse the plurality of second feature maps corresponding one-to-one to the plurality of preprocessed images to obtain a third feature map; fuse the third feature maps corresponding to different views to obtain the first feature map of the corresponding task.

[0077] In some embodiments of the present application, the model processing unit 302 may specifically be configured to: perform fully connected processing on the plurality of preprocessed images respectively to obtain a plurality of global feature maps corresponding one-to-one to the plurality of preprocessed images; fuse each global feature map and the corresponding preprocessed image to obtain a plurality of fused feature maps corresponding one-to-one to the plurality of preprocessed images; convert the fused feature maps to the BEV view to obtain a plurality of the second feature maps corresponding one-to-one to the plurality of preprocessed images.

[0078] In some embodiments of the present application, the above result determination unit 303 may be specifically configured to: for each task among the multiple tasks, based on the corresponding relationship between each task operation of the task and the feature channels, perform the corresponding task operation on the sub-feature maps of different feature channels respectively to obtain the task processing result of the task; and generate a multi-task processing result based on the task processing results of each task among the multiple tasks.

[0079] In some embodiments of the present application, the multiple tasks include an obstacle detection task and a lane line recognition task. The physical perception range of the obstacle detection task is within a first preset distance around the vehicle, and the physical perception range of the lane line recognition task is outside a second preset distance in front of the vehicle.

[0080] It should be noted that for the convenience and brevity of description, the specific working process of the above multi-task processing device 300 may refer to Figures 1 to 2 the corresponding process of the method, which will not be elaborated here.

[0081] As Figure 4 shown, it is a schematic diagram of an electronic device provided by an embodiment of the present application. Specifically, the electronic device 4 may include: a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40, such as a multi-task processing program. When the processor 40 executes the computer program 42, it implements the steps in the above-mentioned embodiments of various multi-task processing methods, such as Figure 1 the steps S101 to S103 shown. Alternatively, when the processor 40 executes the computer program 42, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 3 the functions of the image acquisition unit 301, the model processing unit 302, and the result determination unit 303 shown.

[0082] The computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 41 and executed by the processor 40 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.

[0083] For example, the computer program can be divided into: an image acquisition unit, a model processing unit, and a result determination unit. The specific functions of each unit are as follows: The image acquisition unit is used to acquire an image to be processed; the model processing unit is used to input the image to be processed into a multi-task processing model, the multi-task processing model includes multiple branches corresponding one-to-one to multiple tasks, and each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task. Different physical perception ranges correspond to different pixel mapping relationships, and the pixel mapping relationship represents the position and size of the pixel region corresponding to each pixel on the first feature map in the image to be processed; the result determination unit is used to determine the multi-task processing result according to the first feature maps corresponding one-to-one to the multiple tasks output by the multi-task processing model.

[0084] The electronic device may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 These are merely examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0085] The so-called processor 40 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0086] The memory 41 may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. The memory 41 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 41 may also include both the internal storage unit and the external storage device of the electronic device. The memory 41 is used to store the computer program and other programs and data required by the electronic device. The memory 41 may also be used to temporarily store the data that has been output or will be output.

[0087] It should be noted that for the convenience and conciseness of description, the structure of the above-mentioned electronic device can also refer to the specific description of the structure in the method embodiments, which will not be elaborated here.

[0088] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0089] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0090] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0091] In the embodiments provided in the present application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are only illustrative. For example, the division of the above-mentioned module or unit is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0092] The unit described as a separation component may or may not be physically separated. The component presented as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0093] In addition, in each embodiment of this application, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0094] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0095] The above-mentioned embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.

Claims

1. A multitasking method, characterized in that, Including: Obtain an image to be processed; Input the image to be processed into a multi-task processing model, where the multi-task processing model includes multiple branches corresponding one-to-one to multiple tasks, and each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task. Different physical perception ranges correspond to different pixel mapping relationships, and the pixel mapping relationship represents the position and size of the pixel region corresponding to each pixel on the first feature map in the image to be processed; Determine the multi-task processing result according to the first feature maps corresponding one-to-one to the multiple tasks output by the multi-task processing model.

2. The multitasking method according to claim 1, characterized in that, The image to be processed includes perception images of multiple perspectives of a vehicle; The step of outputting a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task includes: Determine second feature maps corresponding one-to-one to the multiple perspectives in the BEV perspective according to the physical perception range of the corresponding task and the perception image of each perspective; Determine the first feature map of the corresponding task according to the second feature map.

3. The multitasking method according to claim 2, characterized in that, Before outputting the first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task, it includes: For each of the multiple perspectives, process the perception image into multiple preprocessed images with different resolutions; The step of determining second feature maps corresponding one-to-one to the multiple perspectives in the BEV perspective according to the physical perception range of the corresponding task and the perception image of each perspective includes: For each of the multiple perspectives, perform feature extraction on the multiple preprocessed images respectively according to the physical perception range of the corresponding task to obtain multiple second feature maps corresponding one-to-one to the multiple preprocessed images in the BEV perspective.

4. The multitasking method according to claim 3, characterized in that, The step of determining the first feature map of the corresponding task according to the second feature map includes: For the same perspective, fuse the multiple second feature maps corresponding one-to-one to the multiple preprocessed images to obtain a third feature map; Fuse the third feature maps corresponding to different perspectives to obtain the first feature map of the corresponding task.

5. The multitasking method according to claim 3, characterized in that, The step of performing feature extraction on the multiple preprocessed images respectively according to the physical perception range of the corresponding task to obtain multiple second feature maps corresponding one-to-one to the multiple preprocessed images in the BEV perspective includes: Perform fully connected processing on the multiple preprocessed images respectively to obtain multiple global feature maps corresponding one-to-one to the multiple preprocessed images; Fuse each global feature map with the corresponding preprocessed image to obtain multiple fused feature maps corresponding one-to-one to the multiple preprocessed images; Convert the fused feature map to the BEV perspective to obtain multiple second feature maps corresponding one-to-one to the multiple preprocessed images.

6. The multitasking method according to any one of claims 1 to 5, characterized in that, The first feature map corresponding to a single task contains sub-feature maps of multiple feature channels; The step of determining the multi-task processing result according to the first feature maps corresponding one-to-one to the multiple tasks output by the multi-task processing model includes: For each of the multiple tasks, according to the corresponding relationship between the respective task operations of the task and the feature channels, the corresponding task operations are respectively performed based on the sub-feature maps of different feature channels to obtain the task processing result of the task; Based on the task processing results of each of the multiple tasks, a multi-task processing result is generated.

7. The multitasking method according to any one of claims 1 to 5, characterized in that, The multiple tasks include an obstacle detection task and a lane line recognition task. The physical perception range of the obstacle detection task is within a first preset distance around the vehicle, and the physical perception range of the lane line recognition task is outside a second preset distance in front of the vehicle.

8. A multitasking device, characterized in that, Comprising: An image acquisition unit for acquiring an image to be processed; A model processing unit for inputting the image to be processed into a multi-task processing model. The multi-task processing model includes multiple branches corresponding one-to-one to multiple tasks. Each branch is used to output a first feature map of the corresponding task in the BEV perspective according to the physical perception range of the corresponding task. Different physical perception ranges correspond to different pixel mapping relationships, and the pixel mapping relationship represents the position and size of the pixel region corresponding to each pixel on the first feature map in the image to be processed; A result determination unit for determining a multi-task processing result according to the first feature maps corresponding one-to-one to the multiple tasks output by the multi-task processing model.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the multi-task processing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the multi-task processing method according to any one of claims 1 to 7 are implemented.