Point cloud multitask processing method and device, electronic equipment and computer storage medium

By using pre-trained point cloud multitasking models for feature extraction and task mapping in the field of intelligent driving technology, the problem of point cloud multitasking processing accuracy reduction in different detection ranges is solved, and efficient task processing is achieved.

CN119942311APending Publication Date: 2025-05-06BEIJING VOYAGER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311444422.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to handle point cloud multitasking in different detection ranges at the same time, resulting in a decrease in task processing accuracy.

Method used

By obtaining point cloud data under the corresponding detection range of different tasks, using the pre-trained point cloud multi-task model for feature extraction and task mapping, the task output results of each task are determined. The model is processed through the backbone network and the first network corresponding to the task, and introduces a second network to adjust based on the spatial attention mechanism during the training process.

Benefits of technology

Multitasking of point cloud data under different detection ranges is realized, which improves task processing accuracy and reduces interference from invalid point clouds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942311A_ABST
    Figure CN119942311A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a point cloud multi-task processing method and device, electronic equipment and a computer storage medium, and the method comprises the steps: obtaining point cloud data with task identifiers of corresponding tasks in detection ranges corresponding to different tasks, inputting each point cloud data into a backbone network in a pre-trained point cloud multi-task model for feature extraction, determining an initial feature map, and inputting the initial feature map into a first network corresponding to each task in the point cloud multi-task model, and performing mapping based on the task identifiers corresponding to the feature points in the initial feature map, and determining a task output result of the corresponding task. Therefore, the task processing result of the corresponding task is determined by performing feature extraction on the point cloud data of the different tasks in the different detection ranges and performing mapping based on the task identifier, the tasks in the different detection ranges can be processed at the same time, and the task processing precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to a point cloud multi-task processing method, device, electronic device and computer storage medium. Background Art

[0002] In the field of intelligent driving technology, target detection using point clouds can identify the bounding boxes of obstacles on the road (such as vehicles, pedestrians, two-wheeled vehicles, etc.), and semantic segmentation using point clouds can identify which point clouds belong to a certain type of obstacle.

[0003] However, due to the huge amount of point cloud data, processing point cloud data under different tasks (such as target detection, semantic segmentation, etc.) through different network models to determine the corresponding output results will take up a lot of resources and computing power. At the same time, although the existing point cloud multi-task technology can process different tasks at the same time, it mostly outputs the results of different tasks under the point cloud data of the same detection range, and cannot process multiple tasks with different detection ranges at the same time. In addition, for the same batch of point cloud data, since the detection range corresponding to the point cloud semantic segmentation task in actual applications will be smaller than the target detection range, if point cloud data with a larger detection range is used for multi-task processing at the same time, the accuracy of one or more tasks will be reduced. Summary of the invention

[0004] In view of this, an object of an embodiment of the present invention is to provide a point cloud multi-task processing method to simultaneously process tasks in different detection ranges and improve task processing accuracy.

[0005] In a first aspect, an embodiment of the present invention is to provide a point cloud multi-task processing method, the method comprising:

[0006] Acquire point cloud data within detection ranges corresponding to different tasks, wherein the point cloud data has a task identifier corresponding to the task;

[0007] Input each of the point cloud data into the backbone network of the pre-trained point cloud multi-task model to extract features and determine an initial feature map;

[0008] The initial feature map is input into the first network corresponding to each of the tasks in the point cloud multi-task model, so as to map the task identifiers corresponding to the feature points in the initial feature map to determine the task output results of the corresponding tasks.

[0009] Furthermore, the point cloud multi-task model is trained by the following steps:

[0010] Acquire a training data set, wherein the training data set includes a plurality of samples, wherein the samples include point cloud data under detection ranges corresponding to different tasks, and the point cloud data has a corresponding task identifier;

[0011] Inputting the point cloud data in each of the samples into a preset backbone network for feature extraction, and determining a sample feature map of each of the samples;

[0012] Input each of the sample feature maps into a first network corresponding to each of the tasks and a second network corresponding to at least one of the tasks for mapping, and determine the corresponding task output result and the task expected result; wherein the second network maps the initial feature map based on a spatial attention mechanism;

[0013] The backbone network parameters of the corresponding task branches are adjusted according to the output results of each task and the expected results of the task.

[0014] Furthermore, the second network maps the sample feature map based on a spatial attention mechanism to determine the expected result of the corresponding task, including:

[0015] Obtaining a sample feature map output by the backbone network and a shallow feature map corresponding to the sample feature map;

[0016] The upsampled sample feature map and the shallow feature map are fused to determine a first intermediate feature map;

[0017] Processing the first intermediate feature map based on a preset spatial attention network to determine a weight corresponding to each position point in the first intermediate feature map;

[0018] Performing weighted summation on the upsampled sample feature maps according to the weights to determine a second intermediate feature map;

[0019] Perform feature mapping on the second intermediate feature map to determine the corresponding expected result of the task.

[0020] Furthermore, adjusting the backbone network parameters of the corresponding task branches according to the output results of each task and the expected results of the task includes:

[0021] Determine the difference between the expected result of the task and the task output result of the corresponding task branch;

[0022] The backbone network parameters of the corresponding task branch are adjusted according to the difference.

[0023] Furthermore, the loss function of the point cloud multi-task model is determined according to the loss function of the task branch corresponding to each task.

[0024] Furthermore, obtaining point cloud data within the detection range corresponding to different tasks includes:

[0025] Get the original point cloud collected by the sensor;

[0026] The original point cloud is screened according to the detection ranges corresponding to different tasks to determine the point cloud data corresponding to each task.

[0027] Furthermore, the original point cloud is screened according to the detection ranges corresponding to different tasks to determine the point cloud data corresponding to each task, including:

[0028] Determine the coordinate range of each task under the detection range;

[0029] Determine the data corresponding to the point cloud whose position coordinates in the original point cloud are within the coordinate range as the point cloud data of the corresponding task;

[0030] The position coordinates corresponding to the point cloud in the original point cloud whose position coordinates are outside the coordinate range are reset.

[0031] Furthermore, the method further comprises:

[0032] voxelize each of the point cloud data to determine the corresponding point cloud column;

[0033] Inputting each of the point cloud data into the backbone network of the pre-trained point cloud multi-task model for feature extraction, and determining the initial feature map includes:

[0034] The point cloud column is input into the backbone network of the pre-trained point cloud multi-task model for feature extraction to determine the initial feature map.

[0035] Furthermore, voxelizing each of the point cloud data to determine a corresponding point cloud column includes:

[0036] The point cloud data is input into a preset point voxel neural network for voxelization processing to determine the corresponding point cloud column.

[0037] In a second aspect, an embodiment of the present invention is intended to provide a point cloud multi-task processing device, the device comprising:

[0038] An acquisition unit, used for acquiring point cloud data in a detection range corresponding to different tasks, wherein the point cloud data has a task identifier corresponding to the task;

[0039] A processing unit is used to input each of the point cloud data into a backbone network in a pre-trained point cloud multi-task model for feature extraction and determine an initial feature map; input the initial feature map into a first network corresponding to each of the tasks in the point cloud multi-task model to map the task identifiers corresponding to the feature points in the initial feature map to determine the task output result of the corresponding task.

[0040] In a third aspect, an embodiment of the present invention aims to provide a computer program product, wherein the computer program product comprises a computer program / instructions, and when the computer program / instructions are executed by a processor, the method as described in any one of the above items is implemented.

[0041] In a fourth aspect, an embodiment of the present invention aims to provide an electronic device, comprising a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement a method as described in any one of the above items.

[0042] In a fifth aspect, an embodiment of the present invention aims to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps as described in any one of the above items are implemented.

[0043] The technical solution of the embodiment of the present invention obtains point cloud data with task identifiers of corresponding tasks under the detection ranges corresponding to different tasks, inputs each of the point cloud data into the backbone network of the pre-trained point cloud multi-task model for feature extraction to determine the initial feature map, and inputs the initial feature map into the first network corresponding to each of the tasks in the point cloud multi-task model to map based on the task identifiers corresponding to the feature points in the initial feature map to determine the task output result of the corresponding task. Thus, by extracting features from point cloud data of different tasks under different detection ranges and mapping based on the task identifiers to determine the task processing results of the corresponding tasks, tasks in different detection ranges can be processed simultaneously, thereby improving the task processing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0045] Figure 1 is a schematic diagram of a point cloud multi-task model according to an embodiment of the present invention;

[0046] Figure 2 is a flow chart of a point cloud multi-task processing method according to an embodiment of the present invention;

[0047] Figure 3 is a flow chart of determining point cloud data corresponding to each task according to an embodiment of the present invention;

[0048] Figure 4 is a schematic diagram of a point cloud multi-task model training process according to an embodiment of the present invention;

[0049] Figure 5 is a flow chart of a point cloud multi-task model training method according to an embodiment of the present invention;

[0050] Figure 6 is a schematic diagram of determining an expected result of a task according to an embodiment of the present invention;

[0051] Figure 7 is a schematic diagram of a spatial attention network according to an embodiment of the present invention;

[0052] Figure 8 is a schematic diagram of a point cloud multi-task processing device according to an embodiment of the present invention;

[0053] Fig. 9 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.

[0055] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.

[0056] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".

[0057] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0058] The solutions described in this specification and in the examples, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.

[0059] The point cloud task processing model in the prior art can only process a single point cloud task or point cloud tasks with the same point cloud detection range, but cannot process point cloud tasks with different detection ranges at the same time. Moreover, when processing different tasks at the same time, the task processing accuracy of different tasks is reduced due to the differences in the actual detection ranges of different tasks. In view of this, the embodiment of the present invention aims to provide a point cloud multi-task processing model and method to achieve simultaneous processing of different point cloud tasks with different detection ranges and improve the processing accuracy of point cloud tasks.

[0060] In this embodiment, the autonomous driving vehicle is taken as the target object, and the point cloud multi-tasking process related to the use of the autonomous driving vehicle is taken as an example for explanation. It should be understood that the method in this embodiment can also be applied to point cloud multi-tasking scenarios of other target objects (such as intelligent logistics robots, etc.), without strictly limiting the specific usage scenarios of the method.

[0061] Figure 1 Schematic diagram of a point cloud multi-task model according to an embodiment of the present invention. Figure 1 The point cloud multi-task model shown processes the point cloud data under different detection ranges corresponding to different tasks to determine the task output results of the corresponding tasks. The point cloud multi-task model includes a backbone network A and first networks B1-Bn corresponding to different tasks, and the number n of the first networks B is the same as the number of tasks. When performing point cloud multi-task processing, in this embodiment, the point cloud data under different detection ranges corresponding to different tasks are input into the backbone network A, and then the feature map obtained after processing by the backbone network A is input into the first network corresponding to each task for processing, thereby obtaining the task output results corresponding to each task. For example, the point cloud data 1, point cloud data 2, ..., and point cloud data n corresponding to task 1, task 2, ..., and task n are respectively input into the point cloud multi-task model for processing, and then the task output results 1, task output results 2, ..., and task output results n of each corresponding task are determined by the first networks B1-Bn corresponding to each task. Therefore, in this embodiment, different tasks corresponding to point cloud data with different detection ranges can be processed simultaneously through the point cloud multi-task model, which is conducive to improving the efficiency of point cloud multi-task processing. At the same time, since each task can determine the task output result based on the point cloud data within the corresponding detection range, the processing accuracy of different tasks can be improved.

[0062] Optionally, since the detection ranges corresponding to different tasks may overlap, in this embodiment, when the point cloud data under different detection ranges corresponding to different tasks are input into the backbone network A for processing, the point cloud data under the detection range corresponding to each task can be input into the backbone network A separately, that is, the point cloud data under the detection range corresponding to one task is input once, and the point cloud data corresponding to multiple different tasks need to be input into the backbone network A multiple times in total, and the number of inputs is the same as the number of tasks. Among them, the points in the point cloud under the detection range corresponding to each task and the data corresponding to each point have corresponding task identifiers; the point cloud data under the detection range corresponding to different tasks can also be integrated first, and all the integrated point cloud data are input into the backbone network A at the same time, that is, the point cloud data corresponding to multiple different tasks only need to be input into the backbone network A once. Among them, the points in the point cloud corresponding to each task are distinguished by different task identifiers, and the points in multiple detection ranges have multiple task identifiers, and the corresponding relationship between the points in the point cloud and the tasks is represented by the task identifier.

[0063] It should be understood that the point cloud data n (n=1, 2, ...) of this embodiment corresponds to all point clouds within the detection range of the corresponding task n. Each point cloud has position coordinates, attribute information (such as intensity, color, etc.) and a corresponding task label. That is, the point cloud data n is a collection of a large number of discrete point clouds and related data of the point clouds within the detection range corresponding to the task n.

[0064] It should be noted that the detection range corresponding to each task in this embodiment can be set according to the specific usage scenario. The detection ranges corresponding to different tasks can be the same or different.

[0065] Figure 2 is a flow chart of the point cloud multi-task processing method according to an embodiment of the present invention. Figure 2 As shown, the point cloud multi-task processing method in this embodiment includes the following steps.

[0066] In step S100, point cloud data in detection ranges corresponding to different tasks are obtained, wherein the point cloud data has a task identifier corresponding to the task.

[0067] In this embodiment, different tasks are used to characterize point cloud tasks with different detection ranges or point cloud tasks with different task types. Point cloud task types include target detection, semantic segmentation, etc., and the number and combination of tasks can be set according to the actual usage scenario. Among them, target detection is used to mark and identify specific target objects in the point cloud, such as cars, pedestrians, buildings, etc., and each point cloud point can be marked as belonging to a certain target category. Semantic segmentation is used to assign each point in the point cloud to different semantic categories, such as roads, buildings, trees, etc. Each point cloud point is marked as belonging to a certain semantic category to achieve pixel-level segmentation of the point cloud.

[0068] Optionally, the point cloud data in this embodiment is acquired by a sensor (e.g., a radar sensor). When acquiring point cloud data in the detection range corresponding to different tasks, this embodiment obtains the original point cloud acquired by the sensor, filters the original point cloud according to the detection range corresponding to different tasks, determines the point cloud data corresponding to each task, and assigns the corresponding task identifier to the point cloud data corresponding to each task.

[0069] Furthermore, in this embodiment, when determining the point cloud data corresponding to each task, the point cloud data corresponding to each task is determined according to the position coordinates of each point cloud point in the original point cloud and the detection range coordinates corresponding to the task. Figure 3 The steps shown are to determine the point cloud data corresponding to each task.

[0070] In step S310, the coordinate range of each task corresponding to the detection range is determined.

[0071] In this embodiment, in order to improve the accuracy of the task processing results, the point cloud data uses three-dimensional data. By constructing a universal three-dimensional coordinate system (such as an xyz coordinate system), the coordinate ranges of different tasks corresponding to the detection range under the same coordinate system and the position coordinates of each point cloud point in the point cloud data are determined, so as to determine the point cloud data corresponding to each task according to the position coordinates of the point cloud point and the coordinate range corresponding to the detection range. Among them, the coordinate origin in the three-dimensional coordinate system is the location of the target object, such as the location of the autonomous driving vehicle. The position coordinates of the point cloud point are determined according to the relative position of the point cloud point and the vehicle.

[0072] In step S320, the point cloud within the coordinate range corresponding to each task is determined.

[0073] In this embodiment, the point cloud data corresponding to each task is determined by judging whether each point cloud point in the original point cloud is within the coordinate range. Furthermore, in this embodiment, the position coordinates of each point cloud point are compared with the coordinate range of each detection range to judge whether the corresponding point cloud point is within at least one coordinate range corresponding to each task. Specifically, when all coordinates of the point cloud point are within the coordinate value range corresponding to the task, it is determined that the point cloud point is within the coordinate range of the corresponding task; when at least one of the coordinates of the point cloud point is outside the coordinate value range corresponding to the task, it is determined that the point cloud point is outside the coordinate range of the corresponding task.

[0074] Assume that there are two tasks. Task 1 is target detection, which aims to identify and determine specific objects in the point cloud and use a bounding box (such as a 3D box) to frame the object's position; Task 2 is semantic segmentation, which aims to determine the bounding box of the object in the point cloud and determine the object category of the object (such as vehicles, pedestrians, flowers, etc.); Usually, the detection range of target detection is larger than that of semantic segmentation. For example, the detection range corresponding to Task 1 is 100m (that is, the detection range threshold is 100m), and the corresponding coordinate range is [-100, -100, -1.5, 100, 100, 4.5], that is, the coordinates x, y and z of the point cloud points in the detection range of Task 1 are respectively in the intervals [-100, 100], [-100, 100] and [-1.5, 4.5]. The detection range corresponding to Task 2 is 40m, and the corresponding coordinate range is [-40, -40, -1.5, 40, 40, 4.5]. That is, the coordinates x, y, and z of the point cloud points within the detection range corresponding to Task 2 are in the intervals [-40, 40], [-40, 40], and [-1.5, 4.5], respectively.

[0075] Next, take the example of judging whether the three point cloud points in the point cloud (including point cloud point 1, point cloud point 2 and point cloud point 3) are located in the detection range corresponding to tasks 1 and 2. Assume that the coordinates of point cloud point 1 are (80, 30, 4), the coordinates of point cloud point 2 are (30, 20, 2), and the coordinates of point cloud point 3 are (110, 20, 2). Since the coordinates x, y and z of point cloud point 1 are all located in the coordinate interval corresponding to task 1, that is, point cloud point 1 is located in the coordinate range corresponding to task 1, then point cloud point 1 is a point cloud point located in the coordinate range of task 1; since the coordinates of point cloud point 2 are x, y and z are all located in the coordinate intervals corresponding to task 1 and task 2 at the same time, that is, point cloud point 2 is located in the coordinate ranges corresponding to task 1 and task 2 at the same time, then point cloud point 2 is a point cloud point located in the coordinate ranges of task 1 and task 2 at the same time; since the coordinate x of point cloud point 3 is not located in the coordinate intervals corresponding to task 1 and task 2, that is, point cloud point 3 is neither located in the coordinate range corresponding to task 1 nor in the coordinate range corresponding to task 2, then point cloud point 3 is neither a point cloud point in the coordinate range corresponding to task 1 nor a point cloud point in the coordinate range corresponding to task 2.

[0076] Optionally, when determining the point cloud within the coordinate range corresponding to each task, in this embodiment, one point cloud point can be locked in turn, and the coordinates of the current point cloud point can be compared with the coordinate range corresponding to each task, to determine the coordinate range to which the current point cloud point belongs, and then determine the point cloud within the coordinate range corresponding to each task according to the task corresponding to the coordinate range to which each point cloud point belongs; or one task can be locked in turn, and the coordinates of each point cloud point can be compared with the coordinate range corresponding to the current task, to determine the point cloud point within the coordinate range corresponding to the current task, and then determine the point cloud within the coordinate range corresponding to each task. Therefore, in this embodiment, by providing different comparison methods, it is possible to make the coordinate comparison of each point cloud point with each task more convenient, speed up the determination of whether each point cloud point is within the coordinate range of each task, and help improve the efficiency of determining the point cloud data corresponding to each task.

[0077] Furthermore, in this embodiment, the point cloud corresponding to each task is determined by the Range mask module. By pre-setting the detection range or coordinate range corresponding to the task in the Rangemask module, the original point cloud is input into the Range mask module corresponding to different tasks for screening, and the point cloud corresponding to each task can be determined. Among them, the Range mask module is a functional module in Adobe Lightroom and Adobe Camera Raw software, which can be used to accurately control the masking and selective editing of the image. Therefore, the point cloud corresponding to each task can be determined by selecting the point cloud in the specific detection range corresponding to each task in the original point cloud through the Range mask module.

[0078] At the same time, in order to clearly distinguish the point clouds within the detection range corresponding to different tasks and the point clouds outside the detection range, in this embodiment, after determining the point clouds within the coordinate range corresponding to each task by the above method, the relevant data of different point clouds are distinguished and processed through the following steps.

[0079] In step S330, data corresponding to the point cloud whose position coordinates in the original point cloud are within the coordinate range is determined as point cloud data of the corresponding task.

[0080] In this embodiment, after determining the point clouds whose position coordinates are within the coordinate range corresponding to each task, that is, after determining the point clouds within the coordinate range corresponding to each task, the point clouds within each coordinate range are determined as the point clouds of the corresponding tasks, and the task identifier of the task is added to the point cloud points in the corresponding point cloud to distinguish the point clouds corresponding to different tasks, so as to facilitate the subsequent corresponding task processing based on the task data corresponding to the point cloud, thereby improving the corresponding task processing efficiency.

[0081] Optionally, in order to better express the point cloud data corresponding to each task and facilitate subsequent processing, in this embodiment, after determining the point cloud data corresponding to each task, the point cloud data corresponding to each task will be voxelized (Voxelization) to determine the corresponding point cloud column (Pillar). Among them, the point cloud column has the same data content as the point cloud data, and only differs in the expression form. The point cloud data before voxelization is usually represented by an array or a similar data structure. Each point cloud point in the point cloud corresponding to the point cloud has position coordinates and attribute information; voxelization can convert the point cloud data into a uniform voxel grid, each voxel represents a volume unit of the same size, and forms a point cloud column. Compared with the point cloud data before voxelization, the point cloud column generated after voxelization is usually represented by multidimensional data or a sparse data structure, and stored in the form of a three-dimensional grid. The data density is more uniform, the data structure identifier is more compact, and the storage and calculation efficiency are higher. Therefore, by voxelizing the point cloud data and inputting the voxelized point cloud column into the pre-trained point cloud multi-task model for processing, the point cloud multi-task model can process the point cloud data more easily and efficiently, which is conducive to further improving the task processing efficiency.

[0082] Furthermore, the voxelization method in this embodiment can adopt one of grid voxelization, octree voxelization, distance field voxelization, depth map voxelization and surface reconstruction voxelization. It should be understood that the above voxelization methods mentioned in this embodiment are only examples, and the specific usage method can be set according to the specific usage scenario, which is not limited here.

[0083] In step S340, the position coordinates corresponding to the point clouds in the original point cloud whose position coordinates are outside the coordinate range are reset.

[0084] In this embodiment, after determining the point cloud in the original point cloud that is outside the coordinate range of the corresponding task, the point cloud outside the detection range is distinguished and displayed by resetting the position coordinates of the point cloud points (for example, resetting the xyz coordinates to 0 or -1), so as to reduce the interference of invalid point clouds under the corresponding task during task processing, which is conducive to improving task processing efficiency.

[0085] In step S200, each point cloud data is input into the backbone network of the pre-trained point cloud multi-task model for feature extraction to determine an initial feature map.

[0086] In this embodiment, the backbone network and the model parameters of the backbone network in the point cloud multi-task model are shared between different tasks. After determining the point cloud data under the detection range corresponding to each task, each point cloud data is input into the backbone network in the pre-trained point cloud multi-task model for feature extraction, and an initial feature map is determined. Optionally, after voxelizing the point cloud data to generate the corresponding point cloud column, in this embodiment, the point cloud column corresponding to each point cloud data is input into the pre-trained point cloud multi-task model for feature extraction, and the corresponding initial feature map is determined.

[0087] Optionally, the backbone network in this embodiment preferably adopts the PointPillar model, and may also adopt the centerpoints model or other models. Among them, PointPillars is a point cloud perception model based on two-dimensional convolution. By projecting the point cloud data into a two-dimensional feature space and then processing it using a two-dimensional convolutional neural network, it can have high efficiency and speed when processing large-scale point cloud data, and show good performance in target detection tasks. The CenterPoint model is a point cloud target detection model based on center point prediction. It realizes target detection and positioning by predicting the center point of the object in the point cloud and related attributes (such as category, size, orientation, etc.). It combines the sparsity of point cloud data and the local characteristics of the object. By using a specific neural network architecture and loss function for training, it can achieve accurate target detection on point cloud data. At the same time, since the PointPillars and CenterPoint models are mainly designed for point cloud target detection tasks, when the tasks to be processed include semantic segmentation, the classification tags (task identifiers) corresponding to each point cloud in the point cloud data and some specific model architectures, such as PointNet, PointNet++, PointCNN, KPConv, etc., can be used to extract and aggregate point clouds, and then use convolutional neural networks (CNN) or similar structures to perform point-level classification predictions, thereby achieving semantic segmentation of point clouds.

[0088] In step S300, the initial feature map is input into the first network corresponding to each task in the point cloud multi-task model, so as to map the task identifiers corresponding to the feature points in the initial feature map to determine the task output results of the corresponding tasks.

[0089] In this embodiment, after the backbone network in the point cloud multi-task model determines the initial feature map based on the point cloud data of each task, the initial feature map is transmitted to the first network corresponding to each task respectively. Each first network maps the features in the feature map based on the task identifier corresponding to the feature point in the initial feature map, generates a corresponding prediction result, and then determines the task output result of the corresponding task.

[0090] Optionally, the first network in this embodiment uses a head network to output the task output results of each task. When the processing task is a target detection task, the corresponding first network is used to generate a target detection result based on the initial feature map, and the target detection result includes the position (bounding box) and category prediction of the target (such as a vehicle, pedestrian, etc.). Furthermore, the first network used for target detection in this embodiment includes a region of interest (Region of Interest, RoI) pooling layer, a classifier (Classifier) ​​and a bounding box regressor (Bounding Box Regressor). Among them, the region of interest pooling layer is used to map the candidate box in the extracted feature map (based on the position information of the candidate box) to a feature map of a fixed size; the classifier is used to predict the category of each RoI, which is usually implemented using a fully connected layer or a convolutional layer plus a global average pooling layer; the bounding box regressor is used to regress the bounding box position of each RoI, and a fully connected layer is usually used to predict the coordinate offset of the bounding box.

[0091] Alternatively, in this embodiment, when the processing task is a semantic segmentation task, the corresponding first network is used to convert the initial feature map into a pixel-level semantic segmentation result, that is, to assign a semantic category label to each pixel. Furthermore, the first network for semantic segmentation includes a convolution layer, an upsampling layer (Upsampling) and a classifier (Classifier). Among them, the convolution layer is used to perform further convolution operations on the feature map to capture semantic information of different scales. The upsampling layer is used to enlarge the low-resolution feature map to the same resolution as the input image through interpolation or deconvolution operations. The classifier is used to predict the category of each pixel, which is usually implemented using a convolution layer or a fully connected layer.

[0092] It should be understood that the selection of the network structure of the first network corresponding to different tasks in this embodiment can be adjusted according to factors such as task requirements (such as target detection, semantic segmentation, etc.), processing data characteristics (such as point cloud data density, distribution and noise, etc.) and performance requirements (model calculation efficiency, model size, loss function type, etc.) to better adapt to the task processing process, which is conducive to further improving the task processing efficiency.

[0093] The technical solution of the embodiment of the present invention obtains point cloud data with task identifiers of corresponding tasks under the detection range corresponding to different tasks, inputs each point cloud data into the backbone network of the pre-trained point cloud multi-task model for feature extraction to determine the initial feature map, and inputs the initial feature map into the first network corresponding to each task in the point cloud multi-task model, so as to map the task identifiers corresponding to the feature points in the initial feature map to determine the task output results of the corresponding tasks. Thus, by extracting features from point cloud data of different tasks under different detection ranges and mapping based on the task identifiers to determine the task processing results of the corresponding tasks, it is possible to simultaneously process different tasks corresponding to point cloud data under different detection ranges, thereby improving the point cloud multi-task processing efficiency. At the same time, since the output results of each task are determined based on the point cloud data under the corresponding detection range, the interference of invalid point clouds can be reduced during the processing, thereby improving the accuracy of point cloud multi-task processing.

[0094] Figure 4 Schematic diagram of the point cloud multi-task model training process of an embodiment of the present invention. Figure 4 As shown in the figure, since the point cloud data used in task processing has spatial attributes, and in order to improve the accuracy of task processing results, in the point cloud multi-task model training stage, this embodiment introduces a second network based on the original point cloud multi-task model structure. The second network is an auxiliary network introduced in the point cloud multi-task model training process to complete the model training. The processing flow of the second network is similar to that of the first network corresponding to each task. The difference is that a spatial attention mechanism is added to learn the importance of input data in the spatial dimension through the spatial attention mechanism, and selectively focus on and highlight the features of different positions during the processing process, effectively capture the point cloud position information and spatial relationship, and can further improve the task processing performance.

[0095] At the same time, it should be noted that the second network in this embodiment is only used in the training process of the point cloud multi-task model, and does not appear in the process of point cloud multi-task processing by the point cloud multi-task model obtained after training. This can improve the processing accuracy of point cloud multi-task without increasing the network model inference calculation time and computing power consumption.

[0096] For ease of understanding, this embodiment will be combined with Figure 4 Continue to introduce the training process of the point cloud multi-task model.

[0097] Figure 5 : is a flow chart of the point cloud multi-task model training method according to an embodiment of the present invention. Figure 5 As shown, in this embodiment, it is obtained by training through the following steps.

[0098] In step S510, a training data set is obtained.

[0099] In this embodiment, the training data set includes a plurality of samples, and the samples include point cloud data in detection ranges corresponding to different tasks, and the point cloud data has a corresponding task identifier.

[0100] In this embodiment, the point cloud data in the training data set is determined by processing the original data. Among them, the original point cloud includes all the point clouds collected by the sensor. After obtaining the original point cloud, the original point cloud is input into the Range mask module 10 corresponding to each task, and the original point cloud is screened by the Range mask module 10 based on the detection range threshold preset by the corresponding task, and the point cloud and point cloud data corresponding to each task are determined, and the corresponding task identifier is added to each point cloud (the point cloud points in each point cloud also have the task identifier of the corresponding task). Afterwards, the point cloud data corresponding to each task is simultaneously input into the voxelization network 20 for voxelization processing to convert the original point cloud data into the corresponding point cloud column.

[0101] In step S520, the point cloud data in each sample is input into the backbone network in the preset point cloud multi-task model for feature extraction to determine the sample feature map of each sample.

[0102] In this embodiment, the point cloud columns formed by voxelizing the point cloud data in each sample are respectively input into the backbone network 30 in the preset point cloud multi-task model for feature extraction to determine the sample feature map of the corresponding sample. The loss function of the point cloud multi-task model is determined according to the loss function of the task branch corresponding to each task.

[0103] Optionally, since the point clouds in each point cloud data have corresponding task labels during the model training stage, in this embodiment, the loss function of the corresponding task branch can be determined according to the output results of each task and the task labels of the point clouds under the corresponding tasks, and the loss functions of each task branch are added together to determine the multi-task joint loss function corresponding to the point cloud task model.

[0104] For example, in the same training batch, assuming that the task labels of all tasks exist, the loss values ​​loss1 and loss2 corresponding to task A and task B are determined respectively, and then the corresponding loss function LossA is determined according to the loss value loss1 of task A, and the corresponding loss function LossB is determined according to the loss value loss2 corresponding to task B, and the loss functions LossA and LossB are added to determine the multi-task joint loss function Loss=LossA+LossB; assuming that only the task label of task A exists, the corresponding loss function LossA is determined according to the loss value loss1 corresponding to task A, the loss function LossB corresponding to task B is set to 0, and the multi-task joint loss function Loss=LossA is determined; assuming that only the task label of task B exists, the corresponding loss function LossB is determined according to the loss value loss2 corresponding to task B, the loss function LossA corresponding to task A is set to 0, and the multi-task joint loss function Loss=LossB is determined. Therefore, in this embodiment, the multi-task joint loss function corresponding to the point cloud multi-task model can be quickly determined by the above method, so that the task output result determined by the point cloud multi-task model based on the multi-task joint loss function has a certain accuracy and reliability, which is conducive to improving the use value of the task output result. At the same time, this embodiment can make the trained point cloud multi-task model have the ability to handle single tasks and multiple tasks at the same time through the above method, which is conducive to improving the task adaptability of the point cloud multi-task model and expanding the scope of application of the point cloud multi-task model.

[0105] It should be noted that since the model processing structure of the backbone network usually includes multiple processing levels, each processing level can output the feature map of the corresponding level, and the processing level close to the input side of the backbone network is usually defined as a shallow layer, the closer to the input side, the shallower the level, and the closer to the output side, the deeper the level. Accordingly, the feature map obtained by the level close to the input side of the backbone network is defined as a shallow feature map, and the feature map obtained by the level close to the output side of the backbone network is defined as a deep feature map. At the same time, in this embodiment, the feature map obtained by the output layer of the backbone network is defined as a sample feature map.

[0106] In step S530, each sample feature map is input into the first network corresponding to each task and the second network corresponding to at least one task for mapping, and the corresponding task output result and the task expected result are determined. The second network maps the sample feature map based on the spatial attention mechanism.

[0107] In this embodiment, since the task processing process corresponding to the point cloud task, such as target detection and semantic segmentation, involves position or other spatial information, in this embodiment, a second network is set up in the model training stage, and the sample feature map is mapped based on the spatial attention mechanism in the second network to output the expected task result with higher accuracy. Model training and parameter adjustment are performed based on the difference between the expected task result of the second network data under the same task and the task processing result output by the general first network, which can improve the processing accuracy of the trained point cloud multi-task model.

[0108] Optionally, considering the different requirements of different point cloud tasks for point cloud spatial information analysis, for example, semantic segmentation has more detailed requirements for spatial analysis than target detection, in this embodiment, a corresponding second network can be set for part or all of the task training processes, and the number of second networks is less than or equal to the number of first networks. Therefore, through the above-mentioned setting of the second network, while training to improve the processing accuracy of the point cloud multi-task model, the training process of the point cloud multi-task model can be accelerated, and the training efficiency of the point cloud multi-task model can be improved.

[0109] Furthermore, in this embodiment, Figure 4 The training process shown in FIG. 1 takes the second network corresponding to semantic segmentation as an example to illustrate the method for determining the expected result of the task. The method for determining the expected result of the task corresponding to the task based on the second network is as follows: Figure 6 As shown, the specific steps include the following steps.

[0110] In step S610, a sample feature map output by the backbone network and a shallow feature map corresponding to the sample feature map are obtained.

[0111] In this embodiment, the second network 50 receives the shallow feature map and the sample feature map output by the backbone network 30. The shallow feature map is the processing result of the layer close to the input side of the backbone network, and the sample feature map is the processing result output by the output layer of the backbone network. Compared with the sample feature map, the shallow feature map has a smaller receptive field and a larger feature map size, and contains more detailed spatial information, including more fine-grained features and local information.

[0112] Optionally, in this embodiment, the sample feature maps input to the second network 50 and each first network 40 through the backbone network 30 may be the same, or may be feature maps of different sizes obtained after processing the same point cloud data.

[0113] In step S620, the upsampled sample feature map and the shallow feature map are fused to determine a first intermediate feature map.

[0114] In this embodiment, the sizes of the shallow feature map and the sample feature map received by the second network 50 are usually different, and the size of the sample feature map is smaller than that of the shallow feature map. Therefore, after obtaining the sample feature map and the shallow feature map, the sample feature map will be upsampled first, the size of the sample feature map will be enlarged to be the same as the size of the shallow feature map, and the upsampled sample feature map will be fused with the shallow feature map of the same size to determine the first intermediate feature map, so that the first intermediate feature map can contain more point cloud information.

[0115] In step S630, the first intermediate feature map is processed based on a preset spatial attention network to determine the weights corresponding to each position point in the first intermediate feature map.

[0116] Optionally, the second network 50 in this embodiment includes a spatial attention network. After determining the first intermediate feature map, the first intermediate feature map is processed by a spatial attention mechanism based on a preset spatial attention network to determine the weights corresponding to each position point (i.e., pixel point) in the intermediate feature map. Among them, the spatial attention mechanism is a mechanism for emphasizing or suppressing different spatial positions in a neural network, which can learn the importance of input data in the spatial dimension and weight or selectively focus on features at different positions during processing.

[0117] Optionally, Figure 7 Schematic diagram of a spatial attention network according to an embodiment of the present invention. Figure 7 As shown, the spatial attention network 51 in this embodiment includes a maximum pooling layer (Maxpool) and an average pooling layer (Avgpool). Among them, the average pooling layer performs weighted averaging of features according to the importance score of each position to generate global pooling features. The maximum pooling layer performs weighted maximization of features according to the importance score of each position to highlight the features of important areas.

[0118] Furthermore, if Figure 7As shown, when the first intermediate feature map F is processed by the spatial attention network in this embodiment, the average pooling operation is first performed on the input first intermediate feature map F through the average pooling layer to obtain the average pooling result AvgPool(F), and then the maximum pooling operation is performed on the first intermediate feature map F to obtain the maximum pooling result MaxPool(F). The average pooling result AvgPool(F) and the maximum pooling result MaxPool(F) are respectively input into two independent multi-layer perceptrons MLP. Each multi-layer perceptron performs a series of linear and nonlinear transformations on the input to generate a corresponding intermediate feature representation. Finally, the intermediate feature representations of the two multi-layer perceptrons are added, and nonlinear mapping is performed through the activation function σ to obtain the final output result Mc(F). Among them, the dimension of the output result is the same as the dimension of the first intermediate feature map, and each spatial position corresponds to an attention weight. Specifically, the output result of the spatial attention network in this embodiment can be expressed by adopting the following formula:

[0119]

[0120] Among them, M c (F) is the output of the spatial attention network; F is the input of the spatial attention network, that is, the first intermediate feature map; MLP is the perceptron; σ represents the activation function (sigmoid), W0∈R C / r×C , W1∈R C×C / r .

[0121] Therefore, in this embodiment, the spatial attention network enables the model to adaptively focus on important areas in the feature map and extract key features, so that the point cloud multi-task model can effectively capture location information and spatial relationships, thereby improving the performance of the task.

[0122] In step S640, weighted summation is performed on the upsampled sample feature maps according to the weights to determine a second intermediate feature map.

[0123] In this embodiment, after determining the weights corresponding to each position point in the first intermediate feature map through the spatial attention network 51, the upsampled sample feature map is weighted summed according to each weight to determine the second intermediate feature map, so as to highlight the important features in the sample feature map through the second intermediate feature map.

[0124] In step S650, feature mapping is performed on the second intermediate feature map to determine the corresponding expected result of the task.

[0125] Optionally, in this embodiment, the second network 50 includes the first network in addition to the spatial attention network 51. After determining the second intermediate feature map, the first network in the second network 50 maps the second intermediate feature map to determine the expected result of the task corresponding to the semantic segmentation task (i.e., out3 as shown in 4). Compared with the conventional task processing results, since the expected result of the task involves more spatial information, the information expression is closer to the actual information and has higher accuracy.

[0126] In step S540, the backbone network parameters of the corresponding task branches are adjusted according to the output results of each task and the expected results of the task.

[0127] In this embodiment, the output result of each task is determined by the first network corresponding to the task in the point cloud multi-task model, and the task expectation is determined by the second network output. The method of determining the task output result by the first network corresponding to each task and determining the task expected result of the corresponding task by the second network has been introduced before and will not be repeated here.

[0128] Furthermore, in this embodiment, after determining the output results corresponding to each task, since the expected results of the task are more accurate than the corresponding task output results, the backbone network parameters of the corresponding task branches will be adjusted according to the task output results and the expected results corresponding to each task, so that the trained point cloud multi-task model has better processing performance and outputs task processing results that are more in line with the actual situation, thereby improving the point cloud multi-task processing accuracy.

[0129] Optionally, when adjusting the backbone network parameters of the corresponding task branch according to the task output results and the expected results of the task, the difference between the expected task result and the task output result of the corresponding task branch will be determined, and then the backbone network parameters of the corresponding task branch will be adjusted according to the difference, so that the task output result obtained by the trained point cloud multi-task model after processing the same point cloud data of the same task is consistent with the corresponding task expected result, thereby improving the processing performance of the point cloud multi-task model and the point cloud multi-task processing accuracy.

[0130] Furthermore, in this embodiment, the difference between the task output result and the task expected result is represented by KL divergence. KL divergence can characterize the difference between the distribution of the task output result and the task expected result, that is, relative entropy.

[0131] For ease of understanding, in this embodiment, Figure 4The training process shown in the figure illustrates the method of adjusting network parameters according to the task output results and the task expected results. Assume that the task output result of task A (target detection) is determined to be out1 after being processed by the backbone network 30 and the corresponding first network 40 in the point cloud multi-task model, and the task output result of task B (semantic segmentation) is determined to be out2 after being processed by the point cloud multi-task model, and the task expected result of task B is determined to be out3 after being processed by the backbone network 30 and the second network 50 in the point cloud multi-task model. Therefore, in this embodiment, the backbone network parameters of the corresponding task branch can be adjusted by the difference KL-loss corresponding to the task expected result out3 corresponding to task B and the task output result out2.

[0132] It should be noted that the task output result mentioned in this embodiment tends to be consistent with the expected task result means that when processing the same point cloud data corresponding to the same task, the task output result determined by the point cloud multi-task model is exactly the same as the expected task result determined by the second network or the difference is less than a preset threshold, and the preset threshold can be set according to the actual usage scenario.

[0133] The technical solution of this embodiment obtains a training data set for training the point cloud multi-task model by screening the original point cloud collected by the sensor based on the detection range corresponding to different tasks and voxelizing the point cloud data corresponding to each task obtained by the screening, and inputs the point cloud data in the obtained training data set into the backbone network in the preset point cloud multi-task model for feature extraction, determines the shallow feature map and the sample feature map, and then maps the sample feature map through the first network corresponding to each task in the point cloud multi-task model to determine the task output result of the corresponding task, and maps the shallow feature map and the sample feature map through the second network based on the spatial attention mechanism to determine the task expected result of the corresponding task, and then adjusts the backbone network parameters of the corresponding task branch according to the task output result and the task expected result, so that the trained point cloud multi-task model has better processing performance, outputs task processing results that are more in line with the actual situation, and further improves the point cloud multi-task processing accuracy.

[0134] Figure 8 Schematic diagram of a point cloud multi-task processing device according to an embodiment of the present invention. Figure 8As shown, the point cloud multi-task processing device in this embodiment includes an acquisition unit 1 and a processing unit 2. The acquisition unit 1 is used to acquire point cloud data under the detection range corresponding to different tasks, and the point cloud data has a task identifier of the corresponding task. The processing unit 2 is used to input each point cloud data into the backbone network in the pre-trained point cloud multi-task model for feature extraction and determine the initial feature map; input the initial feature map into the first network corresponding to each task in the point cloud multi-task model, so as to map the task identifier corresponding to the feature point in the initial feature map and determine the task output result of the corresponding task.

[0135] Optionally, when acquiring point cloud data under the detection range corresponding to different tasks, the acquisition unit 1 in this embodiment is also used to acquire the original point cloud collected by the sensor, screen the original point cloud according to the detection range corresponding to different tasks, and determine the point cloud data corresponding to each task. Further, when determining the point cloud data corresponding to each task, it is also used to determine the coordinate range under the detection range corresponding to each task, determine the data corresponding to the point cloud whose position coordinates in the original point cloud are within the coordinate range as the point cloud data of the corresponding task, and reset the position coordinates corresponding to the point cloud whose position coordinates in the original point cloud are outside the coordinate range.

[0136] Furthermore, the acquisition unit 1 in this embodiment is also used to voxelize each point cloud data, determine the corresponding point cloud column, input each point cloud data into the backbone network in the pre-trained point cloud multi-task model for feature extraction, and determine the initial feature map.

[0137] Optionally, the point cloud multi-task processing device in this embodiment also includes a training unit 3. The training unit 3 is used to obtain a training data set including multiple samples, each sample including point cloud data with a corresponding task identifier under the detection range corresponding to different tasks; the point cloud data in each sample is input into a preset backbone network for feature extraction, the sample feature map in the corresponding sample is determined, each sample feature map is input into the first network corresponding to each task and the second network corresponding to at least one task for mapping, the corresponding task output result and the task expected result are determined, and the backbone network parameters of the corresponding task branch are adjusted according to each of the task output results and the task expected results, so as to train and obtain a point cloud multi-task model for point cloud multi-task processing. The loss function of the point cloud multi-task model is determined according to the loss function of the task branch corresponding to each of the tasks. The second network maps the sample feature map based on the spatial attention mechanism.

[0138] Optionally, the training unit 3 in this embodiment is also used to obtain the sample feature map output by the backbone network and the shallow feature map corresponding to the sample feature map when determining the expected task result of the corresponding task; fuse the upsampled sample feature map and the shallow feature map to determine a first intermediate feature map; process the first intermediate feature map based on a preset spatial attention network to determine the weights corresponding to each position point in the first intermediate feature map; perform weighted summation of the upsampled sample feature map according to each weight to determine a second intermediate feature map; perform feature mapping on the second intermediate feature map to determine the corresponding expected task result.

[0139] Furthermore, when the training unit 3 adjusts the backbone network parameters of the corresponding task branch according to the task output results and the expected task results, it is also used to determine the difference between the expected task results and the task output results of the corresponding task branch; and adjust the backbone network parameters of the corresponding task branch according to the difference.

[0140] Fig. 9 Schematic diagram of an electronic device according to an embodiment of the present invention. Fig. 9 The electronic device shown is a general address query device, which includes a general computer hardware structure, which at least includes a processor 91 and a memory 92. The processor 91 and the memory 92 are connected via a bus 93. The memory 92 is suitable for storing instructions or programs executable by the processor 91. The processor 91 can be an independent microprocessor or a collection of one or more microprocessors. Thus, the processor 91 executes the instructions stored in the memory 92, thereby executing the method flow of the embodiment of the present invention as described above to realize the processing of data and the control of other devices. The bus 93 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to the display controller 94 and the display device and the input / output (I / O) device 95. The input / output (I / O) device 95 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices known in the art. Typically, the input / output device 95 is connected to the system via an input / output (I / O) controller 96.

[0141] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices (equipment) or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may adopt a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0142] The present application is described with reference to flowcharts of methods, apparatuses (devices) and computer program products according to embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0143] These computer program instructions may be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device that implements the process Figure 1 A function specified in a process or multiple processes.

[0144] These computer program instructions may also be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce the instructions for implementing the process Figure 1 A device that specifies functions in a process or multiple processes.

[0145] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used for a computer to execute part or all of the above method embodiments.

[0146] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by specifying relevant hardware through a program, and the program is stored in a storage medium, including several instructions for a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0147] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A point cloud multi-task processing method, characterized in that: The method comprises: Acquire point cloud data within detection ranges corresponding to different tasks, wherein the point cloud data has a task identifier corresponding to the task; Input each of the point cloud data into the backbone network of the pre-trained point cloud multi-task model to extract features and determine an initial feature map; The initial feature map is input into the first network corresponding to each of the tasks in the point cloud multi-task model, so as to map the task identifiers corresponding to the feature points in the initial feature map to determine the task output results of the corresponding tasks.

2. The method according to claim 1, characterized in that The point cloud multi-task model is trained by the following steps: Acquire a training data set, wherein the training data set includes a plurality of samples, wherein the samples include point cloud data under detection ranges corresponding to different tasks, and the point cloud data has a corresponding task identifier; Inputting the point cloud data in each of the samples into a preset backbone network for feature extraction, and determining a sample feature map of each of the samples; Input each of the sample feature graphs into a first network corresponding to each of the tasks and a second network corresponding to at least one of the tasks for mapping, and determine the corresponding task output result and the task expected result; wherein the second network maps the sample feature graph based on a spatial attention mechanism; The backbone network parameters of the corresponding task branches are adjusted according to the output results of each task and the expected results of the task.

3. The method according to claim 2, characterized in that The second network maps the sample feature map based on the spatial attention mechanism to determine the expected result of the corresponding task, including: Obtaining a sample feature map output by the backbone network and a shallow feature map corresponding to the sample feature map; The upsampled sample feature map and the shallow feature map are fused to determine a first intermediate feature map; Processing the first intermediate feature map based on a preset spatial attention network to determine a weight corresponding to each position point in the first intermediate feature map; Performing weighted summation on the upsampled sample feature maps according to the weights to determine a second intermediate feature map; Perform feature mapping on the second intermediate feature map to determine the corresponding expected result of the task.

4. The method according to claim 2, characterized in that: The adjusting of the backbone network parameters of the corresponding task branches according to the output results of each task and the expected results of the task includes: Determine the difference between the expected result of the task and the task output result of the corresponding task branch; The backbone network parameters of the corresponding task branch are adjusted according to the difference.

5. The method according to claim 1, characterized in that The loss function of the point cloud multi-task model is determined according to the loss function of the task branch corresponding to each task.

6. The method according to claim 1, characterized in that The acquisition of point cloud data within the detection range corresponding to different tasks includes: Get the original point cloud collected by the sensor; The original point cloud is screened according to the detection ranges corresponding to different tasks to determine the point cloud data corresponding to each task.

7. The method according to claim 6, characterized in that The screening of the original point cloud according to the detection ranges corresponding to different tasks to determine the point cloud data corresponding to each task includes: Determine the coordinate range of each task under the detection range; Determine the data corresponding to the point cloud whose position coordinates in the original point cloud are within the coordinate range as the point cloud data of the corresponding task; The position coordinates corresponding to the point cloud in the original point cloud whose position coordinates are outside the coordinate range are reset.

8. The method according to claim 1, characterized in that: The method further comprises: voxelize each of the point cloud data to determine the corresponding point cloud column; Inputting each of the point cloud data into the backbone network of the pre-trained point cloud multi-task model for feature extraction, and determining the initial feature map includes: The point cloud column is input into the backbone network of the pre-trained point cloud multi-task model for feature extraction to determine the initial feature map.

9. The method according to claim 8, characterized in that The voxelizing the point cloud data to determine the corresponding point cloud column comprises: The point cloud data is input into a preset point voxel neural network for voxelization processing to determine the corresponding point cloud column.

10. A point cloud multi-task processing device, characterized in that: The device comprises: An acquisition unit, used for acquiring point cloud data in a detection range corresponding to different tasks, wherein the point cloud data has a task identifier corresponding to the task; A processing unit is used to input each of the point cloud data into a backbone network in a pre-trained point cloud multi-task model for feature extraction and determine an initial feature map; input the initial feature map into a first network corresponding to each of the tasks in the point cloud multi-task model to map the task identifiers corresponding to the feature points in the initial feature map to determine the task output result of the corresponding task.

11. A computer program product, characterized in that The computer program product comprises a computer program / instructions, and when the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 9 are implemented.