Task execution method and device

By obtaining the target feature map in a multi-task perception network, generating heat map tags and adjusting the detection loss function, the problem of inefficiency of the target detection task in the prior art is solved, and efficient object detection without using the anchor box is achieved.

CN120032229APending Publication Date: 2025-05-23BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311575710.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When performing object detection tasks, existing multi-task aware networks need to establish complex anchor boxes in advance, resulting in inefficient execution of object detection tasks.

Method used

By obtaining the first target feature map of the original image passing through the first network and the second network, a heat map label of the object detection task is generated, and the detection task loss function of the current detection model is adjusted to obtain the object detection model, and then the object detection task is performed without using the anchor box.

Benefits of technology

The object detection is realized through a multi-task-aware network without using anchor boxes, which improves the execution efficiency of the object detection task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004567048760000091
    Figure BDA0004567048760000091
  • Figure BDA0004567048760000101
    Figure BDA0004567048760000101
  • Figure HDA0004567048870000011
    Figure HDA0004567048870000011
Patent Text Reader

Abstract

The embodiment of the invention provides a task execution method and device. According to the embodiment of the invention, a first target feature map of an original image passing through a first network and a second network is obtained; generating a thermodynamic diagram label of a target detection task based on a first target feature map, and determining at least one target detection frame and a target category value of the at least one target detection frame from the first target feature map by a current detection model, according to the at least one target detection frame and the difference value between the target category of the at least one target detection frame and the thermodynamic diagram label, continuously adjusting a current detection task loss function to obtain a target detection model, and executing a target detection task based on the target detection model and the first target feature map to obtain at least one target detection object, the task of target detection is realized through the multi-task sensing network under the condition that an anchor frame is not used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic technology, and in particular to a task execution method and device. Background Art

[0002] In the field of autonomous driving, a perception network with corresponding functions can be established according to the task type. Current task types include driving area segmentation, lane line detection, target detection, etc. In this context, different perception networks are relatively independent. If you want to achieve parallel processing of tasks of different task types, you need to deploy and run multiple perception network models on the vehicle side at the same time.

[0003] In the related technology, the above problems can be solved by designing a multi-task perception network. In the current multi-task perception network, when performing the task of target detection, it is necessary to establish an anchor frame in advance and determine the target detection object in the anchor frame. However, the process of establishing the anchor frame is relatively complicated, resulting in low execution efficiency of the target detection task. Summary of the invention

[0004] The present invention provides a task execution method and device to solve the problem that the multi-task perception network in the related art cannot realize the task of target detection.

[0005] In a first aspect, an embodiment of the present application provides a task execution method, comprising:

[0006] Obtain a first target feature map of the original image after passing through the first network and the second network;

[0007] Generate a heat map label for the target detection task based on the first target feature map;

[0008] Determine at least one target detection frame and a target category value of the at least one target detection frame from the first target feature map based on the current detection model;

[0009] Based on the at least one target detection frame and the difference value between the target category of the at least one target detection frame and the heat map label, adjusting the detection task loss function of the current detection model to obtain a target detection model;

[0010] The target detection task is performed based on the target detection model and the first target feature map to obtain at least one target detection object.

[0011] In a second aspect, an embodiment of the present application provides a task execution device, including:

[0012] An acquisition module, used to acquire a first target feature map of the original image after passing through the first network and the second network;

[0013] A generation module, used to generate a heat map label for a target detection task based on the first target feature map;

[0014] A determination module, configured to determine at least one target detection frame and a target category value of the at least one target detection frame from the first target feature map based on a current detection model;

[0015] An adjustment module, configured to adjust the detection task loss function of the current detection model based on the at least one target detection frame and the difference between the target category of the at least one target detection frame and the heat map label to obtain a target detection model;

[0016] An execution module is used to perform the target detection task based on the target detection model and the first target feature map to obtain at least one target detection object.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processing component, a storage component and a display component, wherein the storage component stores one or more computer instructions, and the one or more computer instructions are used to be called and executed by the processing component to implement the task execution method described in the first aspect.

[0018] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which, when executed by a computer, implements the task execution method described in the first aspect.

[0019] An embodiment of the present application provides a task execution method and device. In the embodiment of the present application, a first target feature map of an original image is obtained through a first network and a second network; a heat map label of a target detection task is generated based on the first target feature map, and a current detection model determines at least one target detection box and a target category value of the at least one target detection box from the first target feature map, and continuously adjusts the current detection task loss function through the difference between the at least one target detection box and the target category of the at least one target detection box and the heat map label to obtain a target detection model, and performs a target detection task based on the target detection model and the first target feature map to obtain at least one target detection object, so as to achieve the target detection task through a multi-task perception network without using an anchor frame. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application, and do not constitute an improper limitation on the present application.

[0021] Figure 1 A schematic diagram showing a process flow of a task execution method provided in an embodiment of the present application is shown;

[0022] Figure 2 A schematic diagram of the structure of a task execution network provided in an embodiment of the present application is shown;

[0023] Figure 3 A schematic diagram showing the structure of another task execution network provided in an embodiment of the present application is shown;

[0024] Figure 4 A schematic diagram showing the structure of another task execution network provided in an embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram of the structure of a task execution device provided in an embodiment of the present application is shown;

[0026] Figure 6 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] In order to make those of ordinary skill in the art better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples consistent with some aspects of the present application as detailed in the attached claims.

[0029] Before introducing the solution of this application, let me first introduce the background technology of this application:

[0030] In the field of autonomous driving, the existing technical route mainly establishes a perception network with corresponding functions according to the task type (such as drivable area segmentation, lane detection, target detection, etc.). In this context, different networks are relatively independent. If you want to achieve parallel processing of multiple tasks, you need to deploy and run multiple perception network models on the vehicle side at the same time.

[0031] With the advancement of deep learning technology, different types of perception networks have been successfully applied to the field of autonomous driving and have achieved practical results. Among them, the network models commonly used for drivable area segmentation and lane detection include the DeepLab series and the BiseNet series. The network models used for target detection include the anchor-based YOLO series and the anchor-free CornerNet and CenterNet. However, the above network models are all single-task models. Different models cannot share information (network weights, labels) and reuse input images. The total inference time will increase significantly with the increase in the number of tasks.

[0032] Considering the similarities between different tasks in the field of autonomous driving, some studies have proposed using the same backbone network to extract features from input images, and then designing corresponding decoding networks based on the characteristics of each task. Among them, the more representative multi-task networks include the YOLOP series, HybridNet, etc. Compared with single-task models, multi-task networks can, on the one hand, use a shared backbone network to reduce the number of network parameters and significantly increase the network inference speed, and on the other hand, improve the generalization ability of the network by sharing information between different tasks.

[0033] In the current multi-task perception network, when performing target detection tasks, it is necessary to establish an anchor frame in advance and determine the target detection object in the anchor frame. However, the process of establishing the anchor frame is relatively complicated, resulting in low execution efficiency of the target detection task.

[0034] In order to solve the above technical problems, an embodiment of the present application provides a task execution method and device. In the embodiment of the present application, a first target feature map of an original image is obtained after passing through a first network and a second network; a heat map label of a target detection task is generated based on the first target feature map, and a current detection model determines at least one target detection box and a target category value of the at least one target detection box from the first target feature map, and continuously adjusts the current detection task loss function through the difference value between at least one target detection box and the target category of at least one target detection box and the heat map label to obtain a target detection model, and performs a target detection task based on the target detection model and the first target feature map to obtain at least one target detection object, so as to achieve the target detection task through a multi-task perception network without using an anchor frame.

[0035] In order to solve the problems of the prior art, the embodiments of the present application provide a task execution method, device, electronic device and storage medium. The task execution method provided by the embodiments of the present application is first introduced below.

[0036] Figure 1 FIG. 1 is a flow chart showing a task execution method provided by an embodiment of the present application. Figure 1 As shown, the execution device of the method may be an on-board terminal in a vehicle, and the method may include:

[0037] S101, obtaining a first target feature map of an original image after passing through a first network and a second network.

[0038] The original image may be captured by an image acquisition device in the vehicle, and further sent by the image acquisition device to the vehicle-mounted terminal, wherein the vehicle-mounted terminal is provided with a first network and a second network.

[0039] The first network may be a backbone network, which is a downsampling network, i.e., used for feature extraction, dimensionality reduction, and learning high-level abstract representations. A downsampling network is usually composed of multiple downsampling layers, each of which reduces the spatial size of the input data and increases the number of features. The sampling rate of the first network can be flexibly set according to actual conditions, which will not be elaborated here.

[0040] The second network can be a feature pyramid structure network, which is an upsampling network. The second network is used to realize feature fusion of different scales and extract higher-level semantic features. The sampling rate of the second network can be flexibly set according to actual conditions, which will not be elaborated here.

[0041] In some embodiments, after the original image is acquired, data enhancement processing is performed on the original image, and the data enhancement mainly includes conventional image enhancement methods such as random rotation, horizontal flipping, and cropping.

[0042] In an optional embodiment, the sampling rate of the first network can be 32 times the downsampling rate, and the sampling rate of the second network can be 4 times the upsampling rate. In this embodiment, the first target feature map is one quarter of the original image size.

[0043] S102: Generate a heat map label for the target detection task based on the first target feature map.

[0044] In some embodiments, heat map labels can be used in various computer vision tasks such as target detection, image segmentation, and image classification to help us understand the decision-making process and focus areas of the model. The generation process of heat map labels can include the following steps:

[0045] Map the extracted features to the heatmap space. This can be done using techniques such as transposed convolution or inverse pooling. These operations restore the spatial dimensions of the features and increase the resolution.

[0046] Heatmap generation: The mapped features are further processed to obtain the final heatmap labels. Common methods include using activation functions to transform features into probability distributions and performing smoothing operations to reduce noise.

[0047] Visualization: Finally, the generated heatmap labels can be visualized with the original input data to better understand how much the model is paying attention to the input data. This usually involves overlaying or multiplying the heatmap onto the input image, creating the effect of highlighting areas of interest.

[0048] S103. Determine at least one target detection box and a target category value of at least one target detection box from the first target feature map based on the current detection model.

[0049] The vehicle terminal is also equipped with a current detection block, which is used to predict at least one target detection frame and a target category value of at least one target detection frame based on the first target feature map. The target category value is the probability distribution of the category to which the target detection frame belongs.

[0050] It can be understood that the first target feature map may include image features of multiple categories. For example, the original image includes three cats and two dogs, so five target detection frames are determined, and the category value of each target detection frame can be the probability distribution of whether the target detection frame belongs to a cat or a dog.

[0051] S104. Based on at least one target detection box and the difference value between the target category value of at least one target detection box and the heat map label, adjust the detection task loss function of the current detection model to obtain a target detection model.

[0052] It can be understood that the current detection model includes a detection task loss function, the heat map label can be regarded as the label of the current detection model, and the labels of at least one target detection box and at least one target category value of the target detection box can be determined from the heat map label. Therefore, we can gradually adjust the detection task loss function of the current detection model based on the difference value between at least one target detection box and at least one target category value of the target detection box and the heat map label until the difference value converges to obtain a target detection model including the target detection task loss function. The difference value convergence can be that the difference value is minimum or the difference value no longer changes, etc.

[0053] S105. Perform a target detection task based on the target detection model and the first target feature map to obtain at least one target detection object.

[0054] Furthermore, the first target feature map can be input into the target detection model to obtain a target frame of at least one target detection object in the first target feature map, and the target frame can be solved to obtain at least one corresponding target detection object in the original image, wherein the solution process includes a compensation process, a regression process, and the like.

[0055] The embodiment of the present application provides a task execution method and device. The embodiment of the present application obtains a first target feature map of an original image through a first network and a second network; generates a heat map label for a target detection task based on the first target feature map, and the current detection model determines at least one target detection box and a target category value of at least one target detection box from the first target feature map, and continuously adjusts the current detection task loss function through the difference between at least one target detection box and the target category of at least one target detection box and the heat map label to obtain a target detection model, and performs the target detection task based on the target detection model and the first target feature map to obtain at least one target detection object, so as to achieve the target detection task through a multi-task perception network without using an anchor frame.

[0056] In some embodiments, since it is a multi-task perception network, the network is also used to perform lane line segmentation tasks, and the method also includes: generating a first label for the lane line segmentation task based on the first target feature map; determining a lane line segmentation marking map from the first target feature map based on the current lane line segmentation model; adjusting the lane line segmentation loss function of the current lane line segmentation model based on the lane line segmentation marking map and the first label to obtain a target lane line segmentation model; performing the lane line segmentation task based on the target lane line segmentation model and the first target feature map to obtain at least one lane line.

[0057] In some embodiments, the multi-task perception network is also used to perform a drivable area segmentation task, and the method also includes: obtaining a second target feature map of the original image through the first network; generating a second label for the drivable area segmentation task based on the second target feature map; determining a drivable area labeling map from the second target feature map based on the current drivable area model; adjusting the drivable area segmentation loss function of the current drivable area model based on the drivable area labeling map and the second label to obtain a target drivable area model; performing the drivable area segmentation task based on the target drivable area model and the second target feature map to obtain at least one drivable area.

[0058] Since the drivable area task does not require an overly deep network structure for feature extraction in experimental verification, the present invention directly connects the drivable area segmentation network with the backbone network, that is, the second target feature map can be input into the current drivable area network, so that the drivable area labeling map can be determined from the second target feature map, and based on the difference value between the drivable area labeling map and the second label, the drivable area segmentation loss function of the drivable area model is gradually adjusted to obtain a target drivable area model including a target drivable segmentation loss function, and the drivable area segmentation task is performed based on the target drivable area model and the second target feature map to obtain at least one drivable area.

[0059] At least one drivable area is directly determined from the second target feature map, which reduces the sampling process of the second network. While ensuring the correctness of reasoning, it can effectively reduce the number of network parameters and greatly improve the network reasoning speed.

[0060] The lane line segmentation task requires a deep network structure for feature extraction in experimental verification. Therefore, it is necessary to connect the lane line segmentation network with the feature pyramid structure network, that is, the first target feature map needs to be input into the lane line segmentation network so that the lane line segmentation marking map can be determined from the second target feature map, and based on the difference value between the lane line marking map and the second label, the lane line segmentation loss function of the lane line model is gradually adjusted to obtain the target lane line model including the target drivable segmentation loss function, and the lane line segmentation task is performed based on the target lane line model and the second target feature map to obtain at least one drivable lane line.

[0061] The first label and the second label can be implemented as a MASK image.

[0062] In an optional embodiment, a format conversion tool may be used to convert the json format labels in the lane segmentation and drivable area segmentation tasks into a MASK image format. The format conversion tool may be an open CV tool.

[0063] In an optional embodiment, the network used to perform the above three tasks in the embodiment of the present application is a further improvement on the Multitask CenterNet multi-task network, and a new multi-task network without anchor frames is proposed, which will be referred to as Multitask CenterNet v2, abbreviated as MCN v2, Figure 2 A schematic diagram of the structure of a task execution network provided in an embodiment of the present application is shown. Figure 2 As shown in the figure, the network includes a backbone network (BACKBONE), a feature pyramid structure network (NECK) and three decoding networks: a target detection network with detection points to be compensated, a drivable area segmentation decoding network and a lane line segmentation network.

[0064] Among them, the sampling rate of the backbone network is 32 times the downsampling rate, and the sampling rate of the feature pyramid structure network is 4 times the upsampling rate.

[0065] exist Figure 2In the method, the original image is input into the backbone network, which performs scale transformation on the original image and generates feature images of one quarter, one eighth, one sixteenth and one thirty-second sizes of the original image respectively. Finally, the feature image of one thirty-second size of the original image is input into the feature pyramid structure network, which upsamples the feature image of one thirty-second size to obtain feature images of one eighth and one sixteenth sizes of the original image respectively. The feature images of one eighth and one sixteenth sizes of the original image are fused with the feature images of one eighth and one sixteenth sizes of the original image in the feature pyramid structure network to obtain fused feature fused images of one eighth and one sixteenth sizes.

[0066] The feature image with a size of one thirty-second in the backbone network is input into the drivable area segmentation network, and the feature fusion image with a size of one eighth of the original image in the feature pyramid structure network is input into: the target detection network with key points and the lane line segmentation network so that each network can perform its corresponding tasks respectively.

[0067] In some embodiments, S101, obtaining a first target feature map of the original image after passing through the first network and the second network includes the following steps: extracting features of the original image through the first network to obtain a feature image; reducing the size of the feature image through the first network to obtain a second target feature map; and enlarging the second target feature map through the second network to obtain the first target feature map.

[0068] The feature image includes image features at different locations.

[0069] In some embodiments, the first network is a backbone network, the backbone network includes downsampling units of different sampling scales, each downsampling unit includes multiple sampling blocks, and the first network is used to determine the second target feature map based on image features at different positions, including: for any sampling block, inputting feature images at different positions into a first normalization layer to obtain a first normalized feature map, the first normalized feature image includes normalized image features at different positions; processing the normalized feature map through an attention mechanism to obtain a first fused feature map, the first fused feature map includes: normalized image features at different positions and weight features after fusion Fusion features; fusing fusion features at different positions with image features to obtain a second fusion feature map; inputting the first fusion feature map into a second normalization layer to obtain a second normalized feature map, the second normalized feature map including normalized fusion features at different positions; inputting the second normalized map into multiple convolutional layers to obtain a third fusion feature map, the third fusion feature map including: features obtained by fusing fusion features at different positions with image features and convolutional features at different positions obtained by multiple convolutional layers; obtaining a second target feature map obtained by reducing the size of the feature image through downsampling units of different sampling scales.

[0070] In some embodiments, the normalized feature map is processed through an attention mechanism to obtain a first fused feature map, including: obtaining context information of different positions; determining weight features of different positions based on the context information; and fusing the normalized feature maps of different positions with the weight features to obtain a first fused feature map.

[0071] In order to explain the processing flow of the backbone network in detail, Figure 3 FIG. 4 shows a schematic diagram of another task execution network provided in an embodiment of the present application. Figure 3 As shown in the figure, MCN v2 selects the MSCAN downsampling network used in SegNeXt as the backbone network. The backbone network is composed of downsampling units with different sampling scales, and the sampling scales are 4 times the downsampling rate, 8 times the downsampling rate, 16 times the downsampling rate and 32 times the downsampling rate.

[0072] Each downsampling unit includes multiple sampling blocks (Block blocks), each Block block consists of an Attention module, a 1*1 convolution layer (conv), a depth-wise separable convolution layer (dconv) with a convolution kernel size of 3*3, and a Batch Norm normalization layer. Among them, the use of a depth-wise separable convolution layer can effectively reduce the number of model parameters.

[0073] Specifically, Figure 3As shown, for any sampling block, feature images at different positions are input into the first normalization layer to obtain a first normalized feature map, which includes normalized image features at different positions; the normalized feature map is processed through the attention mechanism to obtain a first fused feature map, which includes: fused features after the normalized image features at different positions and weight features are fused; the fused features at different positions are fused with the image features to obtain a second fused feature map; the first fused feature map is input into the second normalization layer to obtain a second normalized feature map, which includes the normalized fused features at different positions; the second normalized map is input into multiple convolutional layers to obtain a third fused feature map, which includes: features after the fusion features at different positions are fused with the image features and features obtained by fusing the convolutional features at different positions obtained through multiple convolutional layers.

[0074] Furthermore, the third fusion feature is input into other Block blocks to obtain the second target feature map obtained by reducing the size of the feature image through downsampling units of different sampling scales.

[0075] Among them, the general implementation of the Attention mechanism includes the following steps:

[0076] Calculate attention weights: Calculate the correlation between different positions or channels through some calculation methods, such as inner product, dot product, weighted sum, etc., and get attention weights. These weights represent the importance of each position or channel.

[0077] Weighted summation: Multiply the original data (such as sequence, feature map) by the attention weight to obtain weighted data. This allows the model to pay more attention to important positions or channels in subsequent processing.

[0078] Normalization: Normalize the weighted data to ensure that its numerical range is reasonable and help the stability and convergence of the model.

[0079] It should be noted that the number of Block blocks in different downsampling units is as follows: Table 1 is a reference table for setting the number of Block blocks in different downsampling units in an embodiment of the present application.

[0080]

[0081]

[0082] Table 1

[0083] In some embodiments, the first target feature map includes a thermal map of at least one type of target detection object and multiple detection feature maps, each thermal map corresponds to a category value, S103, determining at least one target detection frame and a target category value of at least one target detection frame from the first target feature map based on the current detection model includes the following steps: for any target detection frame, determining the thermal map category value corresponding to the target detection frame as the target category value of the target detection frame, and the target category value is the probability distribution of the category to which the target detection frame belongs; determining a position where the thermal force is greater than a preset thermal threshold as a detection point to be compensated; determining the target detection frame height, target detection frame width, target detection point horizontal coordinate compensation and target detection point vertical coordinate compensation through multiple detection feature maps, and the multiple detection feature maps include: a height detection map, a width detection map, a target detection point horizontal coordinate compensation map and a target detection point vertical coordinate compensation map; determining the target detection point based on the detection point to be compensated, the target detection point horizontal coordinate compensation and the target detection point vertical coordinate compensation, the target detection point is the center point of the target detection frame, and the target category value of the target detection point is consistent with that of the target detection frame; determining the target detection frame based on the target detection point, the target detection frame height and the target detection frame width.

[0084] The detection point to be compensated is the center point of the target detection frame before compensation. The process of determining the detection point to be compensated may include the following steps:

[0085] Thresholding: In order to determine the detection points to be compensated, the generated heat map can be thresholded. By setting a suitable threshold, the heat map pixels below the threshold are set to 0, and the heat map pixels above the threshold are set to 1 or other fixed values. In this way, the heat map can be converted into a binary image.

[0086] The position where the thermal force is greater than a preset thermal force threshold is determined as a detection point to be compensated, wherein the thermal force threshold can be flexibly set according to actual conditions.

[0087] Clustering or connected region detection: For binary heat maps, clustering algorithms or connected region detection algorithms can be used to identify the locations of the detection points to be compensated. Common clustering algorithms include K-means clustering, DBSCAN, etc., while connected region detection algorithms can use methods such as connected component analysis. These algorithms can regard areas with high pixel density or connected areas as the locations of the detection points to be compensated.

[0088] Filtering and screening: Depending on the requirements of a specific task, some filtering and screening operations may be required for clusters or connected areas. For example, the number, size, shape and other characteristics of the detection points to be compensated can be eliminated or retained. This can further improve the accuracy and robustness of the detection of the detection points to be compensated.

[0089] Coordinate calibration of the detection point to be compensated: Finally, for the determined detection point to be compensated, it can be mapped to the coordinate space of the original input data by taking the center point of its pixel position or other defined methods to obtain the exact position of the detection point to be compensated in the input image.

[0090] It is understandable that after the target detection point is determined, the position and size of the target detection frame can be determined based on the coordinates of the target detection point, the height of the target detection frame, and the width of the target detection frame.

[0091] In order to explain the process of determining the target detection frame in detail, Figure 4 FIG. 4 shows a schematic diagram of another structure of a task execution network according to an embodiment of the present application. Figure 4 As shown in the figure, the target detection network of MCN v2 uses the decoding network with key point detection in the CenterNet network. The decoding network with key point detection includes: a heat map network for detecting the category and coordinates of the detection points to be compensated, a network for compensating the coordinates of the detection points to be compensated, a heat map network for detecting the category and center point coordinates of the target box, a network for compensating the center point coordinates of the target box, and a network for regressing the length and width of the target box.

[0092] After the upsampling layer, the decoding network with detection of the detection point to be compensated determines the detection result based on the first target feature map. The result of the heat map network for detecting the category and coordinates of the detection point to be compensated is expressed as (C_kpt, 1 / nSize, 1 / nSize), the result of the network for compensating the coordinates of the detection point to be compensated is expressed as (C_kpt, 1 / nSize, 1 / nSize), the result of the heat map network for detecting the category and center point coordinates of the target frame is expressed as (C, 1 / nSize, 1 / nSize), the network for compensating the center point coordinates of the target frame is expressed as (C, 1 / nSize, 1 / nSize), and the result of the network for regressing the length and width of the target frame is expressed as (C, 1 / nSize, 1 / nSize) where C_kpt and C are the number of categories of the detection point to be compensated and the target frame, respectively, and Size is the size of the original input image. Each network is composed of a fully convolutional network.

[0093] Assuming that the sampling rate of the backbone network is 32 times the downsampling rate and the sampling rate of the feature pyramid structure network is 4 times the upsampling rate, then n is 8. In general, the detection point to be compensated is consistent with the center point of the detection target box. Assuming that the result of the heat map network for detecting the category and coordinates of the detection point to be compensated is expressed as (2, 1, 2), it means that the number of categories of the detection point is 2, and the coordinates of the first feature map are (1, 2).

[0094] It should be noted that unlike the original CenterNet network, the object detection network of MCN v2 is connected to the feature pyramid structure network structure instead of directly connected to the backbone network. The above design can enable the decoding network to obtain richer feature information on the one hand, and on the other hand, adding the feature pyramid structure network will reduce the mutual interference between various tasks.

[0095] Compared with the target detection network, the two segmentation networks of MCN v2 are relatively simple in design. The up-convolutional layer composed of a convolutional layer and an upsampling layer is used to restore the feature map to the same size as the original image, and the segmentation result MASK image is obtained after conversion by the argmax layer.

[0096] In some embodiments, the detection task loss function of the current detection model includes: a current detection box category loss function, a current detection point category loss function, a current detection box size loss function and a current detection box position compensation loss function. S104: Based on at least one target detection box and the difference value between the target category of at least one target detection box and the heat map label, the detection task loss function of the current detection model is adjusted to obtain the target detection model, including the following steps: obtaining the detection point category value, the detection box category value, the detection box size and the detection box center position in the heat map label; gradually adjusting the weights corresponding to the current detection box category loss function, the current detection point category loss function, the current detection box size loss function and the current detection box position compensation loss function; adjusting the detection point category value and the target category value based on the first difference value, the detection point category value and the target category value, the detection box size loss function and the current detection box position compensation loss function; adjusting the detection task loss function of the current detection model based on the first difference value, the detection point category value and the target category value, the detection box size and the detection box center position; gradually adjusting the detection point category loss function, the current detection box size loss function and the current detection box position compensation loss function; adjusting the detection task loss function of the current detection model based on the first difference value, the detection point category value and the target category value, the detection box size and the detection box center position ... When the second difference between the detection box category value and the target category value, the third difference between the detection box size and the target detection box size, and the fourth difference between the detection box position, the target detection box center position, and the target detection point position all meet their corresponding convergence conditions, the current detection box category loss function, the current detection box category loss function, the current detection box size loss function, and the current detection box position compensation loss function are determined as the target detection box category loss function, the target detection box category loss function, the target detection box size loss function, and the target detection box position compensation loss function; the target detection task loss function is determined based on the detection box category loss function, the target detection box category loss function, the target detection box size loss function, and the target detection box position compensation loss function; and a target detection model including the target detection task loss function is obtained.

[0097] In some embodiments, the target detection task loss function is determined based on the detection box category loss function, the target detection box category loss function, the target detection box size loss function, and the target detection box position compensation loss function, satisfying formula (1):

[0098] L OD-kpt =L hm +λ size L size +λoff L off +L hm-kpt +η off L off (1)

[0099] Among them, L OD-kpt is the target detection task loss function, L hm is the target detection box category loss function and L hm-kpt is the target detection point category loss function, L size is the target detection box size loss function, L off is the target detection box position compensation loss function, λ size is the first hyperparameter of the target detection box size loss function, λ off and η off are the second and third hyperparameters of the target detection box position compensation loss function respectively.

[0100] The methods for adjusting each loss function mainly include the following aspects:

[0101] Weight parameter adjustment: For different loss functions, their contribution to the overall loss function can be changed by adjusting their weight parameters. The weight parameters can be set manually according to the requirements of the task and the characteristics of the dataset, or they can be automatically adjusted through methods such as cross-validation.

[0102] Loss function design: Depending on the specific task, a more appropriate loss function form can be designed. For example, in the task of detecting the detection points to be compensated, some additional constraints such as joint length and angle can be designed based on the spatial relationship between the detection points to be compensated, and taken into account in the loss function to improve the accuracy of the detection points to be compensated.

[0103] Data preprocessing: Appropriate preprocessing of the input data can improve the effect of the loss function. For example, in the object detection task, the image can be rotated, scaled, etc. to increase the diversity and robustness of the data, thereby improving the performance of the loss function.

[0104] Network structure design: Reasonable network structure design can help improve the effect of loss function. For example, in the target detection task, a multi-scale feature fusion mechanism can be adopted to utilize feature information of different scales, thereby improving the accuracy and robustness of target detection.

[0105] Data augmentation: During the training process, data augmentation methods can be used to expand the training set and improve the generalization ability of the model. For example, in the object detection task, operations such as random cropping and image flipping can be used to generate more diverse and rich training samples.

[0106] Each loss function is adjusted in the above manner so that the first difference value, the second difference value, the third difference value and the fourth difference value all meet their corresponding convergence conditions, wherein the convergence condition may be that the difference value is minimum or that the difference value no longer changes, etc.

[0107] In some embodiments, the above method further comprises:

[0108] Determine the total loss function based on the target drivable area segmentation loss function, the target lane line segmentation loss function and the target detection task loss function;

[0109] Based on the target model including the total loss function and the original image, at least one target detection object, at least one drivable area and at least one lane line are determined.

[0110] In some embodiments, a total loss function is determined based on the target drivable area segmentation loss function, the target lane line segmentation loss function and the target detection task loss function, satisfying formula (2):

[0111] L total =αL FS +βL LD +L OD-kpt (2)

[0112] Among them, L total is the total loss function, α and β are the hyperparameters for adjusting the weights of the three subtask loss functions. FS is the drivable area segmentation loss function, α is the fourth hyperparameter of the drivable area segmentation loss function; L LD is the lane segmentation loss function, β is the fifth hyperparameter of the lane segmentation loss function, L OD-kpt is the target detection loss function.

[0113] By using the trained network (which has converged) for reasoning, the prediction results of the target box, target detection points and segmentation tasks can be obtained. Among them, the MASK map of the lane line segmentation and drivable area segmentation results can be directly given by the corresponding network. The target box detection result with the detection point to be compensated needs to be obtained after solving the output result of the target detection network. The specific solution method of this solution refers to the post-processing process of the CenterNet network, that is, the category and center point coordinates of the target box and target detection points are calculated through the heat map network, the length and width of the target box are obtained through the network that regresses the length and width of the target box, and the network that compensates the center point coordinates of the target box is used to compensate for the accuracy loss caused by the change in the size of the feature map. Thus, at least one target detection object is obtained.

[0114] like Figure 5 As shown, the task execution device provided in the embodiment of the present application may include: an acquisition module 501, a generation module 502, a determination module 503, an adjustment module 504 and an execution module 505.

[0115] An acquisition module 501 is used to acquire a first target feature map of an original image after the first network and the second network;

[0116] A generating module 502, configured to generate a heat map label for a target detection task based on the first target feature map;

[0117] A determination module 503 is used to determine at least one target detection box and a target category value of at least one target detection box from the first target feature map based on the current detection model;

[0118] An adjustment module 504 is used to adjust the detection task loss function of the current detection model based on at least one target detection box and a difference value between the target category of at least one target detection box and the heat map label to obtain a target detection model;

[0119] The execution module 505 is used to perform the target detection task based on the target detection model and the first target feature map to obtain at least one target detection object.

[0120] An embodiment of the present application provides a task execution device. In the embodiment of the present application, a first target feature map of an original image is obtained through a first network and a second network; a heat map label of a target detection task is generated based on the first target feature map, and a current detection model determines at least one target detection box and a target category value of at least one target detection box from the first target feature map, and continuously adjusts the current detection task loss function through the difference between at least one target detection box and the target category of at least one target detection box and the heat map label to obtain a target detection model, and performs a target detection task based on the target detection model and the first target feature map to obtain at least one target detection object, so as to achieve the target detection task through a multi-task perception network without using an anchor frame.

[0121] In some embodiments, since it is a multi-task perception network, the network is also used to perform the lane segmentation task.

[0122] The generating module 502 is further used to generate a first label for the lane segmentation task based on the first target feature map;

[0123] The determination module 503 is further used to determine the lane line segmentation mark map from the first target feature map based on the current lane line segmentation model;

[0124] The adjustment module 504 is further used to adjust the lane line segmentation loss function of the current lane line segmentation model based on the lane line segmentation mark image and the first label to obtain a target lane line segmentation model;

[0125] The execution module 505 is further used to perform the lane line segmentation task based on the target lane line segmentation model and the first target feature map to obtain at least one lane line.

[0126] In some implementations, the multi-task perception network also performs the drivable area segmentation task.

[0127] The acquisition module 501 is also used to acquire a second target feature map of the original image after passing through the first network;

[0128] The generating module 502 is further used to generate a second label for the drivable area segmentation task based on the second target feature map;

[0129] The determination module 503 is further used to determine the driving area marking map from the second target feature map based on the current driving area model;

[0130] The adjustment module 504 is further used to adjust the drivable area segmentation loss function of the current drivable area model based on the drivable area marking map and the second label to obtain a target drivable area model;

[0131] The execution module 505 is further used to perform a drivable area segmentation task based on the target drivable area model and the second target feature map to obtain at least one drivable area.

[0132] In some embodiments, the acquisition module 501 is specifically used to:

[0133] Extract features from the original image through the first network to obtain a feature image, where the feature image includes image features at different positions;

[0134] Reduce the size of the feature image through the first network to obtain a second target feature map;

[0135] The second target feature map is enlarged by the second network to obtain the first target feature map.

[0136] In some embodiments, the first target feature map includes a heat map of at least one type of target detection object and multiple detection feature maps, each heat map corresponds to a category value, and the determination module 503 is specifically used to:

[0137] For any target detection frame, the heat map category value corresponding to the target detection frame is determined as the target category value of the target detection frame, and the target category value is the probability distribution of the category to which the target detection frame belongs;

[0138] A position where the thermal force is greater than a preset thermal force threshold is determined as a detection point to be compensated;

[0139] Determine the target detection frame height, target detection frame width, target detection point horizontal coordinate compensation and target detection point vertical coordinate compensation through multiple detection feature maps, the multiple detection feature maps include: a height detection map, a width detection map, a target detection point horizontal coordinate compensation map and a target detection point vertical coordinate compensation map;

[0140] The target detection point is determined based on the detection point to be compensated, the horizontal coordinate compensation of the target detection point, and the vertical coordinate compensation of the target detection point. The target detection point is the center point of the target detection frame, and the target category value of the target detection point is consistent with that of the target detection frame.

[0141] The target detection box is determined based on the target detection point, the target detection box height and the target detection box width.

[0142] In some embodiments, the detection task loss function of the current detection model includes: a current detection box category loss function, a current detection point category loss function, a current detection box size loss function and a current detection box position compensation loss function. The adjustment module 504 is specifically used to:

[0143] Get the detection point category value, detection box category value, detection box size and detection box center position in the heat map label;

[0144] Gradually adjust the weights corresponding to the current detection box category loss function, the current detection point category loss function, the current detection box size loss function, and the current detection box position compensation loss function;

[0145] When the first difference value between the detection point category value and the target category value, the second difference value between the detection box category value and the target category value, the third difference value between the detection box size and the target detection box size, and the fourth difference value between the detection box position, the target detection box center position and the target detection point position all meet their respective corresponding convergence conditions, the current detection box category loss function, the current detection box category loss function, the current detection box size loss function and the current detection box position compensation loss function are determined as the target detection box category loss function, the target detection box category loss function, the target detection box size loss function and the target detection box position compensation loss function.

[0146] Determine the target detection task loss function based on the detection box category loss function, the target detection box category loss function, the target detection box size loss function and the target detection box position compensation loss function;

[0147] Obtain the target detection model including the target detection task loss function.

[0148] In some embodiments, the target detection task loss function is determined based on the detection box category loss function, the target detection box category loss function, the target detection box size loss function, and the target detection box position compensation loss function, satisfying the relationship:

[0149] L OD-kpt =L hm +λ size L size +λ off L off +L hm-kpt +η off L off

[0150] Among them, L OD-kpt is the target detection task loss function, L hm is the target detection box category loss function and L hm-kpt is the target detection point category loss function, L size is the target detection box size loss function, L off is the target detection box position compensation loss function, λ size is the first hyperparameter of the target detection box size loss function, λ off and η off are the second and third hyperparameters of the target detection box position compensation loss function respectively.

[0151] In some embodiments, the determination module 503 is also used to: determine the total loss function based on the target drivable area segmentation loss function, the target lane line segmentation loss function and the target detection task loss function; and determine at least one target detection object, at least one drivable area and at least one lane line based on the target model including the total loss function and the original image.

[0152] In some embodiments, a total loss function is determined based on the target drivable area segmentation loss function, the target lane line segmentation loss function, and the target detection task loss function, satisfying the relationship:

[0153] L total =αL FS +βL LD +L OD-kpt

[0154] Among them, L total is the total loss function, α and β are the hyperparameters for adjusting the weights of the three subtask loss functions. FS is the drivable area segmentation loss function, α is the fourth hyperparameter of the drivable area segmentation loss function; L LD is the lane segmentation loss function, β is the fifth hyperparameter of the lane segmentation loss function, L OD-kpt is the target detection loss function.

[0155] Figure 6 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6As shown, the electronic device may include a processor 601 and a memory 602 storing computer programs or instructions.

[0156] Specifically, the processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.

[0157] The memory 602 may include a large capacity memory for data or instructions. For example, but not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 602 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid-state memory. The memory may include a read-only memory (ROM), a random access memory (RAM), a disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described in the task execution method provided in the above-mentioned embodiments.

[0158] The processor 601 implements any one of the task execution methods in the above embodiments by reading and executing computer program instructions stored in the memory 602 .

[0159] In one example, the electronic device may further include a communication interface 603 and a bus 610. Figure 6 As shown, the processor 601, the memory 602, and the communication interface 603 are connected via a bus 610 and communicate with each other.

[0160] The communication interface 603 is mainly used to implement the communication between the modules, devices, units and / or devices in the embodiment of the present invention.

[0161] Bus 610 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 610 may include one or more buses. Although the embodiment of the present invention describes and shows a specific bus, the present invention considers any suitable bus or interconnection.

[0162] The electronic device can execute the task execution method in the embodiment of the present invention, thereby realizing the task execution method described in the above embodiment.

[0163] In addition, in combination with the task execution method in the above embodiment, the embodiment of the present invention can provide a readable storage medium to implement the task execution method in the above embodiment. The readable storage medium stores program instructions, and when the program instructions are executed by the processor, any one of the task execution methods in the above embodiment is implemented.

[0164] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0165] The functional blocks shown in the above structured block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0166] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in a different order from the embodiments, or several steps can be performed simultaneously.

[0167] The above reference is according to the method of the embodiment of the present application, the flow chart of the device (system) and the computer program product and / or the block diagram described various aspects of the present application.It should be understood that each square box in the flow chart and / or the block diagram and the combination of each square box in the flow chart and / or the block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more square boxes of the flow chart and / or the block diagram.Such a processor can be but is not limited to a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It can also be understood that each square box in the block diagram and / or the flow chart and the combination of the square boxes in the block diagram and / or the flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.

[0168] The above is only a specific implementation of the present invention. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be covered within the protection scope of the present invention.

Claims

1. A task execution method, It is characterized in that include: Obtain a first target feature map of the original image after passing through the first network and the second network; Generate a heat map label for the target detection task based on the first target feature map; Determine at least one target detection frame and a target category value of the at least one target detection frame from the first target feature map based on the current detection model; Based on the at least one target detection frame and the difference value between the target category of the at least one target detection frame and the heat map label, adjusting the detection task loss function of the current detection model to obtain a target detection model; The target detection task is performed based on the target detection model and the first target feature map to obtain at least one target detection object.

2. The method according to claim 1, It is characterized in that The method further comprises: Generate a first label for the lane segmentation task based on the first target feature map; Determine a lane segmentation marking map from the first target feature map based on the current lane segmentation model; Based on the lane line segmentation mark image and the first label, adjusting the lane line segmentation loss function of the current lane line segmentation model to obtain a target lane line segmentation model; The lane line segmentation task is performed based on the target lane line segmentation model and the first target feature map to obtain at least one lane line.

3. The method according to claim 1, It is characterized in that The method further comprises: Obtain a second target feature map of the original image after passing through the first network; Generate a second label for the drivable area segmentation task based on the second target feature map; Determine a driving area marking map from the second target feature map based on the current driving area model; Based on the driving area marking map and the second label, adjusting the driving area segmentation loss function of the current driving area model to obtain a target driving area model; The drivable area segmentation task is performed based on the target drivable area model and the second target feature map to obtain at least one drivable area.

4. The method according to claim 1, It is characterized in that The method of obtaining a first target feature map of the original image after passing through the first network and the second network includes: Extract features from the original image through a first network to obtain a feature image, wherein the feature image includes image features at different positions; Reducing the size of the feature image through the first network to obtain a second target feature map; The second target feature map is enlarged by a second network to obtain a first target feature map.

5. The method according to claim 1, It is characterized in that The first target feature map includes a heat map of at least one type of target detection object and multiple detection feature maps, each heat map corresponds to a category value, and determining at least one target detection box and a target category value of the at least one target detection box from the first target feature map based on the current detection model includes: For any target detection frame, determine the heat map category value corresponding to the target detection frame as the target category value of the target detection frame, and the target category value is the probability distribution of the category to which the target detection frame belongs; A position where the thermal force is greater than a preset thermal force threshold is determined as a detection point to be compensated; Determine the target detection frame height, target detection frame width, target detection point horizontal coordinate compensation and target detection point vertical coordinate compensation through the multiple detection feature maps, wherein the multiple detection feature maps include: a height detection map, a width detection map, a target detection point horizontal coordinate compensation map and a target detection point vertical coordinate compensation map; Determine a target detection point based on the detection point to be compensated, the horizontal coordinate compensation of the target detection point, and the vertical coordinate compensation of the target detection point, wherein the target detection point is the center point of the target detection frame, and the target category value of the target detection point is consistent with that of the target detection frame; A target detection frame is determined based on the target detection point, the target detection frame height and the target detection frame width.

6. The method according to claim 5, It is characterized in that The detection task loss function of the current detection model includes: a current detection box category loss function, a current detection point category loss function, a current detection box size loss function and a current detection box position compensation loss function. The detection task loss function of the current detection model is adjusted based on the at least one target detection box and the difference between the target category of the at least one target detection box and the heat map label to obtain the target detection model, including: Obtain the detection point category value, detection box category value, detection box size and detection box center position in the heat map label; Gradually adjusting the weights corresponding to the current detection box category loss function, the current detection point category loss function, the current detection box size loss function, and the current detection box position compensation loss function; When the first difference value between the detection point category value and the target category value, the second difference value between the detection box category value and the target category value, the third difference value between the detection box size and the target detection box size, and the fourth difference value between the detection box position, the target detection box center position, and the target detection point position all meet their respective corresponding convergence conditions, the current detection box category loss function, the current detection box category loss function, the current detection box size loss function, and the current detection box position compensation loss function are determined as the target detection box category loss function, the target detection box category loss function, the target detection box size loss function, and the target detection box position compensation loss function. Determine the target detection task loss function based on the detection box category loss function, the target detection box category loss function, the target detection box size loss function and the target detection box position compensation loss function; A target detection model including the target detection task loss function is obtained.

7. The method according to claim 6, It is characterized in that The target detection task loss function is determined based on the detection box category loss function, the target detection box category loss function, the target detection box size loss function and the target detection box position compensation loss function, satisfying the relationship: L OD-kpt =L hm +λ size L size +λ off L off +L hm-kpt +n off L off Among them, L OD-kpt is the target detection task loss function, L hm is the target detection box category loss function and L hm-kpt is the target detection point category loss function, L size is the target detection box size loss function, L off is the target detection box position compensation loss function, λ size is the first hyperparameter of the target detection box size loss function, λ off and η off are the second and third hyperparameters of the target detection box position compensation loss function respectively.

8. The method according to any one of claims 1 to 7, It is characterized in that The method further comprises: Determine a total loss function based on the target drivable area segmentation loss function, the target lane line segmentation loss function and the target detection task loss function; Based on the target model including the total loss function and the original image, the at least one target detection object, the at least one drivable area and the at least one lane line are determined.

9. The method according to claim 8, It is characterized in that The total loss function is determined based on the target drivable area segmentation loss function, the target lane line segmentation loss function and the target detection task loss function, satisfying the relationship: L total =αL FS +βL LD +L OD-kpt Among them, L total is the total loss function, α and β are hyperparameters for adjusting the weights of the three subtask loss functions. FS is the drivable area segmentation loss function, α is the fourth hyperparameter of the drivable area segmentation loss function; L LD is the lane segmentation loss function, β is the fifth hyperparameter of the lane segmentation loss function, L OD-kpt is the target detection loss function.

10. A task execution device, It is characterized in that include: An acquisition module, used to acquire a first target feature map of the original image after passing through the first network and the second network; A generation module, used to generate a heat map label for a target detection task based on the first target feature map; A determination module, configured to determine at least one target detection frame and a target category value of the at least one target detection frame from the first target feature map based on a current detection model; An adjustment module, configured to adjust the detection task loss function of the current detection model based on the at least one target detection frame and the difference between the target category of the at least one target detection frame and the heat map label to obtain a target detection model; An execution module is used to perform the target detection task based on the target detection model and the first target feature map to obtain at least one target detection object.