Multi-task sensing method and device based on multispectral image, medium, program product and terminal
By combining a feature pyramid network of visible light and infrared images with a multi-task learning model, the problems of poor performance of autonomous driving perception methods under extreme weather conditions and multi-dimensional detection are solved, achieving efficient multi-task perception and improving the safety and efficiency of autonomous driving.
Patent Information
- Application Number
- CN202410610591.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
Existing autonomous driving perception methods based on visible light images perform poorly in extreme weather conditions, cannot output detection results from multiple dimensions simultaneously, and multi-task perception models suffer from overfitting and inference burden.
A multi-task perception method using multispectral images is adopted. Visible light and infrared images are acquired, and feature pyramid networks are used for feature extraction and fusion. Combined with E-Stage network and hollow spatial pyramid pooling network, multi-task feature maps are generated. Then, a multi-task learning model is used for training and prediction to generate region segmentation results and 3D detection boxes.
It improves the model's performance in extreme environments, enhances the accuracy and robustness of multi-dimensional detection, reduces the risk of overfitting, and improves the inference performance and processing speed of key autonomous driving tasks.
Smart Images

Figure CN120976882A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image recognition, and in particular to a multi-task perception method and device based on multi-spectral images, a medium, a program product and a terminal. BACKGROUND
[0002] In recent years, with the rise of artificial intelligence technology based on deep learning, deep learning methods have been widely applied in various fields, and in particular, deep learning methods based on automatic driving perception technology have become a current research hotspot. In the automatic driving perception task, drivable area segmentation, two-dimensional target detection and three-dimensional target detection on the road are relatively key tasks. Perception data usually comes from two types of sensors, namely image acquisition devices and radars. Due to the high cost of radars and other reasons, the industry tends to adopt a pure visual perception scheme based only on images.
[0003] For image-based perception schemes, target detection methods that rely only on visible light images as input data have certain limitations. First, visible light images alone cannot adapt to various weather conditions, especially in extreme weather conditions such as night, heavy fog and rainy weather. Second, in terms of model design, it is not possible to output multiple dimensional detection results simultaneously based on the input two-dimensional visible light images. Specifically, although there are many two-dimensional target detection algorithms for drivable area segmentation based on convolutional neural networks, and high accuracy has been achieved. However, in three-dimensional target detection, three-dimensional space is involved, and directly predicting three-dimensional information from two-dimensional images will result in the lack of depth information, making model learning more difficult and limiting the accuracy and performance of the model. Existing multi-task perception algorithms cannot perfectly output multiple dimensional detection results, and the reason for this is that even if a multi-task perception algorithm is used, the correlation between tasks may become ambiguous and unclear due to the lack of sufficient depth information in the input data, thereby reducing the effectiveness and reliability of the algorithm. In addition, while multi-task perception algorithms avoid overfitting to a particular task through mutual learning, the indiscriminate addition of redundant and ineffective tasks not only increases the inference burden, but also may cause the overall model to underfit. SUMMARY
[0004] In view of the above-mentioned shortcomings of the prior art, the present application aims to provide a multi-task perception method and device based on multi-spectral images, a medium, a program product and a terminal, which solves the problem of poor performance in extreme weather conditions caused by using only visible light as input in existing two-dimensional multi-task perception models, the inability to output multiple dimensional detection results, and the difficulty in designing branch structures for multi-dimensional output using multi-task perception models to balance the inference burden while preventing overfitting.
[0005] To achieve the above object and other related objects, the first aspect of the present application provides a multi-task perception method based on multi-spectral images, comprising: acquiring a visible light image and an infrared image of a target region respectively, and performing field of view matching on the visible light image and the infrared image; performing feature extraction on the visible light image and the infrared image after field of view matching based on a feature pyramid network, to generate a multi-task feature map; performing feature fusion on the multi-task feature map based on the feature pyramid network, and inputting the multi-task feature map into a multi-task learning model for training and prediction, to generate a region segmentation result and a three-dimensional detection box respectively.
[0006] In some embodiments of the first aspect of the present application, the process of performing feature extraction on the infrared image after field of view matching comprises: performing a two-fold downsampling convolution operation on the infrared image to generate a first feature map; inputting the first feature map into an E-Stage network module for two-fold downsampling to generate a second feature map with a size of one-eighth of the visible light image; inputting the second feature map into the E-Stage network module for two-fold downsampling to generate a third feature map with a size of one-sixteenth of the visible light image.
[0007] In some embodiments of the first aspect of the present application, the process of performing feature extraction on the visible light image after field of view matching comprises: performing four-fold downsampling on the visible light image and inputting into an E-Stage network module for two-fold downsampling to generate a fourth feature map with a size of one-eighth of the visible light image; inputting the fourth feature map into the E-Stage network module for two-fold downsampling to generate a fifth feature map with a size of one-sixteenth of the visible light image; performing weighted fusion on the third feature map and the fifth feature map to generate a weighted fusion feature map; inputting the weighted fusion feature map into the E-Stage network module for two-fold downsampling to generate a sixth feature map with a size of one-thirty-second of the visible light image; inputting the sixth feature map into a dilated spatial pyramid pooling network to generate a multi-task feature map.
[0008] In some embodiments of the first aspect of the present application, the process of performing feature extraction on the visible light image after field of view matching comprises: performing four times down-sampling on the visible light image and inputting into the E-Stage network module to perform two times down-sampling, to generate a fourth feature map with a size of one-eighth of the visible light image; inputting the fourth feature map into the E-Stage network module to perform two times down-sampling, to generate a fifth feature map with a size of one-sixteenth of the visible light image; performing weighted fusion on the third feature map and the fifth feature map, to generate a weighted fusion feature map; inputting the weighted fusion feature map into the E-Stage network module to perform two times down-sampling, to generate a sixth feature map with a size of one-thirty-second of the visible light image; inputting the sixth feature map into a dilated spatial pyramid pooling network, to generate a multi-task feature map.
[0009] In some embodiments of the first aspect of the present application, the E-Stage network module comprises a main branch and a residual structure; wherein the main branch comprises one or more ordinary convolution layers and one or more deformable convolution layers.
[0010] In some embodiments of the first aspect of the present application, the process of inputting the multi-task feature map into a multi-task learning model for training and prediction to generate a three-dimensional detection box comprises: constructing an auxiliary learning task; based on the loss function of the multi-task learning model and the loss function of the auxiliary learning task, training and optimizing the multi-task learning model; inputting the multi-task feature map into the trained multi-task learning model, to generate a three-dimensional detection box.
[0011] To achieve the above object and other related objects, the second aspect of the present application provides a multi-task perception device based on multi-spectral images, comprising: an image acquisition module: configured to acquire a visible light image and an infrared image of a target region respectively, and perform field of view matching on the visible light image and the infrared image; a feature extraction module: configured to perform feature extraction on the visible light image and the infrared image after field of view matching based on a feature pyramid network, to generate a multi-task feature map; a region segmentation prediction module: configured to perform feature fusion on the multi-task feature map based on the feature pyramid network, to generate a region segmentation result; a detection box prediction module: configured to input the multi-task feature map into a multi-task learning model for training and prediction, to generate a three-dimensional detection box.
[0012] To achieve the above object and other related objects, the third aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the multi-task perception method based on multi-spectral images.
[0013] To achieve the above object and other related objects, the fourth aspect of the present application provides a computer program product, which comprises computer program codes, and when the computer program codes are run on a computer, the computer is caused to implement the multi-task perception method based on multi-spectrum images.
[0014] To achieve the above object and other related objects, the fifth aspect of the present application provides a computer electronic terminal, which comprises a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the multi-task perception method based on multi-spectrum images.
[0015] As described above, the present application has the following beneficial effects: the present application proposes a multi-task learning model taking infrared images and visible light images as inputs at the same time, aiming to improve the performance of the model in extreme environments such as night, rainy day and heavy fog. Specifically, by designing a parallel feature fusion backbone network structure, image features are effectively extracted, and an improved U-Net and DeepLabV3 design is used on the semantic segmentation branch to optimize the multi-scale feature extraction process. In addition, the model introduces a multi-task learning mechanism of 3D monocular detection and semantic segmentation, which improves the learning efficiency and overall accuracy of the model, and reduces the risk of overfitting. The auxiliary learning branch further improves the robustness of the model in the three-dimensional frame detection task, and does not require additional labeling. In addition, the present application is an end-to-end multi-task processing method, which only needs image and camera intrinsic parameter input, and can simultaneously process key tasks of autonomous driving such as drivable area segmentation, two-dimensional target detection and three-dimensional target detection, which significantly improves the inference performance and processing speed compared with single-task models. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A flowchart of an embodiment of the multi-task perception method of multi-spectrum images of the present application is shown.
[0017] Figure 2 A flowchart of another embodiment of the multi-task perception method of multi-spectrum images of the present application is shown.
[0018] Figure 3 A structure diagram of a hollow space pyramid in an embodiment of the multi-task perception method of multi-spectrum images of the present application is shown.
[0019] Figure 4 A structure diagram of an E-Stage network module in an embodiment of the multi-task perception method of multi-spectrum images of the present application is shown.
[0020] Figure 5 A structure diagram of deformable convolution in an embodiment of the multi-task perception method of multi-spectrum images of the present application is shown.
[0021] Figure 6 Figure 1 shows a structural schematic diagram of an embodiment of the D-Block in the multi-task perception method of the multispectral image of the present application.
[0022] Figure 7 Figure 2 shows a structural schematic diagram of an embodiment of the multi-task perception device of the multispectral image of the present application.
[0023] Figure 8 Figure 3 shows a structural schematic diagram of an embodiment of the multi-task perception terminal of the multispectral image of the present application. DETAILED DESCRIPTION
[0024] The present application will be described in detail below through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosure of the present specification. The present application can also be implemented or applied in other different specific embodiments, and various modifications or changes can be made to the details in the specification based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0025] Before the present application is further described, the terms and nouns involved in the embodiments of the present application are explained, and the terms and nouns involved in the embodiments of the present application are applicable to the following explanations:
[0026] <1> FPN (Feature Pyramid Network): FPN is a network structure for target recognition, which realizes multi-scale target recognition by establishing a feature pyramid on different levels.
[0027] <2> Backbone Network: Backbone network is a basic network structure in a deep learning model, which is used to extract feature representation of input data.
[0028] <3> Detection Head: Detection head is the part of the network model structure responsible for predicting results and class labels.
[0029] <4> ASPP (Atrous Spatial Pyramid Pooling): ASPP is an image feature extraction module that expands the receptive field without losing resolution (without downsampling) through multi-scale atrous convolution to extract global and local features.
[0030] <5> Bounding Box: Bounding box is a two-dimensional rectangular bounding box and a three-dimensional bounding box used to represent the position of the target in the target detection task.
[0031] <6> Semantic Segmentation: Semantic segmentation is a computer vision task that aims to assign each pixel in an image to its corresponding semantic class.
[0032] <7> LABELME: LABELME is a tool and platform for labeling images and generating semantic segmentation labels.
[0033] <8> Joint Calibration: Obtain the coordinate transformation relationship between the laser radar spatial coordinate system and the camera spatial coordinate system, convert the real world accurate 3D annotation information obtained by the laser radar to the camera space, and obtain the 3D frame annotation in the camera coordinate system.
[0034] <9> Deformable Convolution: Deformable convolution is an operation in convolutional neural networks that allows the convolution kernel to perform flexible spatial transformations on the input feature map to adapt to the shape and pose of the target.
[0035] <10> Camera Intrinsics: In 3D detection algorithms, it is used to realize the conversion between camera 3D coordinate system and image pixel coordinate system.
[0036] The technical scheme of the embodiments of the present application combines the multi-task perception method of visible light and infrared light image data, which can be applied to the field of autonomous driving. This scheme helps the autonomous vehicle to efficiently recognize the lane lines, traffic signs on the road, detect obstacles such as vehicles and pedestrians, perceive the status of the surrounding environment and understand the scene, and realize functions such as automatic parking and automatic stopping, and is also suitable for target detection tasks in various extreme scenarios. Through this technical scheme, the autonomous vehicle can more accurately perceive and understand the surrounding environment, so as to make correct decisions and improve driving safety and efficiency. In addition, this scheme is also suitable for medical image analysis, agricultural field, industrial detection, urban planning and architectural design, and environmental monitoring fields. Based on image data analysis, this scheme can help doctors diagnose diseases, farmers monitor crop growth, industrial enterprises improve production efficiency, urban planners and architects plan and design, and environmental protection departments monitor environmental conditions. By realizing more accurate understanding and analysis of image data, this technical scheme can realize more efficient and accurate decision-making and application.
[0037] In order to understand the multi-task perception method based on multi-spectral images in the embodiments of the present application, first of all, the multi-task perception method based on multi-spectral images in the embodiments of the present application is combined with Figure 1 detailed description. Figure 1 A flowchart of a multi-task perception method based on multi-spectral images in the embodiments of the present application is shown. The multi-task perception method based on multi-spectral images in the embodiments of the present application mainly includes the following steps:
[0038] Step S11: Obtain the visible light image and the infrared image of the target region respectively, and perform field of view matching on the visible light image and the infrared image.
[0039] In an embodiment of the present application, the visible light image contains three main channels of red, green and blue, corresponding to red, green and blue light in the visible light wavelength range respectively. Through the combination of the three channels, the real color and appearance of the object can be presented. The infrared radiation information captured by a single infrared sensor is presented in a grayscale image, and different grayscale levels or colors represent different temperatures or heat intensities. Brighter (numerical value) areas represent higher temperature or stronger heat areas, while darker (smaller numerical value) areas represent lower temperature or weaker heat areas.
[0040] Step S12: Perform feature extraction on the visible light image and the infrared image after field of view matching based on the feature pyramid network to generate a multi-task feature map.
[0041] It should be noted that the input of the feature pyramid network is an image, and the output is a multi-scale feature map obtained after multi-scale feature extraction and fusion. In the feature pyramid network, each pyramid level corresponds to a feature map of a different scale, and these feature maps can contain feature information of different scales, from coarse global information to detailed local information. In the training process, the feature pyramid network optimizes the network parameters through the backpropagation algorithm to minimize the loss function and improve the performance of the network. After training is completed, the feature pyramid network can accept an input image and obtain multiple scale feature maps as output through the forward propagation algorithm to adapt to multiple tasks such as object detection, image segmentation, image classification, etc., thereby realizing more accurate and comprehensive image analysis and understanding.
[0042] It is worth noting that in actual application, the resolution of the infrared camera is usually lower than that of the visible light camera, and the feature information contained in the infrared image is also lower than that of the visible light image. Therefore, if the infrared image is directly resized to match the resolution of the visible light image and then fused, it will introduce additional useless feature information. Therefore, in an embodiment of the present application, before performing feature extraction on the visible light image and the infrared image after field of view matching, a size matching operation is also performed, which includes the following process: the infrared light image is cropped and scaled to 1:2 size to realize that the image field of view obtained by infrared light shooting and the image field of view obtained by visible light shooting are completely consistent. For example, if the infrared light image and the visible light image obtained by shooting are both 640x640, the infrared image is resized to 320x320 or the visible light image is enlarged to 1280x1280. In actual image acquisition scenarios, a high-resolution camera is generally used to obtain the visible light image, and the resolution of the image captured by infrared light is relatively low.
[0043] In an embodiment of the present application, the process of feature extraction performed on the infrared image after field of view matching includes: performing a two-fold downsampling convolution operation on the infrared image to generate a first feature map; inputting the first feature map into an E-Stage network module for two-fold downsampling to generate a second feature map with a size of one-eighth of the size of the visible light image; inputting the second feature map into the E-Stage network module for two-fold downsampling to generate a third feature map with a size of one-sixteenth of the size of the visible light image.
[0044] Figure 2 The leftmost side shows the process of feature extraction performed on the infrared image, which includes: using a two-fold downsampling convolution operation on the infrared image to generate a first feature map with a size of 1*240*320 after convolution, and then inputting the first feature map into an E-Stage1 network module for two-fold downsampling to generate a second feature map with a size of 1*120*160. Then inputting the second feature map into an E-Stage2 network module for two-fold downsampling to generate a third feature map with a size of 1*60*80.
[0045] In an embodiment of the present application, the process of feature extraction performed on the visible light image after field of view matching includes: performing four-fold downsampling on the visible light image and inputting it into an E-Stage network module for two-fold downsampling to generate a fourth feature map with a size of one-eighth of the size of the visible light image; inputting the fourth feature map into the E-Stage network module for two-fold downsampling to generate a fifth feature map with a size of one-sixteenth of the size of the visible light image; performing weighted fusion on the third feature map and the fifth feature map to generate a weighted fusion feature map; inputting the weighted fusion feature map into the E-Stage network module for two-fold downsampling to generate a sixth feature map with a size of one-thirty-second of the size of the visible light image; inputting the sixth feature map into a dilated spatial pyramid pooling network to generate a multi-task feature map.
[0046] As shown in Figure 3 The leftmost side shows the process of feature extraction performed on the infrared image, which includes: using a two-fold downsampling convolution operation on the infrared image to generate a first feature map with a size of 1*240*320 after convolution, and then inputting the first feature map into an E-Stage1 network module for two-fold downsampling to generate a second feature map with a size of 1*120*160. Then inputting the second feature map into an E-Stage2 network module for two-fold downsampling to generate a third feature map with a size of 1*60*80.
[0047] It should be noted that the application designs a feature pyramid network that can simultaneously input visible light images and infrared images, and designs parallel infrared backbone networks and visible light backbone networks. For infrared images, only a lighter infrared backbone network branch is used for feature extraction, wherein the infrared backbone network branch includes two E-Stage network modules. For visible light images containing more feature data, a visible light backbone network branch is used to extract feature information, wherein the visible light backbone network branch includes two E-Stage network modules. The design of such parallel backbone structure takes into account the structural characteristics of infrared light images and visible light images. Among them, the channel number of the infrared image is less, the original infrared resolution is also lower, and the feature data contained is less, while the channel number of the visible light is more, and the feature information contained is more abundant. The number of network layers of the visible light backbone network branch and the infrared backbone network branch is the same, and the lightness of the infrared backbone network lies in that the channel number of the feature map output by the branch is only half of the channel number of the feature map output by the visible light backbone network branch. The infrared feature map (E-Stage2 output) extracted by the infrared backbone network branch and the visible light feature map (E-Stage4 output) extracted by the visible light backbone network branch are weighted and fused, and then a E-Stage module is used for feature extraction once again, so as to generate a fusion feature map containing infrared image features and visible light image features, and realize the balance of infrared light image feature information and visible light image feature information, balance the feature information of two different formats, and maximize the preservation of the unique feature information of the two images.
[0048] In an embodiment of the application, the E-Stage network module includes a main branch and a residual structure; wherein the main branch includes one or more ordinary convolution layers and one or more deformable convolution layers.
[0049] It is worth noting that the E-Stage network module is a network architecture designed and optimized by the present application. Figure 4 The structural diagram of the E-Stage network module in an embodiment of the application is shown. It should be noted that the backbone network plays a role in extracting image features in the feature pyramid. Through multiple convolution and pooling operations, the input image is gradually abstracted into different levels of feature representation, so that different scale feature information of the image can be fully represented and utilized. Among them, the input feature vector contains four parameters, which are batch size, channel number, height and width. For example, Figure 4As shown, in this embodiment, the input features are sequentially subjected to a 3*3 convolution layer, a deformable convolution layer and a 3*3 convolution layer to generate convolution features, then the input features are subjected to a 1*1 convolution and fused with the above-mentioned convolution features, and then the fused features are input into a 3*3 convolution layer and a 1*1 convolution layer to obtain a final feature vector containing rich feature information. After the E-Stage network module processing, the output feature image channel number will be doubled compared with the input feature image channel number, and the image size will be halved compared with the input feature image size.
[0050] Further, the E-Stage network module adopts a residual structure, and ordinary convolution and deformable convolution are used in the main branch, and the structure of a residual network is introduced. For the E-Stage network module creatively proposed in the present application, the advantages of simultaneously introducing deformable convolution and a residual network are that the advantages of both can be combined to further improve the performance of the model. Deformable convolution can learn specific deformation while learning the convolution kernel, thereby improving the adaptability of the model to complex structures and deformation in the image; and the residual network can help the model to learn deeper features, alleviate the gradient vanishing problem and accelerate the convergence of the model. Therefore, the combination of deformable convolution and a residual network can better process complex image data and improve the performance and generalization ability of the model.
[0051] Figure 5 A structural diagram of deformable convolution in an embodiment of the present application is shown. The structure of deformable convolution can be divided into two parts: the upper half is to generate offset of feature point position in x and y directions based on the input feature map using ordinary convolution, and the parameters for calculating the offset will be learned together with the network; the lower half is to calculate a new feature map after position offset based on the feature map and the offset, and then perform ordinary convolution on the new feature map, thereby realizing convolution on an arbitrary contour image.
[0052] Step S13: performing feature fusion on the multi-task feature map based on the feature pyramid network, and inputting the multi-task feature map into a multi-task learning model for training and prediction to respectively generate a region segmentation result and a three-dimensional detection frame.
[0053] In an embodiment of the present application, the process of generating a region segmentation result based on the feature pyramid network for the multi-task feature map includes: inputting the multi-task feature map into a first fusion layer and performing two times of upsampling to generate a first fusion feature; inputting the first fusion feature and the weighted fusion feature map into a second fusion layer and performing two times of upsampling to generate a second fusion feature; inputting the second fusion feature and the second feature map into a third fusion layer and performing two times of upsampling to generate a third fusion feature; inputting the third fusion feature and the first feature map into a fourth fusion layer and performing two times of upsampling to generate a fourth fusion feature; inputting the fourth fusion feature into a fifth fusion layer and performing two times of upsampling to generate a fifth fusion feature; and performing a dimension adjustment operation on the fifth fusion feature to generate a region segmentation result.
[0054] In an embodiment of the present application, as shown in Figure 2 The multi-task feature map output by the ASPP network is sequentially subjected to five times of D-Block 2 times upsampling network from bottom to top. The D-Block 2 times upsampling network refers to performing D-Block network on the input feature map first, and then performing 2 times upsampling. Figure 6 The structure diagram of the D-Block network in the FPN-neck in the embodiment is shown. The upsampling module in the FPN is usually referred to as a neck, which is used to connect feature maps of different levels and implement a feature pyramid network. The neck module usually includes upsampling, fusion, feature size adjustment and other operations, which are used to extract and integrate feature information of different scales to improve the performance of target detection and segmentation and other tasks.
[0055] As shown in Figure 6 The D-Block network proposed in the present application includes two branches. The first branch includes a 3*3 convolution layer, a 1*1 convolution layer, a 3*3 convolution layer and a 1*1 convolution layer. The second branch includes a 1*1 convolution layer. Then the results of the first branch and the second branch are weighted and fused to generate the output feature map of the final D-Block network. The size of the feature map does not change before and after the input of the D-Block, and the number of channels is reduced. Finally, the high-dimensional feature map is restored to the size of the visible light Figure 1 and the restored feature map is taken as the final output.
[0056] In an embodiment of the present application, the process of inputting the multi-task feature map into the multi-task learning model for training and prediction to generate a three-dimensional detection frame includes: constructing an auxiliary learning task; training and optimizing the multi-task learning model based on the loss function of the multi-task learning model and the loss function of the auxiliary learning task; and inputting the multi-task feature map into the trained multi-task learning model to generate a three-dimensional detection frame.
[0057] In an embodiment of the present application, the multi-task perception method based on multi-spectral images comprises multiple detection heads, including a region segmentation detection head, a three-dimensional frame detection head, and an auxiliary learning detection head. The region segmentation detection head is used to output a region segmentation result; the three-dimensional frame detection head is used to output a three-dimensional detection frame; and the auxiliary learning detection head is used to improve the training accuracy of a multi-task learning model of the three-dimensional detection frame. An additional auxiliary learning task is introduced during training, and the training accuracy of the model is improved by means of auxiliary learning of eight corner projection points of the three-dimensional frame. The auxiliary learning task is discarded during prediction, and therefore does not increase the time consumption of the training and inference process.
[0058] As shown in Figure 2 , the three-dimensional frame detection head comprises two processes, training and prediction. The prediction mainly includes the following five outputs: a heat map, a target depth, a target orientation, an offset of a two-dimensional center frame of the target in the image to a three-dimensional center frame, and dimensions of the three-dimensional frame. The heat map is used to provide the approximate position of the target center point, the depth is used to restore the three-dimensional shape of the target, the orientation is used to provide the pose of the target, the offset is used to accurately position the position of the target, and the dimensions determine the size of the target.
[0059] Figure 6 The flowchart of an embodiment of the multi-task perception method based on multi-spectral images of the present application is shown. In this embodiment, the following processes are included: first, the visible image and the infrared image are subjected to field of view matching operation, and the visible image and the infrared image subjected to the field of view matching operation are input into the multi-task network to generate a three-dimensional frame detection result and a drivable area segmentation result. A loss function of three-dimensional frame detection is calculated based on the three-dimensional frame detection result and a three-dimensional target detection label. A loss function of region segmentation is calculated based on the drivable area segmentation result and a semantic segmentation label. The total loss function of the entire multi-task perception network is calculated according to the loss function of three-dimensional frame detection and the loss function of region segmentation.
[0060] Further, the loss function of region segmentation comprises a cross-entropy loss function, and the specific calculation method is shown in formula 1.
[0061]
[0062] The calculation method of the loss function of three-dimensional frame detection is shown in formula 2.
[0063] Loss 3ddet =β1Loss heatmap +β2Loss dim +β3Loss depth+ β4Loss offset + β5Loss orientation
[0064] + β6Loss wh + β7Loss auxiliary (Formula 2)
[0065] wherein, Loss 3ddet represents a loss function of three-dimensional box detection, which is composed of seven parts, Loss heatmap represents a heat map loss, Loss dim represents a dimension loss, Loss depth represents a depth loss, Loss offset represents a two-dimensional box center offset loss, Loss orientation represents a direction loss, Loss wh represents a length-width loss of a two-dimensional box, Loss auxiliary represents an auxiliary learning loss. β1 to β7 represent the weight of each loss function in the loss function of three-dimensional box detection. The weight can be flexibly adjusted according to actual needs.
[0066] In an embodiment of the present application, the heat map loss and the auxiliary learning loss adopt a focal loss function (Focal Loss). The focal loss function reduces the weight of easy-to-classify samples and improves the attention of the model to difficult-to-classify samples, so as to solve the problem of class imbalance. The two-dimensional box center offset loss and the direction loss adopt an absolute value loss function (L1 Loss). The absolute value loss function is used to measure the absolute error between the predicted value and the true value. At the same time, the depth loss adopts an uncertainty loss function (Uncertainty Loss). The uncertainty loss function encourages the model to give higher uncertainty to samples that are difficult to predict, thereby improving the robustness of the model to noise and outliers. The dimension loss adopts an IOU-oriented loss function. The intersection over union (IOU) refers to the ratio between the predicted box and the true box. At this time, the calculation method of the loss function of three-dimensional box detection is shown in Formula 3 and Formula 4:
[0067]
[0068] The intersection and union regions of the boundary box in the dimension are calculated by the width w, the height h and the length l of the boundary box. In the dimension loss, the IOU-oriented loss function gives greater weight to the parameters that contribute more to the IOU. Therefore, according to the derivative ratio, each item in w, h and l is normalized, so as to map the width, height and length of the boundary box to the same order of magnitude, so as to generate a new loss function containing weights. The normalized dimension loss function is shown in Formula 4, and its expansion is shown in Formula 5.
[0069]
[0070]
[0071] wherein y represents a true value, represents a predicted value, w represents a width of the true value, represents a width of the predicted value; h represents a height of the true value, represents a height of the predicted value; l represents a length of the true value, represents a length of the predicted value. Here, the standard loss is normalized by dividing the true value, and multiplied by a compensation weight composed of w, h, and l. The formula of the compensation weight weight is shown in Formula 6. The dimension loss function with the compensation weight is equivalent to adding a normalization process to the original loss function, and uniformly scaling with the same weight. This makes the width, height, and length mutually related when using the gradient descent algorithm for optimization calculation, thereby avoiding the problem of uneven optimization caused by individual optimization.
[0072]
[0073] Further, the total loss function of the entire multi-task perception network is shown in Formula 7, wherein a represents the weight of the region segmentation loss function, and β represents the weight of the three-dimensional frame loss function. The region segmentation loss function and the three-dimensional frame loss function jointly constitute the total loss function in the model, and the weights can be adjusted based on actual needs. Exemplarily, since the difficulty of three-dimensional frame detection is higher than that of region segmentation detection, the weight corresponding to the three-dimensional frame detection is set to 0.6, and the weight corresponding to the region segmentation detection is set to 0.4.
[0074] Loss total = αLoss seg + βLoss 3ddet (Formula 7)
[0075] In an embodiment of the present application, the above-mentioned semantic segmentation label and three-dimensional target detection label are preset label data sets, respectively used for calculating semantic segmentation loss and three-dimensional frame detection loss, and the total loss is the weighted sum of the semantic segmentation loss and the three-dimensional frame detection loss. The format of the semantic segmentation label is a single-channel image consistent with the resolution of the visible light image, wherein the pixel value of "1" represents a drivable area, and the pixel value of "0" represents a non-drivable area. The three-dimensional target detection label includes camera intrinsic matrix and coordinate information of the three-dimensional frame.
[0076] Further, the semantic segmentation label can be automatically labeled by an image labeling tool, including but not limited to LabelMe, CVAT, etc. The three-dimensional target detection label is labeled by joint calibration, and the process of joint calibration includes camera intrinsic calibration, camera-laser radar joint calibration, and 3D target frame labeling and verification. The camera intrinsic matrix K is solved by an optimization algorithm to perform camera intrinsic calibration; the conversion relationship between the camera coordinate system and the laser radar coordinate system is solved by an optimization algorithm to perform joint calibration; finally, the laser radar point cloud is labeled in 3D, and the 3D target frame is converted to the camera coordinate system to complete the labeling of the three-dimensional target detection label.
[0077] In an embodiment of the present application, the auxiliary learning task includes multiple auxiliary learning tasks such as octagonal point projection key point heat map. Taking the octagonal point projection key point heat map as an example, the octagonal point projection key point heat map is a heat map used to represent the projection position of the eight corner points of an object on a two-dimensional image, and the process includes: defining the eight corner points of a three-dimensional frame; projecting the three-dimensional corner points onto the two-dimensional image plane based on the camera intrinsic parameter; and projecting each corner point onto the heat map in a Gaussian distribution manner on the two-dimensional image. The value of each point on the heat map represents the probability or confidence that the position is a corner point.
[0078] In an embodiment of the present application, the multi-task perception method based on multi-spectral images is applied to actual automatic driving perception, so as to perceive the obstacles in the area near the vehicle and perceive the distance between the obstacles and the vehicle. At the same time, it is also judged whether the road is drivable, for example, although there is no obstacle in the non-motor vehicle lane and the like, it is still not drivable. In this embodiment, the obstacles include cars, large vehicles (including buses, trucks, vans, etc.), pedestrians, and riders (people riding bicycles, people riding motorcycles, etc.).
[0079] In this embodiment, the size of the input visible light image (RGB) is (n, 3, 960, 1280), and the size of the input infrared image is (n, 1, 480, 640), where n represents the batch size. First, a 3*3 convolution operation is performed on the infrared image, and the output feature image with a dimension of (n, 16, 240, 320) is obtained, and then the E-Stage network module is used twice to generate an infrared backbone network branch feature image with a dimension of (n, 64, 60, 80).
[0080] Further, the visible light image is subjected to a 3*3 convolution operation, and a feature image with a dimension of (n, 32, 240, 320) is output, and then the feature image is input into two E-Stage network modules to generate a visible light backbone network branch feature image with a dimension of (n, 128, 60, 80); the visible light feature image and the infrared feature image are subjected to a weighted fusion module to obtain a weighted fusion feature image with a dimension of (n, 192, 60, 80), and after being subjected to an E-Stage5 network module, the output dimension of the feature image is a fusion feature image with a dimension of (n, 384, 30, 40). The fusion feature image with a size of (n, 384, 30, 40) is input into an ASPP module.
[0081] Specifically, in the ASPP module, the input fusion feature image is divided into multiple groups, and convolution operation and pooling operation are respectively performed on each group to generate the same size feature map containing multiple different scale features. Subsequently, the same size feature maps of different scale features are spliced and superimposed, and a 1*1 convolution kernel is used to adjust the channel number of the fusion feature map generated after feature superposition to 256, and the dimension of the fusion feature map finally output by the backbone network is (n, 256, 30, 40).
[0082] Further, the backbone network with a dimension of (n, 256, 30, 40) is subjected to multiple upsampling to output a dimension of (n, seg_class, 960, 1280) through a segmentation detection head, wherein seg_class represents the class of the segmentation detection head, and exemplarily, seg_class is set to 2 and is used for calculation of a region segmentation loss function through the class of the segmentation detection head. At the same time, the feature map with a dimension of (n, 256, 30, 40) output by the ASPP module continues to pass through several convolution layers and generates multiple branches. The branches include: a heat map with a size of (n, det_class, 30, 40), a two-dimensional frame with a size of (n, 2, 30, 40), a depth with a size of (n, 1, 30, 40), a bias with a size of (n, 2, 30, 40), a direction with a size of (n, 2, 30, 40), and a dimension with a size of (n, 3, 30, 40). At this time, the eight corner points of the auxiliary learning branch projected on the heat map have a dimension of (n, 8, 30, 40) and are used for calculating a loss function for three-dimensional frame detection and a training process. In prediction, the multispectral image is input into the trained model to generate a drivable area and a 3D detection frame. And based on the three-dimensional detection frame and the camera intrinsic matrix, a two-dimensional detection frame is calculated.
[0083] In the embodiments of the present application, the same items or similar items with basically the same functions and effects are distinguished by using "first", "second", and the like. For example, the first feature map and the second feature map are only used to distinguish different feature maps, and do not limit the order. Those skilled in the art can understand that the "first feature map" and the "second feature map" do not limit the number and execution order, and the "first", "second" and the like do not necessarily mean different.
[0084] It should be noted that in the embodiments of the present application, "exemplary" or "for example" means an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of "exemplary" or "for example" is intended to present the relevant concept in a specific manner.
[0085] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b or c can represent a, b, c, a-b, a-c, b-c or a-b-c, where a, b and c can be single or multiple.
[0086] Figure 7 is a schematic block diagram of a multi-task perception device 700 based on a multi-spectral image provided by the embodiments of the present application. As shown in the figure, the device includes: Figure 7
[0087] The image acquisition module 701 is configured to acquire a visible light image and an infrared image of a target region respectively, and perform field of view matching on the visible light image and the infrared image.
[0088] The feature extraction module 702 is configured to perform feature extraction on the visible light image and the infrared image after field of view matching based on a feature pyramid network, to generate a multi-task feature map.
[0089] The region segmentation prediction module 703 is configured to perform feature fusion on the multi-task feature map based on the feature pyramid network, to generate a region segmentation result.
[0090] The detection box prediction module 704 is configured to input the multi-task feature map into a multi-task learning model for training and prediction, so as to generate a three-dimensional detection box.
[0091] It should be understood that the specific process of each module performing the corresponding steps described above has been described in detail in the method embodiments described above, and for the sake of brevity, will not be repeated here.
[0092] It should also be understood that the division of the modules in the embodiments of the present application is illustrative, and is only a logical functional division. In actual implementation, there can be another division manner. In addition, each functional module in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.
[0093] Figure 8 is a schematic block diagram of an electronic terminal for implementing the multi-task perception method based on a multi-spectrum image provided by the embodiments of the present application. As shown in Figure 8 , the computer device includes at least one processor 801, a memory 802, at least one network interface 803, and a user interface 805. Each component in the device is coupled together through a bus system 804. It can be understood that the bus system 804 is used to realize the connection and communication between the components. In addition to including a data bus, the bus system 804 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all kinds of buses are marked as a bus system in Figure 8 .
[0094] The user interface 805 can include a display, a keyboard, a mouse, a trackball, a click gun, a key, a button, a touchpad, or a touch screen, etc.
[0095] It can be understood that the memory 802 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory). The memory described in the embodiments of the present application is intended to include but not limited to these and any other suitable categories of memory.
[0096] The memory 802 in the embodiments of the present application is configured to store various types of data to support the operation of the electronic terminal 800. Examples of the data include any executable programs for operating on the electronic terminal 800, such as an operating system 8021 and application programs 8022. The operating system 8021 includes various system programs, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks. The application programs 8022 can include various application programs, such as a media player, a browser, and the like, for implementing various application services. The method 88 provided by the embodiments of the present application can be included in the application programs 8022.
[0097] The method disclosed in the embodiments of the present application can be applied to the processor 801 or implemented by the processor 801. The processor 801 can be an integrated circuit chip having a processing capability of signals. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 801. The processor 801 described above can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, and the like. The processor 801 can implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application. The general-purpose processor 801 can be a microprocessor or any conventional processor, and the like. In combination with the steps of the accessory optimization method provided by the embodiments of the present application, the steps can be directly embodied as hardware decoding processor execution, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory. The processor reads the information in the memory and combines the hardware to complete the steps of the above method.
[0098] In the exemplary embodiments, the electronic terminal 800 can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), and the like, for executing the above method.
[0099] According to the method provided by the embodiments of the present application, the present application further provides a computer program product, which comprises computer program code, when the computer program code runs on a computer, so that the computer executes Figures 1 to 6The multi-task perception method based on multi-spectrum images of any of the illustrated embodiments.
[0100] According to the method provided in the embodiments of the present application, the present application further provides a computer readable storage medium, which stores program codes, when the program codes are run on a computer, the computer is caused to execute Figures 1 to 6 The multi-task perception method based on multi-spectrum images of any of the illustrated embodiments.
[0101] The terms "component," "module," "system," and the like are used in the present description to refer to computer-related entities, hardware, firmware, a combination of hardware and software, software, or the like. For example, a component can be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be a component. One or more components can reside within a process and / or thread of execution and a component can be localized, co-resident, and / or distributed among one computer and / or among two or more computers. Also, these components can execute from various computer readable media having various data structures stored thereon. The components can communicate by way of local and / or remote processes such as in accordance with a signal having one or more data packets (e.g., data from one component interacting with another component in a local system, distributed system, and / or across a network such as the Internet with other systems via the signal).
[0102] Those of skill in the art would understand that the various illustrative logical blocks and steps described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. The choice of whether to implement the described functionality in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0104] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0105] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0106] In addition, the functional units in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0107] In the above embodiments, the functions of the functional units can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the whole or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium, or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website site, computer, server or data center to another website site, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as high-density digital video disc (digital video disc, DVD), or semiconductor media (such as solid state disk (solid state disk, SSD), etc.
[0108] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0109] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0110] In summary, the present application provides a multi-task perception method, device, medium, program product and terminal based on multi-spectral images. The present application proposes a multi-task learning model combining infrared images and visible light images, aiming to improve the performance of the model in extreme environments such as night, rainy day and heavy fog. Specifically, by designing a parallel feature fusion backbone network structure, image features are effectively extracted, and an improved U-Net and DeepLabV3 design is used on the semantic segmentation branch, optimizing the multi-scale feature extraction process. In addition, the model introduces a multi-task learning mechanism of 3D monocular detection and semantic segmentation, improving the learning efficiency and overall accuracy of the model, while reducing the risk of overfitting. The auxiliary learning branch further improves the robustness of the model in three-dimensional detection tasks without additional labeling. In addition, the present application is an end-to-end multi-task processing method, which only needs image and camera intrinsic parameter input to simultaneously process key tasks of autonomous driving such as drivable area determination, two-dimensional target detection and three-dimensional target detection. Compared with single-task models, the inference performance and processing speed are significantly improved. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has high industrial utilization value.
[0111] The above embodiments are only illustrative of the principles of the present application and its effects, and are not intended to limit the present application. Any modification or change made by any person skilled in the art without departing from the spirit and scope of the present application shall be covered by the claims of the present application.
Claims
1. A multi-task perception method based on multispectral images, characterized in that, include: The visible light image and infrared image of the target area are acquired respectively, and field-of-view matching is performed on the visible light image and the infrared image; Based on the feature pyramid network, feature extraction is performed on the visible light image and the infrared image after field-of-view matching to generate a multi-task feature map; Based on the feature pyramid network, feature fusion is performed on the multi-task feature map, and the multi-task feature map is input into the multi-task learning model for training and prediction to generate region segmentation results and 3D detection boxes respectively.
2. The multi-task perception method based on multispectral images according to claim 1, characterized in that, The process of performing feature extraction on the infrared image after field-of-view matching includes: The infrared image is subjected to a 2x downsampling convolution operation to generate a first feature map; The first feature map is input into the E-Stage network module for double downsampling to generate a second feature map with a size one-eighth of the visible light image size; The second feature map is input into the E-Stage network module for double downsampling to generate a third feature map with a size one-sixteenth of the visible light image size.
3. The multi-task perception method based on multispectral images according to claim 2, characterized in that, The process of performing feature extraction on the visible light image after field-of-view matching includes: The visible light image is downsampled four times and then input into the E-Stage network module for downsampling twice to generate a fourth feature map with a size one-eighth that of the visible light image. The fourth feature map is input into the E-Stage network module for double downsampling to generate a fifth feature map with a size one-sixteenth of the visible light image size; The third feature map and the fifth feature map are weighted and fused to generate a weighted fused feature map; The weighted fused feature map is input into the E-Stage network module for double downsampling to generate a sixth feature map with a size one-thirty-second of the visible light image size; The sixth feature map is input into the hollow space pyramid pooling network to generate a multi-task feature map.
4. The multi-task perception method based on multispectral images according to claim 3, characterized in that, The process of performing feature fusion on the multi-task feature map based on the feature pyramid network to generate region segmentation results includes: The multi-task feature map is input into the first fusion layer and upsampled by a factor of two to generate the first fused feature. The first fusion feature and the weighted fusion feature map are input into the second fusion layer and upsampled by a factor of two to generate the second fusion feature; The second fusion feature and the second feature map are input into the third fusion layer and upsampled by a factor of two to generate the third fusion feature; The third fusion feature and the first feature map are input into the fourth fusion layer and upsampled by a factor of two to generate the fourth fusion feature; The fourth fusion feature is input into the fifth fusion layer and upsampled by a factor of two to generate the fifth fusion feature; The fifth fusion feature is subjected to dimensional adjustment to generate a region segmentation result.
5. The multi-task perception method based on multispectral images according to claim 2, characterized in that, The E-Stage network module includes a main branch and a residual structure; wherein the main branch includes one or more ordinary convolutional layers and one or more deformable convolutional layers.
6. The multi-task perception method based on multispectral images according to claim 1, characterized in that, The process of inputting the multi-task feature map into a multi-task learning model for training and prediction to generate a 3D detection box includes: Construct auxiliary learning tasks; The multi-task learning model is trained and optimized based on the loss function of the multi-task learning model and the loss function of the auxiliary learning task. The multi-task feature map is input into the trained multi-task learning model to generate a 3D detection box.
7. A multi-task sensing device based on multispectral images, characterized in that, include: Image acquisition module: used to acquire visible light and infrared images of the target area respectively, and perform field-of-view matching on the visible light and infrared images; Feature extraction module: used to perform feature extraction on the visible light image and the infrared image after field-of-view matching based on the feature pyramid network, so as to generate a multi-task feature map; Region segmentation prediction module: used to perform feature fusion on the multi-task feature map based on the feature pyramid network to generate region segmentation results; Detection box prediction module: used to input the multi-task feature map into the multi-task learning model for training and prediction, so as to generate a three-dimensional detection box.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-task perception method based on multispectral images as described in any one of claims 1 to 6.
9. A computer program product, characterized in that, The computer program product includes computer program code, which, when run on a computer, enables the computer to implement the multi-task perception method based on multispectral images as described in any one of claims 1 to 6.
10. An electronic terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the multi-task perception method based on multispectral images as described in any one of claims 1 to 6.