Training method of multi-task processing model, multi-task processing method and related device
By adopting a multi-tasking model, a shared backbone network and an upsampling module in the automatic parking system, the problem of computing resource consumption and processing time of target detection and feasible area segmentation tasks is solved, and more efficient computing and lower complexity deployment is achieved.
Patent Information
- Application Number
- CN202510596637.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120126100A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly relates to a training method for a multi-task processing model, a multi-task processing method, and related devices. Background Art
[0002] During the automatic parking process, the vehicle needs to accurately perceive and identify the surrounding environment to avoid obstacles. Currently, in the automatic parking system, obstacle avoidance perception usually adopts a single-task processing method, that is, target detection and drivable area segmentation are performed separately. Among them, target detection is to capture the images around the vehicle using a fish-eye camera, and use deep learning algorithms or traditional image processing techniques to perform target detection on the fish-eye images; drivable area segmentation is to segment the free areas around the vehicle based on the fish-eye images to help the vehicle identify the drivable areas.
[0003] In the case of single-task processing, target detection and drivable area are usually processed by two independent neural networks. However, separately executing these two tasks requires more computing resources, resulting in increased power consumption and processing time. Summary of the Invention
[0004] In view of at least one of the above technical problems existing in the prior art, the present application is proposed. The present application can reduce the consumption of computing resources and shorten the multi-task processing time.
[0005] In a first aspect, an embodiment of the present application provides a training method for a multi-task processing model. The multi-task processing model includes: a backbone network, a neck network, a detection head, and a segmentation head; the neck network includes: an upsampling module and a downsampling module; the method includes: Obtain fish-eye image samples and sample labels; Input the fish-eye image samples into the backbone network, so that the backbone network extracts features from the fish-eye image samples and outputs first feature maps of multiple sizes; and splice the first feature maps of multiple sizes; Input the feature map formed after splicing the first feature maps of multiple sizes and the first feature maps of multiple sizes into the upsampling module, so that the upsampling module splices the upsampled feature map with the first feature map of the corresponding size and outputs second feature maps of multiple sizes; Input the second feature maps of multiple sizes into the downsampling module, so that the downsampling module splices the downsampled feature map with the second feature map of the corresponding size and outputs third feature maps of multiple sizes; Input the third feature maps of multiple sizes into the detection head, so that the detection head detects the targets in the fish-eye image samples and outputs target detection results; Input the second feature maps of the multiple sizes into the segmentation head, so that the segmentation head identifies the drivable area in the fisheye image sample and outputs an area detection result; Adjust the parameters of the multi-task processing model based on the object detection result, the area detection result and the sample label.
[0006] In a second aspect, an embodiment of the present application provides a multi-task processing method, including: Obtain a fisheye image; Input the fisheye image into a trained multi-task processing model to obtain an object detection result and an area detection result of the fisheye image; Wherein, the multi-task processing model is trained based on the method provided in the above embodiment.
[0007] In a third aspect, an embodiment of the present application provides a training device for a multi-task processing model. The multi-task processing model includes: a backbone network, a neck network, a detection head and a segmentation head; the neck network includes: an upsampling module and a downsampling module; the device includes: A first acquisition module configured to acquire a fisheye image sample and a sample label; A feature extraction module configured to input the fisheye image sample into the backbone network, so that the backbone network extracts features from the fisheye image sample and outputs first feature maps of multiple sizes; and splice the first feature maps of the multiple sizes; A feature fusion module configured to input the feature map formed by splicing the first feature maps of the multiple sizes and the first feature maps of the multiple sizes into the upsampling module, so that the upsampling module splices the upsampled feature map with the first feature map of the corresponding size and outputs second feature maps of multiple sizes; input the second feature maps of the multiple sizes into the downsampling module, so that the downsampling module splices the downsampled feature map with the second feature map of the corresponding size and outputs third feature maps of multiple sizes; A task processing module configured to input the third feature maps of the multiple sizes into the detection head, so that the detection head detects objects in the fisheye image sample and outputs an object detection result; input the second feature maps of the multiple sizes into the segmentation head, so that the segmentation head identifies the drivable area in the fisheye image sample and outputs an area detection result; A training module configured to adjust the parameters of the multi-task processing model based on the object detection result, the area detection result and the sample label.
[0008] In a fourth aspect, an embodiment of the present application provides a multi-task processing device, including: A second acquisition module configured to acquire a fisheye image; A prediction module, configured to input the fisheye image into a trained multi-task processing model to obtain the object detection result and the region detection result of the fisheye image; Wherein, the multi-task processing model is trained based on the method provided in the above embodiment.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program / instructions are executed by a processor, the method described in any of the above embodiments is implemented.
[0010] The training method, multi-task processing method and related devices of the multi-task processing model provided by the embodiments of the present application. For the object detection task and the drivable area segmentation task, the backbone network and the upsampling module in the multi-task processing model are shared, reducing redundant calculations. Compared with being processed by two independent neural networks, the present application can reduce the consumption of computing resources and shorten the multi-task processing time. Unifying the two tasks into one model reduces the deployment complexity and is applicable to embedded devices and low-computing-power platforms. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.
[0012] Figure 1 It is a flowchart of a training method of a multi-task processing model provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a multi-task processing model provided by an embodiment of the present application; Figure 3 It is a flowchart of a multi-task processing method provided by an embodiment of the present application; Figure 4 It is a schematic diagram of a training device of a multi-task processing model provided by an embodiment of the present application; Figure 5 It is a flowchart of a multi-task processing device provided by an embodiment of the present application. Detailed Embodiments
[0013] To enable those skilled in the art to better understand the technical solutions of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0014] In the prior art, the object detection task and the drivable area segmentation task are usually processed by two independent neural networks. For example, separately executing these two tasks by an object detection model and a drivable area segmentation model respectively requires more computing resources, resulting in an increase in power consumption and processing time. Especially on low-computing-power chips, it is difficult to meet the real-time requirements. The prior art cannot make full use of the shared information between models, such as the sharing of image features, which increases the burden on the chip. For low-computing-power chips, it is difficult for the parking system to meet the real-time requirements when running multiple deep learning models simultaneously.
[0015] In view of this, as Figure 1 shown, the embodiments of the present application provide a training method for a multi-task processing model. Among them, as Figure 2 shown, the multi-task processing model includes: a backbone network, a neck network, a detection head, and a segmentation head; the neck network includes: an upsampling module and a downsampling module; the method includes: Step 101: Obtain fisheye image samples and sample labels.
[0016] The fisheye image samples can be synchronously obtained by a four-way surround fisheye camera to ensure that the time information of different perspectives captured at the same moment is consistent. During the preprocessing process, the original fisheye image can be scaled, that is, by changing the width and height (pixel size) of the image, the resolution or display size of the image can be adjusted.
[0017] The sample labels include detection labels corresponding to the object detection task and segmentation labels corresponding to the drivable area segmentation.
[0018] Step 102: Input the fisheye image samples into the backbone network so that the backbone network extracts features from the fisheye image samples and outputs first feature maps of multiple sizes; and splice the first feature maps of the multiple sizes.
[0019] The backbone network can be the backbone network of an object detection model or improved from the backbone network of an object detection model. The backbone network is used to extract features.
[0020] Step 103: Input the feature map formed by splicing the first feature maps of the multiple sizes and the first feature maps of the multiple sizes into the upsampling module, so that the upsampling module splices the upsampled feature map with the first feature map of the corresponding size, and outputs the second feature maps of multiple sizes.
[0021] The upsampling module and the downsampling module are used to fuse features at different levels.
[0022] Step 104: Input the second feature maps of the multiple sizes into the downsampling module, so that the downsampling module splices the downsampled feature map with the second feature map of the corresponding size, and outputs the third feature maps of multiple sizes.
[0023] Step 105: Input the third feature maps of multiple sizes into the detection head, so that the detection head detects the targets in the fisheye image sample and outputs the target detection result.
[0024] The detection head is used to specifically execute the target detection task, and the detection head of an existing target detection model can be adopted.
[0025] Step 106: Input the second feature maps of multiple sizes into the segmentation head, so that the segmentation head identifies the drivable area in the fisheye image sample and outputs the area detection result.
[0026] The segmentation head is used to execute the drivable area segmentation task, and a fully convolutional network or an encoder-decoder structure can be adopted.
[0027] Step 107: Adjust the parameters of the multi-task processing model based on the target detection result, the area detection result, and the sample label.
[0028] In the embodiment of the present application, for the target detection task and the drivable area segmentation task, the backbone network and the upsampling module in the multi-task processing model are shared, reducing redundant calculations. Compared with being processed by two independent neural networks, the present application can reduce the consumption of computing resources and shorten the multi-task processing time. Unifying the two tasks into one model reduces the deployment complexity and is applicable to embedded devices and low-computing-power platforms.
[0029] In an embodiment of the present application, the backbone network includes: a focusing module, a plurality of convolutional layers, a plurality of convolutional-to-feature fusion modules, and a pooling module; The pooling module includes: an initial convolutional layer, a max pooling and splicing layer, and a final convolutional layer; Among them, the pooling kernel used in the max pooling and splicing layer is 3x3.
[0030] This embodiment improves the backbone network of the existing YOLOv8 to adapt to low-compute chips. In this embodiment, the convolutional layers in the existing YOLOv8 are replaced with focus modules, and the focus module is the Focus module, which is used to split image channels to reduce the computational load and maintain the compactness of features.
[0031] At the same time, in order to improve the feature expression ability and reduce the occupation of computing resources, this embodiment replaces the spatial pyramid pooling module of YOLOv8 with a simplified spatial pyramid pooling module, and replaces the traditional large-kernel pooling layer in the simplified spatial pyramid pooling module with multiple 3x3 pooling layers to obtain a pooling module, realizing more refined multi-scale feature aggregation. Among them, the convolutional to feature fusion module is the C2f module.
[0032] By improving the backbone network of YOLOv8, this embodiment can reduce the computational load of the multi-task processing model and make it adapt to low-compute platforms. In actual application scenarios, the backbone network of YOLOv8 or the backbone network of other object detection models can also be directly used.
[0033] In one embodiment of the present application, the upsampling module includes: an upsampling layer, a first splicing layer, and a first C3 module; The upsampling module splices the upsampled feature map with the first feature map of the corresponding size and outputs second feature maps of multiple sizes, including: The upsampling layer upsamples the feature map formed by splicing the first feature maps of multiple sizes, the first splicing layer splices the upsampled feature map with the first feature map of the fourth level, and the first C3 module processes the spliced feature map to obtain the second feature map of the fourth level; The upsampling layer upsamples the second feature map of the fourth level, the first splicing layer splices the upsampled feature map with the first feature map of the third level, and the first C3 module processes the spliced feature map to obtain the second feature map of the third level; The upsampling layer upsamples the second feature map of the third level, the first splicing layer splices the upsampled feature map with the first feature map of the second level, and the first C3 module processes the spliced feature map to obtain the second feature map of the second level; The upsampling layer upsamples the second feature map of the second level, the first splicing layer splices the upsampled feature map with the first feature map of the first level, and the first C3 module processes the spliced feature map to obtain the second feature map of the first level; Among them, the sizes of the first feature maps and the second feature maps corresponding to the first level, the second level, the third level, and the fourth level decrease in sequence.
[0034] The first level, the second level, the third level, and the fourth level respectively correspond to the P1 level, the P2 level, the P3 level, and the P4 level. C3 is a CSP (Cross Stage Partial) bottleneck with three convolutional layers. The concatenated feature map is input into the first C3 module to further extract more detailed features through the first C3 module, thereby improving the accuracy of object detection and drivable area segmentation.
[0035] In existing object detection tasks, usually only two upsamplings are performed. However, in this embodiment, considering that the drivable area segmentation task requires four upsamplings, instead of performing two upsamplings in the object detection branch and four upsamplings in the drivable area segmentation branch, it is better to directly perform four upsamplings, so that the results output by the four upsamplings can be used in both branches.
[0036] In this embodiment, not only does the drivable area segmentation task meet the requirement of four upsamplings, but the object detection task can also increase two upsamplings to improve the detection accuracy. In addition, multiple successive upsampling and concatenation processes can effectively combine high-level features and low-level features to achieve the recognition of objects of different sizes.
[0037] In an actual application scenario, the upsampling module may also only include: an upsampling layer and a first concatenation layer, that is, without the first C3 module.
[0038] In an embodiment of the present application, the downsampling module further includes: a downsampling layer, a second concatenation layer, and a second C3 module; The downsampling module concatenates the feature map obtained by downsampling with the second feature map of the corresponding size and outputs third feature maps of multiple sizes, including: The downsampling layer downsamples the second feature map of the first level, the second concatenation layer concatenates the feature map obtained by downsampling with the second feature map of the second level, and the second C3 module processes the concatenated feature map to obtain the third feature map of the second level; The downsampling layer downsamples the third feature map of the second level, the second concatenation layer concatenates the feature map obtained by downsampling with the second feature map of the third level, and the second C3 module processes the concatenated feature map to obtain the third feature map of the third level; The downsampling layer downsamples the third feature map of the third level, the second concatenation layer concatenates the feature map obtained by downsampling with the second feature map of the fourth level, and the second C3 module processes the concatenated feature map to obtain the third feature map of the fourth level; The downsampling layer downsamples the third feature map of the fourth level, the second splicing layer splices the feature map obtained by downsampling with the second feature map of the fifth level, and the second C3 module processes the spliced feature map to obtain the third feature map of the fifth level; Among them, the sizes of the third feature maps corresponding to the second level, the third level, the fourth level, and the fifth level decrease in sequence. The fifth level corresponds to the P5 level.
[0039] For the object detection branch, after upsampling, it is also necessary to perform downsampling step by step. This process adjusts the feature resolution step by step, enabling the neck network to integrate feature information at multiple levels. Similar to the upsampling process, the spliced feature map is input into the second C3 module to further extract more detailed features through the second C3 module, thereby improving the accuracy of object detection. The structures of the first C3 module and the second C3 module can be the same or different, and the structures of the first splicing layer and the second splicing layer can be the same or different.
[0040] The downsampling module can output third feature maps of different sizes such as the P2 level, the P3 level, the P4 level, and the P5 level. However, in actual application scenarios, considering the requirements of low-computing-power platforms, the third feature maps input to the detection head can only include the sizes corresponding to the P3 level, the P4 level, and the P5 level. Since each third feature map has already integrated the features of different levels respectively, the detection head can achieve the detection of objects of different sizes.
[0041] In an embodiment of the present application, the detection head includes: a bounding box regression module and a classification module; both the bounding box regression module and the classification module include multiple convolutional layers; The detection head detects the objects in the fish-eye image sample and outputs the object detection result, including: The bounding box regression module outputs the bounding box detection result of the object, and the classification module outputs the class detection result of the object.
[0042] In this embodiment, the detection head adopts an anchor-free design, aiming to predict the bounding box coordinates and class probabilities of the object from the multi-scale feature maps. Compared with the traditional anchor-based method, the anchor-free design significantly reduces the complexity of hyperparameter adjustment by eliminating the dependence on predefined anchors, and at the same time improves the generalization ability of the model to different object sizes and input resolutions. The detection head receives the third feature maps (usually three scales, such as P3, P4, P5) from the downsampling module and generates the detection result through parallel convolutional branches.
[0043] The detection head includes two parallel convolutional branches: a bounding box regression module (cv2) and a classification module (cv3).
[0044] Bounding box regression module (cv2): This branch maps the input third feature map to a tensor of shape (B, 4 * reg_max, H, W) through a series of convolutional operations, where B is the batch size, reg_max is the number of discretization channels (default value is 16), and H and W are the height and width of the feature map respectively. cv2 consists of three layers: two convolutional layers with activation functions (Conv(x, c2, 3)) and one convolutional layer without activation (nn.Conv2d(c2, 4 * reg_max, 1)), which are responsible for predicting the distributions of the four edges (left, top, right, bottom) of the bounding box. Among them, the number of output channels in Conv(x, c2, 3) and the number of input channels in nn.Conv2d(c2, 4 * reg_max, 1) are both c2.
[0045] Classification module (cv3): This branch maps the feature map to a tensor of shape (B, nc, H, W), where nc is the number of classes. Similar to cv2, cv3 includes two convolutional layers (Conv(x, c3, 3)) and one output layer (nn.Conv2d(c3, nc, 1)), which are used to predict the class of each grid point. Among them, the number of output channels in (Conv(x, c3, 3) and the number of input channels in nn.Conv2d(c3, nc, 1) are both c3.
[0046] In addition, the detection head can also include a DFL (Distribution Focal Loss) module, which is used to decode the final coordinate values from the bounding box distribution vector. If reg_max = 1, DFL degenerates to a simple identity mapping (nn.Identity()).
[0047] The forward propagation of the detection head is different according to the training and inference phases.
[0048] Training phase: The input list of feature maps x (containing third feature maps of multiple sizes) is processed by cv2 and cv3 respectively, and then concatenated along the channel dimension into a tensor of shape (B, no, H, W), where no = nc + 4 * reg_max, and it is directly returned for loss calculation.
[0049] Inference phase: Feature processing: Apply cv2 and cv3 to the third feature map of each size to generate bounding box and class features, and after concatenation, flatten them into (B, no, N), where N is the total number of all grid points.
[0050] Dynamic anchor generation: If the input shape changes or the dynamic mode is enabled, the make_anchors function is used to dynamically generate anchors and strides according to the stride and grid offset of the third feature map.
[0051] Output decoding: The flattened tensor is split into a bounding box part and a class part. The box is decoded into actual coordinate values through the DFL module, and the cls is converted into class probabilities through the sigmoid function, and finally concatenated into an output of (B, N, 4 + nc). The box is used to represent the position and size of the target object. The cls is the output of the classification task, representing the prediction score of the model for the class.
[0052] DFL is used to optimize bounding box regression. Traditional bounding box prediction directly regresses coordinate values, while DFL treats each coordinate as a discrete distribution (a vector of length reg_max), and calculates the final coordinate by learning the expected value of the distribution. This method not only improves the prediction accuracy but also enhances the detection ability for small targets.
[0053] For the object detection task, this embodiment adopts a method of multi-scale feature fusion and adaptive anchor box generation, enabling the model to capture the bounding box and class information of targets of different sizes. The detection head fuses multi-layer information from each feature map layer by layer, and combines distance-based distribution regression to optimize the accuracy of predicting the bounding box, improving the detection accuracy and the ability to distinguish small targets in complex backgrounds.
[0054] In an embodiment of the present application, the segmentation head includes: an initial convolutional layer, an upsampling layer, and subsequent convolutional layers; The segmentation head identifies the drivable area in the fisheye image sample and outputs the area detection result, including: The initial convolutional layer maps the second feature maps of multiple sizes to a preset number of channels through convolutional operations, the upsampling layer increases the resolution of the second feature map and decreases the number of channels of the second feature map based on transposed convolution, and the subsequent convolutional layers decrease the number of channels of the second feature map through two convolutional operations to obtain the area detection result.
[0055] The segmentation head focuses on the semantic segmentation task, aiming to generate pixel-level segmentation masks from the feature maps. It converts the low-resolution feature maps into high-resolution mask outputs through upsampling and convolutional operations, supporting the precise segmentation of the target area.
[0056] The segmentation head extends the segmentation branch based on the detection head, specifically including: Initial convolutional layer (cv1): Maps the input feature map (usually the lowest-scale feature map, such as the second feature map at the P5 level) to an intermediate number of channels npr (default value is 32) through convolutional operations.
[0057] Upsampling layer (upsample): Use transposed convolution (nn.ConvTranspose2d) to double the resolution of the feature map and halve the number of channels to npr / 2.
[0058] Subsequent convolutional layers (cv2 and cv3): Through two consecutive convolutional operations, gradually adjust the number of channels to nc + 1 (the number of target classes plus the background class) to generate the final segmentation mask.
[0059] The forward propagation process of the segmentation head is as follows: Training stage: Extract the lowest-scale feature map from the input feature map list x, and process it sequentially through cv1, upsample, cv2, and cv3 to output a segmentation mask with a shape of (B, nc + 1, H', W'), where H' and W' are the upsampled resolutions, and directly return it for loss calculation.
[0060] Inference stage: Based on the output in the training stage, further apply the sigmoid function to convert the logits into probability values to generate the final segmentation result. Logits are the raw output values of the last layer of the segmentation head, which are scores without being processed by any activation function.
[0061] The output mask of the segmentation head represents the class (including the background) to which each pixel belongs, and the resolution is restored through the upsampling operation to ensure the spatial alignment of the segmentation result with the input image. This design provides pixel-level accuracy while maintaining computational efficiency.
[0062] In the segmentation task head, to refine the segmentation boundary and ensure the accuracy of spatial information, the segmentation head uses progressive upsampling and feature refinement. Through layer-by-layer upsampling and concatenation operations, the segmentation head restores the high-resolution features, that is, 1 / 2 of the original image resolution, to enhance the capture of the drivable area boundary, and further combines the refined transposed convolution to restore to the original image size (but considering that the transposed convolution is time-consuming on low-computing-power platforms, the transposed convolution is replaced by upsampling), enabling the model to better distinguish the boundaries in complex scenes and ensuring the clarity and accuracy of the segmentation result.
[0063] As Figure 3 shown, the embodiment of the present application provides a multi-task processing method, including: Step 301: Obtain a fisheye image.
[0064] The obtained fisheye image can be adjusted by scaling to make its size consistent with the fisheye image sample to improve the prediction accuracy.
[0065] Step 302: Input the fisheye image into the trained multi-task processing model to obtain the object detection result and region detection result of the fisheye image.
[0066] Among them, the multi-task processing model is trained based on the method of any of the above embodiments.
[0067] Through the pre-trained multi-task processing model in the embodiments of the present application, the object detection task and the drivable area segmentation task can be processed simultaneously. Compared with using neural networks to process tasks separately, the embodiments of the present application can reduce computational consumption and are applicable to low-computing-power chips.
[0068] In an embodiment of the present application, after obtaining the object detection result, post-processing can be performed on it. For example, screening the object detection result to remove bounding boxes with low confidence, and then performing non-maximum suppression on the remaining bounding boxes to reduce overlapping boxes of the same category. In addition, by appropriately scaling and transforming the coordinates of the bounding boxes, it is ensured that the detection result has the correct proportion and position on the original image. Finally, draw the bounding boxes and class labels on the image to provide a visualization result for each detected object.
[0069] In an embodiment of the present application, after obtaining the region detection result, post-processing can also be performed on it. For example, bilinear interpolation upsampling is performed on the region detection result to restore the original image resolution. Then, the sigmoid function is used to map the region detection result to the range of [0,1] to obtain the probability that each pixel belongs to different categories. On this basis, the value with the highest category probability is selected for classification to generate a binary segmentation mask. Finally, the segmentation mask is converted into a color mask and weighted fusion is performed with the fisheye image to distinguish the segmented area from the background area for easy observation and analysis.
[0070] As Figure 4 shown, the embodiments of the present application provide a training device for a multi-task processing model. The multi-task processing model includes: a backbone network, a neck network, a detection head, and a segmentation head; the neck network includes: an upsampling module and a downsampling module; the device includes: A first acquisition module 401, configured to acquire fisheye image samples and sample labels; A feature extraction module 402, configured to input the fisheye image sample into the backbone network, so that the backbone network extracts features from the fisheye image sample and outputs first feature maps of multiple sizes; and splice the first feature maps of the multiple sizes; The feature fusion module 403 is configured to input the feature map formed by splicing the first feature maps of the multiple sizes and the first feature maps of the multiple sizes into the upsampling module, so that the upsampling module splices the upsampled feature map with the first feature map of the corresponding size and outputs the second feature maps of multiple sizes; input the second feature maps of the multiple sizes into the downsampling module, so that the downsampling module splices the downsampled feature map with the second feature map of the corresponding size and outputs the third feature maps of multiple sizes; The task processing module 404 is configured to input the third feature maps of multiple sizes into the detection head, so that the detection head detects the targets in the fisheye image sample and outputs the target detection result; input the second feature maps of multiple sizes into the segmentation head, so that the segmentation head identifies the drivable area in the fisheye image sample and outputs the area detection result; The training module 405 is configured to adjust the parameters of the multi-task processing model based on the target detection result, the area detection result and the sample label.
[0071] As Figure 5 shown, an embodiment of the present application provides a multi-task processing device, including: The second acquisition module 501 is configured to acquire a fisheye image; The prediction module 502 is configured to input the fisheye image into the trained multi-task processing model to obtain the target detection result and the area detection result of the fisheye image; Wherein, the multi-task processing model is trained based on the method of any of the above embodiments.
[0072] An embodiment of the present application provides a computer program product, and when the computer program / instructions are executed by a processor, the method of any of the above embodiments is implemented.
[0073] Although example embodiments have been described herein with reference to the drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as claimed by the appended claims.
[0074] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0075] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0076] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0077] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the methods of the present application should not be construed as reflecting the intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in that the corresponding technical problems can be solved by features less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiments are hereby expressly incorporated into the specific embodiments, where each claim itself serves as a separate embodiment of the present application.
[0078] Those skilled in the art can understand that, except for features that are mutually exclusive, any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0079] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0080] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some of the modules according to the embodiments of the present application. The present application can also be implemented as a device program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
[0081] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0082] As described above, the above are only specific embodiments of the present application or descriptions of specific embodiments, and the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A training method for a multi-task processing model, characterized in that: The multi-task processing model includes: a backbone network, a neck network, a detection head and a segmentation head; the neck network includes: an upsampling module and a downsampling module; the method includes: Get fisheye image samples and sample labels; Inputting the fisheye image sample into the backbone network so that the backbone network performs feature extraction on the fisheye image sample and outputs first feature maps of multiple sizes; and splicing the first feature maps of multiple sizes; Inputting a feature map formed by splicing the first feature maps of the multiple sizes and the first feature maps of the multiple sizes into the upsampling module, so that the upsampling module splices the feature map obtained by upsampling with the first feature maps of the corresponding sizes, and outputs second feature maps of the multiple sizes; Inputting the second feature maps of the multiple sizes into the downsampling module, so that the downsampling module concatenates the feature maps obtained by downsampling with the second feature maps of the corresponding sizes, and outputs third feature maps of the multiple sizes; Inputting the third feature maps of the plurality of sizes into the detection head, so that the detection head detects the target in the fisheye image sample and outputs the target detection result; Inputting the second feature maps of the plurality of sizes into the segmentation head, so that the segmentation head identifies the drivable area in the fisheye image sample and outputs an area detection result; Based on the target detection result, the region detection result and the sample label, the parameters of the multi-task processing model are adjusted.
2. The method according to claim 1, characterized in that The backbone network includes: a focusing module, multiple convolutional layers, multiple convolutional to feature fusion modules and a pooling module; The pooling module includes: an initial convolution layer, a maximum pooling and concatenation layer, and a final convolution layer; The pooling kernel used in the maximum pooling and concatenation layer is 3x3.
3. The method according to claim 1, characterized in that The upsampling module includes: an upsampling layer, a first concatenation layer and a first C3 module; wherein C3 represents a cross-stage local bottleneck with three convolutional layers; The upsampling module splices the feature map obtained by upsampling with the first feature map of the corresponding size, and outputs second feature maps of multiple sizes, including: The upsampling layer upsamples the feature map formed by splicing the first feature maps of the multiple sizes, the first splicing layer splices the feature map obtained by the upsampling with the first feature map of the fourth level, and the first C3 module processes the spliced feature map to obtain the second feature map of the fourth level; The upsampling layer upsamples the second feature map of the fourth level, the first concatenation layer concatenates the feature map obtained by the upsampling with the first feature map of the third level, and the first C3 module processes the concatenated feature map to obtain the second feature map of the third level; The upsampling layer upsamples the second feature map of the third level, the first concatenation layer concatenates the feature map obtained by the upsampling with the first feature map of the second level, and the first C3 module processes the concatenated feature map to obtain the second feature map of the second level; The upsampling layer upsamples the second feature map of the second level, the first concatenation layer concatenates the feature map obtained by the upsampling with the first feature map of the first level, and the first C3 module processes the concatenated feature map to obtain the second feature map of the first level; Among them, the sizes of the first feature map and the second feature map corresponding to the first level, the second level, the third level and the fourth level are reduced in sequence.
4. The method according to claim 3, characterized in that The downsampling module further includes: a downsampling layer, a second splicing layer, and a second C3 module; The downsampling module concatenates the downsampled feature map with the second feature map of the corresponding size, and outputs a third feature map of multiple sizes, including: The downsampling layer downsamples the second feature map of the first level, the second concatenation layer concatenates the feature map obtained by the downsampling with the second feature map of the second level, and the second C3 module processes the concatenated feature map to obtain a third feature map of the second level; The downsampling layer downsamples the third feature map of the second level, the second splicing layer splices the feature map obtained by the downsampling with the second feature map of the third level, and the second C3 module processes the spliced feature map to obtain a third feature map of the third level; The downsampling layer downsamples the third feature map of the third level, the second splicing layer splices the feature map obtained by downsampling with the second feature map of the fourth level, and the second C3 module processes the spliced feature map to obtain the third feature map of the fourth level; The downsampling layer downsamples the third feature map of the fourth level, the second splicing layer splices the feature map obtained by downsampling with the second feature map of the fifth level, and the second C3 module processes the spliced feature map to obtain the third feature map of the fifth level; Among them, the sizes of the third feature maps corresponding to the second level, the third level, the fourth level and the fifth level decrease in sequence.
5. The method according to claim 1, characterized in that The detection head comprises: a bounding box regression module and a classification module; the bounding box regression module and the classification module each comprise a plurality of convolutional layers; The detection head detects the target in the fisheye image sample and outputs the target detection result, including: The bounding box regression module outputs a bounding box detection result of the target, and the classification module outputs a category detection result of the target.
6. The method according to claim 1, characterized in that The segmentation head includes: an initial convolution layer, an upsampling layer and a subsequent convolution layer; The segmentation head identifies the drivable area in the fisheye image sample and outputs the area detection result, including: The initial convolution layer maps the second feature maps of multiple sizes to a preset number of channels through a convolution operation, the upsampling layer increases the resolution of the second feature map and reduces the number of channels of the second feature map based on transposed convolution, and the subsequent convolution layer reduces the number of channels of the second feature map through two convolution operations to obtain the area detection result.
7. A multitasking method, characterized in that: include: Get fisheye image; Inputting the fisheye image into a trained multi-task processing model to obtain an object detection result and an area detection result of the fisheye image; Wherein, the multi-task processing model is trained based on any one of the methods described in claims 1-6.
8. A training device for a multi-task processing model, characterized in that: The multi-task processing model includes: a backbone network, a neck network, a detection head and a segmentation head; the neck network includes: an upsampling module and a downsampling module; the device includes: A first acquisition module, configured to acquire fisheye image samples and sample labels; A feature extraction module is configured to input the fisheye image sample into the backbone network so that the backbone network performs feature extraction on the fisheye image sample and outputs first feature maps of multiple sizes; and splices the first feature maps of multiple sizes; The feature fusion module is configured to input the feature map formed by splicing the first feature maps of the multiple sizes and the first feature maps of the multiple sizes into the upsampling module, so that the upsampling module splices the feature map obtained by upsampling with the first feature map of the corresponding size, and outputs the second feature map of the multiple sizes; input the second feature map of the multiple sizes into the downsampling module, so that the downsampling module splices the feature map obtained by downsampling with the second feature map of the corresponding size, and outputs the third feature map of the multiple sizes; The task processing module is configured to input the third feature maps of the multiple sizes into the detection head so that the detection head detects the target in the fisheye image sample and outputs the target detection result; input the second feature maps of the multiple sizes into the segmentation head so that the segmentation head identifies the drivable area in the fisheye image sample and outputs the area detection result; A training module is configured to adjust parameters of the multi-task processing model based on the target detection result, the region detection result and the sample label.
9. A multi-tasking processing device, characterized in that: include: A second acquisition module is configured to acquire a fisheye image; A prediction module, configured to input the fisheye image into a trained multi-task processing model to obtain a target detection result and a region detection result of the fisheye image; Wherein, the multi-task processing model is trained based on any one of the methods described in claims 1-6.
10. A computer program product, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-task joint sensing network model for traffic road pavement information and detection method
CN115797881A
Image area detection model construction method, system and equipment
CN117636114A
Road environment perception algorithm in dark scene
CN117636286A
Method and device for determining detection frame of fisheye image, equipment and medium
CN119762743A
Multi-task joint perception network model and detection method for traffic road surface information
US20240420487A1