Multi-task network road target detection method based on subtitle perception pre-training
By adopting the subtitle perception pre-training method in the multi-task perception model, the problems of insufficient initial representation of feature extraction and lack of guidance of semantic information are solved, and the detection accuracy of the model in autonomous driving scenarios is significantly improved.
Patent Information
- Application Number
- CN202510261927.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
In the autonomous driving scenario, the existing multitasking perception model has problems such as insufficient initial representation of feature extraction, lack of guidance of semantic information, and insufficient detection capabilities for small targets and edge areas.
Using a multi-task network method based on subtitle perception pre-training, the image encoder and subtitle perception decoder are constructed, and the semantic description of the image is automatically generated is carried out, so as to effectively learn the deep semantic information of the image in the pre-training stage.
It significantly improves the feature extraction capability of the image encoder, enhances the model's perception of semantic information in autonomous driving scenarios, and improves the detection accuracy of small targets and edge areas.
Smart Images

Figure CN120198871A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle autonomous driving, and particularly relates to a multi-task network road target detection method based on subtitle perception pre-training. Background Art
[0002] With the rapid development of deep learning technology, computer vision has made remarkable progress in the field of autonomous driving. One of the core tasks of an autonomous driving system is to accurately and efficiently perceive its surrounding environment. In this process, the visual perception system, as the core of the perception system of an autonomous driving vehicle, collects image or video data of the surrounding environment through visual sensors such as cameras, and processes and analyzes it with the help of computer vision technology, so as to realize the recognition, classification, positioning and understanding of target objects in the environment.
[0003] Current perception models usually rely on single-task processing, such as vehicle detection, pedestrian detection, lane line segmentation or drivable area segmentation, etc. These tasks often need to be processed by independent models respectively, which not only leads to redundant consumption of computing resources, increases the inference time, but also may affect the consistency of the overall perception result due to the lack of feature sharing between tasks. In addition, there are various target objects with inconsistent scales in the actual driving scenario, such as distant vehicles, small obstacles, etc. Single-task models are difficult to take into account the global environment understanding and the fine detection of local features, and there are certain deficiencies in the detection of small targets and edge areas, which affects the reliability of the perception system.
[0004] In recent years, multi-task learning (MTL) has been widely applied to the field of autonomous driving in order to improve the overall performance of the perception system. Multi-task learning completes multiple tasks simultaneously through a shared network feature extraction layer and a joint optimization strategy, thereby achieving feature sharing, computational resource savings, and detection efficiency improvement. For example, Chinese Patent CN116665176B discloses a multi-task network road target detection method for vehicle autonomous driving. The method includes the following steps: collecting a detection data set of a vehicle driving road, dividing it into a training set, a validation set, and a test set, and performing data augmentation on the input image; annotating the data in the data set according to different types in the detection scenario; building a multi-task network model and constructing a loss function; selecting different training methods to train the multi-task network model according to different requirements, and obtaining the best converged model after multiple iterations. However, this patent has the following disadvantages: Traditional multi-task networks mainly rely on convolutional neural networks to encode image features, do not fully consider the initial representation ability of the feature extraction network, directly extract features from the original image, resulting in insufficient initial feature representation of the network and affecting the convergence speed of subsequent multi-task network training; only achieving feature fusion through feature splicing or convolution operations, without introducing semantic information to guide the feature extraction process, resulting in a semantic mismatch problem in the feature fusion method and affecting the detection accuracy of the boundary region; relying on a large amount of manually labeled data for network training, without enhancing the network feature learning ability through unsupervised or weakly supervised methods, resulting in a significant performance decline in scenarios with insufficient labeled data; in addition, although existing multi-task perception models have been proposed, these models have poor detail capture ability in driving scenarios and the perception accuracy is worse than that of single-task models. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a multi-task network road target detection method based on caption-aware pre-training.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] On the one hand, the present invention provides a multi-task network road target detection method based on caption-aware pre-training, including the following steps:
[0008] Construct an image encoder and a caption-aware decoder;
[0009] Perform the first training on the image encoder and the caption-aware decoder to obtain a pre-trained image encoder;
[0010] Construct a multi-task decoder head, and the multi-task decoder head includes an object detection head, a lane line detection head, and a drivable area detection head;
[0011] Perform a second training on the pre-trained image encoder and the multi-task decoder head to obtain the final image encoder and the multi-task decoder;
[0012] Obtain the autonomous driving image data to be detected, and input the autonomous driving image data to be detected into the image encoder and the multi-task decoder after the second training, and output the bounding box for object detection, the position of the lane line, and the segmentation map of the drivable area.
[0013] Further, the first training on the image encoder and the caption-aware decoder specifically includes:
[0014] Obtain a first training dataset, where the first training dataset includes autonomous driving scene images and corresponding text annotation data, and the text annotation data is the natural language description and explanation corresponding to the autonomous driving scene images, including the driving behavior description of the lanes in the images, traffic sign descriptions, and corresponding driving behavior explanations;
[0015] Preprocess and normalize the autonomous driving scene images, and uniformly set the image size to 640x640x3;
[0016] Input the preprocessed image data into the image encoder to obtain the first image feature;
[0017] Preprocess the text annotation data to convert the text data into a text token sequence; where a token is a discrete representation of the text data, and each token corresponds to a basic unit in the text data, and the basic units include characters, words, and sub-words;
[0018] Input the first image feature and the text token sequence into the caption-aware decoder, and continuously output text tokens in an autoregressive manner to obtain the complete first text output;
[0019] Compare the first text output with the text annotation data, and optimize the parameters of the image encoder and the caption-aware decoder through the first loss function and the backpropagation algorithm to obtain the pre-trained image encoder.
[0020] Further, the first loss function is:
[0021]
[0022] where, L FL is the first loss function, T is the number of text tokens in the first text output, p t is the predicted probability at the t-th position. When the text token at the t-th position output by the caption-aware decoder is the same as the text token at the t-th position in the text annotation data, then p t= p, when the text token at the t-th position output by the subtitle perception decoder is different from the text token at the t-th position in the text annotation data, then p t = 1 - p, where p is the prediction probability of the subtitle perception decoder for the text token at the t-th position, and α t is the balance factor, and FL(p t ) is the loss value of the text token at the t-th position output by the subtitle perception decoder, and γ is the focal factor.
[0023] Furthermore, the image encoder includes a backbone network and a neck network;
[0024] The backbone network includes:
[0025] The first layer includes an independent convolutional module;
[0026] The second layer includes multiple convolutional modules and a CSP module, and the second layer continuously performs operations of convolutional modules and CSP modules three times;
[0027] The third layer includes a convolutional module and a CIB module;
[0028] The fourth layer includes an SPPF module;
[0029] The fifth layer includes an SSA module;
[0030] The neck network includes a first downsampling module, a CSP module, and a second downsampling module connected in sequence.
[0031] Furthermore, inputting the preprocessed image data into the image encoder to obtain the first image feature specifically includes:
[0032] Input the preprocessed image data into the first-layer convolutional module of the backbone network of the image encoder, perform two-dimensional convolution, batch normalization, and activation function operations to obtain the feature p1, and the dimension of the feature p1 is 320×320×64×W;
[0033] Input the feature p1 into the second layer of the backbone network of the image encoder, and perform operations of convolutional modules and CSP modules three times continuously to obtain the feature p2, and the dimension of the feature p2 is 80×80×256×W;
[0034] Input the feature p2 into the third layer of the backbone network of the image encoder, and perform operations of convolutional modules and CIB modules in sequence to obtain the feature p3, and the dimension of the feature p3 is 20×20×512×W;
[0035] Input the feature p3 into the fourth - layer SPPF module of the backbone network of the image encoder, perform max - pooling operation, concatenation operation, and convolution - module operation in sequence to obtain the feature p4, and the dimension of the feature p4 is 20×20×512×W×R;
[0036] Input the feature p4 into the fifth - layer SSA module of the backbone network of the image encoder, perform the masked sparse self - attention module MSSA operation and convolution - module operation in sequence to obtain the feature p5, and the dimension of the feature p5 is 20×20×512×W×R;
[0037] Input the feature p5 output by the fifth - layer SSA module of the backbone network into the first down - sampling module of the neck network for down - sampling operation, output the feature f3, concatenate the feature f3 with the feature p2 output by the second layer of the backbone network, input the concatenated feature into the CSP module of the neck network for feature extraction to obtain the feature f2, and input the feature f2 into the second down - sampling module for down - sampling operation to output the first image feature f1.
[0038] Further, input the first image feature and text tokens into the caption - aware decoder, and continuously output text tokens in an autoregressive manner to obtain the complete first text output, specifically including:
[0039] Perform positional encoding on the first image feature, perform positional encoding on the text token sequence, concatenate the position - encoded first image feature and the text token sequence to obtain a visual - text feature vector, and perform positional encoding on the visual - text feature vector;
[0040] Input the position - encoded visual - text feature vector into the caption - aware decoder, and continuously output text tokens in an autoregressive manner to obtain the complete first text output.
[0041] Further, the caption - aware decoder is an improved Transformer - based decoder, and the improvement includes replacing the cross - attention module of the Transformer - based decoder with a sparse masked self - attention module MSAM.
[0042] Further, perform the second training on the pre - trained image encoder and multi - task decoder head to obtain the final image encoder and multi - task decoder, specifically including:
[0043] Obtain a second training dataset, and the second training dataset includes autonomous driving scene images and their corresponding object detection annotation data, lane line detection annotation data, and drivable area segmentation annotation data;
[0044] Preprocess and normalize the autonomous driving images in the second training dataset, and uniformly adjust the image size to 640x640x3;
[0045] Input the preprocessed autonomous driving images into the first trained image encoder to obtain the feature f3 output by the first downsampling module of the neck network of the image encoder, the feature f2 output by the CSP module of the neck network, and the feature f1 output by the second downsampling module of the neck network. Among them, the dimension of feature f1 is 80x80x256xW, the dimension of feature f2 is 40x40x512xW, and the dimension of feature f3 is 20x20x512xWxR;
[0046] Input the feature f1 into the lane line detection head and the drivable area detection head, and respectively output the position of the lane line and the segmentation map of the drivable area;
[0047] Input the feature f2 and the feature f3 into the object detection head, and output the bounding box of object detection;
[0048] According to the output bounding box of object detection, the position of the lane line, the segmentation map of the drivable area, and the object detection annotation data, lane line detection annotation data, and drivable area segmentation annotation data in the second training dataset, perform joint training through the second loss function, and update the parameters of the image encoder and the multi-task decoder head through the backpropagation algorithm to obtain the final image encoder and multi-task decoder.
[0049] Furthermore, the object detection annotation data includes the position and category label of the object in the autonomous driving scene image, the lane line detection annotation data includes the specific position and marking information of the lane line, and the drivable area segmentation annotation data represents the drivable area in the image.
[0050] Furthermore, the second loss function is:
[0051] L all =σ1l det +σ2l ll +σ3l da
[0052] where, L all is the second loss function, σ1, σ2, σ3 are balance factors, and l det is the object detection loss function, l ll is the lane line detection loss function, and l da is the drivable area segmentation loss function;
[0053] The object detection loss function is:
[0054] l det =α1lclass +α2l obj +α3l box
[0055] where α1, α2, and α3 are balance factors, and l class is the classification loss function, and l obj is the target confidence loss function, and l box is the bounding box regression loss function;
[0056] The lane line detection loss function, the drivable area segmentation loss function, the classification loss function, and the target confidence loss function all adopt the cross - entropy loss function l ce ;
[0057] The cross - entropy loss function l ce is:
[0058]
[0059] where C is the number of classes, y i is the true label of the sample, and p i is the predicted probability of the model for class i;
[0060] The bounding box regression loss function is:
[0061]
[0062] where b=(xb, yb, wb, hb) are the coordinates of the true bounding box, is the coordinate of the predicted bounding box, is the intersection - over - union of the true box and the predicted box, is the Euclidean distance between the centers of the two boxes, c is the diagonal length of the smallest closed box enclosing the true box and the predicted box, v is the aspect ratio consistency loss, representing the difference in aspect ratio, and α is a tuning parameter used to control the influence of the aspect ratio loss.
[0063] Compared with the prior art, the present invention has the following advantages:
[0064] (1) Before the multi-task network training, the present invention constructs a caption-aware pre-training network, including an image encoder and a caption-aware head. By generating semantic descriptions of images in an autoregressive manner, deep semantic information of images is effectively learned during the pre-training phase, enabling the image encoder to extract richer and semantically-guided feature representations. The image encoder can master the global semantic information and local detail features of images at the initial stage, effectively avoiding the problems of slow convergence and insufficient feature representation caused by directly training from randomly initialized weights in traditional methods. During the pre-training phase, the image caption generation task guides the image feature extraction process, prompting the encoder to focus on key regions, scene contexts, and semantic relationships in the images, thereby significantly enhancing the feature representation ability of subsequent multi-task detection.
[0065] (2) The present invention builds a unified multi-task perception network model, adopts a shared feature extraction encoder, and jointly optimizes and trains tasks such as object detection, lane line segmentation, and drivable area segmentation. Through the feature sharing method, the problems of redundant computing resources and repeated feature extraction caused by separate modeling of multiple tasks are avoided, and the computational complexity of the perception system is significantly reduced. The joint training method enables different tasks to share underlying visual features, enhances the model's understanding of the semantic associations between tasks, and realizes the consistent expression of global perception information, especially showing superiority in detecting edge regions in complex driving scenarios.
[0066] (3) The present invention innovatively introduces a prefix prompting strategy in the caption-aware decoder and replaces the traditional cross-attention mechanism with a sparse masked self-attention module MSAM, effectively solving the problem of insufficient semantic alignment in the prior art. The prefix prompting strategy can directly inject image feature information into the text generation process at the initial stage of decoding, enhancing the guiding ability of visual features for the generation process and achieving more accurate visual semantic descriptions. The sparse masked self-attention module only selects the key region features in the image for attention calculation, can suppress background noise, and significantly improves the model's perception ability for small targets, occluded targets, and boundary regions.
[0067] (4) The present invention designs a feature alignment strategy. By cascading upsampling and feature fusion methods, the low-level features and high-level features are spatially scaled and aligned, further enhancing the feature expression ability of the edge regions. This strategy effectively solves the problems of blurred boundaries and missed detections caused by inconsistent feature scales in existing multi-task perception models, especially showing significant performance in the lane line and drivable area segmentation tasks.
[0068] (5) The present invention performs unsupervised pre-training on the image encoder through the image caption generation task, and makes full use of a large amount of unlabeled data for visual feature learning. This technical means can significantly improve the initial representation ability of the feature extraction network in scenarios where data annotation is scarce, reduce the dependence on a large amount of manually annotated data, and has significant advantages in small-sample learning and rare-class sample detection tasks. Brief Description of the Drawings
[0069] Figure 1 is a flowchart of the method of the present invention;
[0070] Figure 2 is a training block diagram of the multi-task autonomous driving visual perception system of the present invention;
[0071] Figure 3 is a model diagram of the image encoder of the present invention;
[0072] Figure 4 is a model diagram of the convolutional module, CSP module, and SPPF module of the image encoder of the present invention;
[0073] Figure 5 is a model diagram of the CIB module and SSA module of the image encoder of the present invention;
[0074] Figure 6 is the caption-aware decoder of the present invention. Detailed Description of the Embodiments
[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0076] Embodiment 1:
[0077] This embodiment provides a multi-task network road target detection method based on caption-aware pre-training, as Figure 1 shown, including the following steps:
[0078] Construct an image encoder and a caption-aware decoder;
[0079] Perform the first training on the image encoder and the caption-aware decoder to obtain a pre-trained image encoder;
[0080] Construct a multi-task decoder head, and the multi-task decoder head includes an object detection head, a lane line detection head, and a drivable area detection head;
[0081] Perform a second training on the pre-trained image encoder and the multi-task decoder head to obtain the final image encoder and multi-task decoder;
[0082] Obtain the autonomous driving image data to be detected, and input the autonomous driving image data to be detected into the image encoder and multi-task decoder after the second training, and output the bounding box for object detection, the position of the lane line, and the segmentation map of the drivable area.
[0083] Further, as Figure 2 shown, the first training on the image encoder and the caption-aware decoder specifically includes:
[0084] Obtain the first training dataset, where the first training dataset includes autonomous driving scene images and corresponding text annotation data, and the text annotation data is the natural language description and explanation corresponding to the autonomous driving scene images, including the driving behavior description of the lanes in the images, traffic sign descriptions, and corresponding driving behavior explanations;
[0085] Preprocess and normalize the autonomous driving scene images, and uniformly set the image size to 640x640x3;
[0086] Input the preprocessed image data into the image encoder to obtain the first image features;
[0087] Preprocess the text annotation data to convert the text data into a text token sequence; where a token is a discrete representation of the text data, and each token corresponds to a basic unit in the text data, and the basic units include characters, words, and sub-words;
[0088] Input the first image features and the text token sequence into the caption-aware decoder, and continuously output text tokens in an autoregressive manner to obtain the complete first text output;
[0089] Compare the first text output with the text annotation data, and optimize the parameters of the image encoder and the caption-aware decoder through the first loss function and the backpropagation algorithm to obtain the pre-trained image encoder.
[0090] The benefits of building the pre-training network are as follows:
[0091] 1. Improve the feature extraction ability
[0092] The caption-aware task generates descriptions of images through autoregression, which enables the model to learn the deep semantic information of the images during the pre-training stage. This allows the image encoder to extract richer and more meaningful features, rather than just low-level visual information (such as edges, textures, etc.). These learned high-level semantic features perform better in subsequent downstream tasks such as object detection and lane line detection.
[0093] 2. Improve the generalization ability of the model
[0094] Through the pre-training of the caption-aware task, the image encoder can process different types of scene descriptions and learn more context information, which helps to improve the generalization ability of the model in different scenarios. The pre-training helps the model understand the details and background in the image, thus enhancing the recognition ability for complex environments and enabling it to better adapt to the changing scenarios in the real world.
[0095] 3. Reduce the dependence on labeled data
[0096] Many downstream tasks, especially object detection and semantic segmentation, often require a large amount of labeled data for training. Through pre-training with the caption-aware task, the model can obtain useful visual features through unsupervised learning methods, thereby reducing the need for a large amount of labeled data in subsequent downstream tasks. This is especially beneficial for application scenarios where the cost of data annotation is high.
[0097] 4. Accelerate the convergence speed
[0098] Through pre-training, the image encoder has learned some basic image feature representations and semantic information in the image. Therefore, when training downstream tasks, it can start from a better initial state, avoiding the difficulties of training from scratch. This usually accelerates the training process of downstream tasks and reduces the training time.
[0099] 5. Enhance the robustness to difficult samples
[0100] The caption-aware task not only learns the visual information of the image but also includes the context and semantic meaning of the image. This enables the pre-trained image encoder to perform more robustly when facing difficult samples (such as blur, occlusion, etc.) and can improve the detection accuracy through a more comprehensive understanding of the scene.
[0101] 6. Improve the multi-task learning ability of the model
[0102] The caption-aware task itself is a multi-task learning task that simultaneously considers multiple aspects of the image (such as objects, background, actions, etc.). This multi-task learning method can promote the image encoder to process multiple tasks simultaneously in downstream tasks, thereby improving the performance of the model in multi-task learning.
[0103] The step of performing the first training on the image encoder and the caption-aware decoder mainly focuses on the feature extraction and semantic understanding problems in visual perception tasks in the autonomous driving scenario. Through the visual-text multimodal pre-training method, the initial feature extraction ability of the image encoder is improved, and the model's perception ability of the semantic information in the autonomous driving scenario is enhanced. This pre-training process combines images and natural language descriptions for joint learning, which can effectively make up for the problem of insufficient initial feature representation ability of the image encoder in the existing technology and reduce the dependence on a large amount of labeled data at the same time.
[0104] Further, the first loss function is as follows:
[0105]
[0106] where, L FL is the first loss function, T is the number of text tokens in the first text output, p t is the prediction probability at the t-th position. When the text token at the t-th position output by the caption-aware decoder is the same as the text token at the t-th position in the text annotation data, then p t = p. When the text token at the t-th position output by the caption-aware decoder is different from the text token at the t-th position in the text annotation data, then p t = 1 - p. p is the prediction probability of the caption-aware decoder for the text token at the t-th position output, α t is the balance factor, FL(p t ) is the loss value of the text token at the t-th position output by the caption-aware decoder, and γ is the focus factor.
[0107] The first loss function adopts the focal loss strategy. By introducing the focus factor and the balance factor, it can effectively solve the problem of class imbalance and strengthen the model's attention to difficult-to-predict text tokens. The focal loss helps the model improve the recognition ability of low-frequency and complex texts by increasing the penalty for misclassified samples, avoiding overfitting to common text patterns. At the same time, the application of the balance factor ensures the reasonable weight allocation of text tokens at different positions, thereby improving the model's generalization ability and text generation quality. This loss function design can effectively accelerate the training process and improve the accuracy and robustness of the model for complex and difficult-to-predict text descriptions in the autonomous driving scenario.
[0108] Further, as Figure 3 shown, the image encoder includes a backbone network and a neck network;
[0109] The backbone network includes:
[0110] The first layer includes an independent convolutional module;
[0111] The second layer includes multiple convolutional modules and a CSP module, and the operations of the convolutional module and the CSP module are continuously performed three times in the second layer;
[0112] The third layer includes a convolutional module and a CIB module;
[0113] The fourth layer includes an SPPF module;
[0114] The fifth layer includes an SSA module;
[0115] The neck network includes a first downsampling module, a CSP module, and a second downsampling module connected in sequence.
[0116] As Figure 4 shown, the convolutional module first performs a 2D convolution operation on the input image or feature, then performs batch normalization, and finally passes through the SILU activation function to increase non-linear features;
[0117] The CSP module first passes through a convolutional module, then performs a split operation to split the feature dimension, then passes through two layers of darknet networks for processing, performs a splicing operation, and finally performs a convolutional module processing. The main benefits of the CSP module:
[0118] 1. Improve computational efficiency
[0119] The CSP module divides the feature map into different parts and processes them at different stages of the network. This phased processing method can effectively reduce the computational complexity. By splitting and processing the features in the early stage in parallel, the CSP module enhances the feature expression ability without significantly increasing the computational burden.
[0120] 2. Enhance the gradient flow
[0121] The CSP module enables information to flow between different stages of the network through cross-stage partial connections. This design can effectively alleviate the problem of gradient disappearance in deep networks, ensuring that the model is more stable during training. Especially for deeper network layers, the improvement of the gradient flow helps to accelerate convergence.
[0122] 3. Improve the feature expression ability
[0123] The CSP module enables the model to extract different levels of features in different sub-networks through the splitting and fusion of the feature map, and then enhances the expression ability of these features through fusion. This can improve the network's ability to learn image details while retaining more extensive context information.
[0124] 4. Reduce the risk of overfitting
[0125] Due to the feature splitting and parallel processing of the CSP module, it can reduce the risk of overfitting to a certain extent. Especially on limited datasets, the model can learn image features through different paths, better generalize to new data, and improve the robustness of the model.
[0126] 5. Accelerated Training
[0127] The design of the CSP module helps to accelerate the training process. Through partial connection and phased processing, the training efficiency of the network is higher. Especially when dealing with large-scale data, the amount of calculation can be significantly reduced, and the training speed can be increased.
[0128] 6. Improved Memory Efficiency
[0129] Since the calculation method of the CSP module splits the feature map and processes it separately in each stage, the number of feature maps that need to be stored in each stage is reduced. Therefore, the CSP module helps to improve the efficiency of memory usage. Especially when using larger-sized images, it can effectively save memory resources.
[0130] The described SPPF module first passes through a convolutional module, then performs three max-pooling operations, and then passes through a concatenation operation and a convolutional module.
[0131] As Figure 5 shown, the described CIB module consists of a 3x3 depth convolution and a 1x1 convolution. The design of the CIB module is mainly used to improve the feature transfer efficiency of the network, enabling better interaction and fusion of information between different stages (layers), thereby improving the detection performance, especially in complex backgrounds and diverse scenarios. The shown SSA module adds the MSSA (Masked Sparse Self-Attention module) module on the basis of multiple convolutional modules. Masked sparse self-attention can significantly improve the processing ability of long sequences and large-scale data by reducing computational and memory overhead, improve the training and inference speed, and perform well in multi-task, long sequence, and complex scenarios. By focusing on key regions and reducing redundant calculations, masked sparse self-attention can also improve the model's performance and generalization ability, and is an efficient optimization technique when dealing with large-scale inputs.
[0132] Further, inputting the preprocessed image data into the image encoder to obtain the first image feature specifically includes:
[0133] Inputting the preprocessed image data into the first convolutional module of the backbone network of the image encoder, performing two-dimensional convolution, batch normalization, and activation function operations to obtain feature p1, and the dimension of the feature p1 is 320×320×64×W;
[0134] Input the feature p1 into the second layer of the backbone network of the image encoder, perform operations of three consecutive convolutional modules and CSP modules to obtain the feature p2, and the dimension of the feature p2 is 80×80×256×W;
[0135] Input the feature p2 into the third layer of the backbone network of the image encoder, and successively perform operations of a convolutional module and a CIB module to obtain the feature p3, and the dimension of the feature p3 is 20×20×512×W;
[0136] Input the feature p3 into the SPPF module of the fourth layer of the backbone network of the image encoder, and successively perform max-pooling operation, concatenation operation and convolutional module operation to obtain the feature p4, and the dimension of the feature p4 is 20×20×512×W×R;
[0137] Input the feature p4 into the SSA module of the fifth layer of the backbone network of the image encoder, and successively perform operations of the masked sparse attention module MSSA and the convolutional module to obtain the feature p5, and the dimension of the feature p5 is 20×20×512×W×R;
[0138] Input the feature p5 output by the SSA module of the fifth layer of the backbone network into the first downsampling module of the neck network for downsampling operation, output the feature f3, concatenate the feature f3 with the feature p2 output by the second layer of the backbone network, input the concatenated feature into the CSP module of the neck network for feature extraction to obtain the feature f2, and input the feature f2 into the second downsampling module for downsampling operation to output the first image feature f1.
[0139] The advantage of this design is that through multiple levels of convolution and module combinations, features from low-level to high-level are gradually extracted, and the feature representation is optimized through downsampling operations. Each layer can focus on image information at different scales, while ensuring the information transmission efficiency and reducing the computational complexity. Such a structure enables the image encoder to handle more complex visual tasks, improving the performance of the model in high-demand scenarios such as autonomous driving. Especially when dealing with targets and complex backgrounds at different scales, it can maintain high detection accuracy and robustness.
[0140] Further, inputting the first image feature and the text token into the caption-aware decoder to continuously output text tokens in an autoregressive manner to obtain the complete first text output specifically includes:
[0141] Perform positional encoding on the first image feature, perform positional encoding on the text token sequence, concatenate the positionally encoded first image feature and the text token sequence to obtain a visual-text feature vector, and perform positional encoding on the visual-text feature vector;
[0142] The visually-textual feature vectors after positional encoding are input into the caption-aware decoder to continuously output text tokens in an autoregressive manner, obtaining the complete first text output.
[0143] As Figure 6 shown, the caption-aware decoder innovatively incorporates a prefix prompting strategy, that is, the first image features after positional encoding are vector concatenated with the text token sequence, and the visually-textual feature vectors after positional encoding are input into the caption-aware decoder, replacing the previous cross-attention strategy. As Figure 6 shown, the benefits of the prefix prompting strategy are as follows:
[0144] 1. Reduce computational overhead
[0145] Traditional cross-attention mechanisms usually need to process and calculate weights for the entire input sequence. Especially in generation tasks, the model needs to frequently exchange information between the encoder and the decoder. The prefix prompting strategy reduces the amount of computation required for each generation by only processing the prefix part of the input sequence and focusing attention on a small amount of additional information. This makes the training and inference processes more efficient. Especially when dealing with long texts or long sequences, it can significantly reduce the computational burden.
[0146] 2. Enhance the controllability of the model
[0147] The prefix prompting strategy provides greater controllability by directly embedding task-related information into the input part of the model. The prefix part can contain specific task instructions or background information to help the model focus more on specific tasks or content during the generation process. For example, in text generation tasks, prefix prompts can guide the model to generate content in a specific style or domain, making the output more in line with expectations. This strategy enables more flexible control of the model's output without the need for complex cross-attention adjustments.
[0148] 3. Improve the stability of the model
[0149] In traditional cross-attention mechanisms, there are complex interactions between the encoder and the decoder of the model, which may lead to instability during the training process, especially when dealing with long texts. The prefix prompting strategy simplifies the model structure by reducing this interaction, making the training process more stable. At the same time, by fixing the prefix, the model can focus more on the details of specific tasks, thereby reducing the risk of overfitting.
[0150] 4. Reduce memory consumption
[0151] In the traditional cross-attention strategy, every part of the input participates in the attention calculation, which may lead to extremely high memory consumption. The prefix prompting strategy reduces the amount of memory required for each calculation by adding prompting information only to the prefix part of the input. Especially when dealing with long sequences, prefix prompting can effectively reduce memory consumption and improve the efficiency during training and inference.
[0152] 5. Flexible task adaptability
[0153] The prefix prompting strategy can flexibly adjust the content of the prefix according to task requirements. For example, for different generation tasks, the content of the prefix prompt can be dynamically adjusted according to the task requirements, enabling the model to better adapt to different input-output patterns. This flexibility enhances the advantages of the prefix prompting strategy in multi-task learning.
[0154] The caption-aware decoder uses three layers of positional encoding to perform positional encoding on image features, text token vectors, and image-text vectors respectively. The positional encoding provides the position information of the elements in the sequence, enabling the self-attention model to understand and process the sequential relationship in the sequence data, and enhancing the model's expressive ability and generalization ability. It not only helps to improve the model's performance in long sequences, improve long-distance dependence modeling, but also enhances the accuracy in generation tasks.
[0155] Furthermore, the caption-aware decoder is an improved Transformer-based decoder, and the improvement includes replacing the cross-attention module of the Transformer-based decoder with a sparse masked self-attention module MSAM to improve the efficiency and generalization of text output.
[0156] Furthermore, the second training of the pre-trained image encoder and the multi-task decoder head to obtain the final image encoder and multi-task decoder specifically includes:
[0157] Obtain a second training dataset, where the second training dataset includes autonomous driving scene images and their corresponding object detection annotation data, lane line detection annotation data, and segmentation annotation data of the drivable area;
[0158] Preprocess and normalize the autonomous driving images in the second training dataset, and uniformly adjust the image size to 640x640x3;
[0159] Input the preprocessed autonomous driving image into the first trained image encoder to obtain the feature f3 output by the first downsampling module of the neck network of the image encoder, the feature f2 output by the CSP module of the neck network, and the feature f1 output by the second downsampling module of the neck network, where the dimension of the feature f1 is 80x80x256xW, the dimension of the feature f2 is 40x40x512xW, and the dimension of the feature f3 is 20x20x512xWxR;
[0160] Input the feature f1 into the lane line detection head and the drivable area detection head, and respectively output the position of the lane line and the segmentation map of the drivable area;
[0161] Input the feature f2 and the feature f3 into the object detection head, and output the bounding box of the object detection;
[0162] According to the output bounding box of the object detection, the position of the lane line, the segmentation map of the drivable area, and the object detection annotation data, lane line detection annotation data, and drivable area segmentation annotation data in the second training dataset, perform joint training through the second loss function, and update the parameters of the image encoder and the multi-task decoder head through the backpropagation algorithm to obtain the final image encoder and multi-task decoder.
[0163] Further, the object detection annotation data includes the position of the object in the autonomous driving scene image and its class label, the lane line detection annotation data includes the specific position and marking information of the lane line, and the drivable area segmentation annotation data represents the drivable area in the image.
[0164] Further, the second loss function is:
[0165] L all =σ1l det +σ2l ll +σ3l da
[0166] where, L all is the second loss function, σ1, σ2, σ3 are balance factors, and l det is the object detection loss function, l ll is the lane line detection loss function, and l da is the drivable area segmentation loss function;
[0167] The object detection loss function is:
[0168] l det =α1l class +α2l obj +α3l box
[0169] Among them, α1, α2, and α3 are balance factors, l class is the classification loss function, l ovj is the target confidence loss function, l box is the bounding box regression loss function;
[0170] The lane line detection loss function, the drivable area segmentation loss function, the classification loss function, and the target confidence loss function all adopt the cross-entropy loss function l ce ;
[0171] The cross-entropy loss function l ce is:
[0172]
[0173] Among them, C is the number of categories, y i is the true label of the sample, p i is the predicted probability of the model for category i;
[0174] The bounding box regression loss function is:
[0175]
[0176] Among them, b = (xb, yb, wb, hb) are the coordinates of the true bounding box, are the coordinates of the predicted bounding box, is the intersection over union of the true box and the predicted box, The Euclidean distance between the centers of the two boxes, c is the diagonal length of the smallest closed box enclosing the true box and the predicted box, v is the aspect ratio consistency loss, representing the difference in aspect ratio, and α is a tuning parameter used to control the influence of the aspect ratio loss.
[0177] Example 2:
[0178] The following is an example for actual application, including the following steps:
[0179] Step1, preprocess the autonomous driving images together with the corresponding text annotations and multi-task annotations;
[0180] Step2, divide the images according to a preset ratio, which are randomly divided into 7:1:2 for training, validation, and testing respectively;
[0181] Step3, input the autonomous driving images into the image encoder network to obtain image encoder features;
[0182] Step4, input the image features into the caption-aware decoder to obtain text predictions, calculate the loss function between the text annotations and the text predictions, and perform backpropagation;
[0183] Step 5. The parameters of the model are updated using the Adam optimization strategy, and the pre-trained network is continuously iteratively trained to obtain a pre-trained image encoder.
[0184] Step 6. Continue to train the multi-task decoder head. Input the features of the pre-trained image encoder into the object detection head, lane line detection head, and drivable area segmentation head, calculate the total loss function, and continuously iteratively optimize to obtain multi-task perception results.
[0185] Step 6. Input the test set of the dataset into the trained multi-task visual perception model to obtain perception results.
[0186] After multiple rounds of experiments, we compared the training results of our model with those of other methods. As shown in Table 1, experiments were conducted on the BDD100K dataset.
[0187] Table 1 Autonomous Driving Multi-Task Perception Results
[0188]
[0189] Table 1 shows the evaluation results of our method on three tasks (object detection, lane line detection, and drivable area segmentation). For the object detection task, recall and mean average precision (mAP50) are used as evaluation metrics. The experimental results show that our model reaches 93.1% and 81.3% in these two metrics respectively, surpassing YOLOP, YOLOM, and HybridNet, reaching the SOTA level. In the lane line detection task, accuracy and intersection over union (IoU) are used as evaluation metrics, and our model performs excellently in both metrics. In the drivable area segmentation task, mean intersection over union (mIoU) is used as the evaluation metric, and our model surpasses previous models, showing good segmentation results.
[0190] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0191] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A multi-task network road object detection method based on caption-aware pre-training, characterized in that: The following steps are involved: Build image encoder and caption-aware decoder; Performing a first training on the image encoder and the caption-aware decoder to obtain a pre-trained image encoder; Constructing a multi-task decoder head, wherein the multi-task decoder head includes an object detection head, a lane line detection head, and a drivable area detection head; Performing a second training on the pre-trained image encoder and multi-task decoder heads to obtain the final image encoder and multi-task decoder; Obtain the autonomous driving image data to be detected, input the autonomous driving image data to be detected into the second trained image encoder and multi-task decoder, and output the bounding box of the target detection, the position of the lane line, and the segmentation map of the drivable area.
2. According to claim 1, a multi-task network road object detection method based on subtitle-aware pre-training is characterized in that: The first training of the image encoder and the subtitle-aware decoder specifically includes: Acquire a first training data set, where the first training data set includes an autonomous driving scene image and corresponding text annotation data, where the text annotation data is a natural language description and explanation corresponding to the autonomous driving scene image, including a driving behavior description of a lane in the image, a traffic sign description, and a corresponding driving behavior explanation; Preprocess and normalize the autonomous driving scene images and set the image size to 640x640x3; Inputting the preprocessed image data into an image encoder to obtain a first image feature; Preprocess the text annotation data and convert the text data into a text token sequence; where a token is a discretized representation of the text data, and each token corresponds to a basic unit in the text data, including characters, words, and subwords; Input the first image feature and the text token sequence into a subtitle-aware decoder, and continuously output the text tokens in an autoregressive manner to obtain a complete first text output; The first text output is compared with the text annotation data, and the parameters of the image encoder and the caption-aware decoder are optimized through a first loss function and a back-propagation algorithm to obtain a pre-trained image encoder.
3. According to claim 2, a multi-task network road object detection method based on subtitle-aware pre-training is characterized in that: The first loss function is: Among them, L FL is the first loss function, T is the number of text tokens output by the first text, and p t is the predicted probability of the t-th position. When the text token at the t-th position output by the subtitle-aware decoder is the same as the text token at the t-th position in the text annotation data, then p t =p, when the text token at the tth position output by the subtitle-aware decoder is different from the text token at the tth position in the text annotation data, then p t =1-p, p is the predicted probability of the subtitle-aware decoder outputting the text token at the tth position, α t is the balance factor, FL(p t ) is the loss value of the text token at the tth position output by the subtitle-aware decoder, and γ is the focus factor.
4. According to claim 2, a multi-task network road object detection method based on subtitle-aware pre-training is characterized in that: The image encoder includes a backbone network and a neck network; The backbone network includes: The first layer,consists of an independent convolutional module; The second layer includes a plurality of convolution modules and CSP modules, and the second layer performs the operations of the convolution modules and the CSP modules three times in succession; The third layer includes convolutional modules and CIB modules; The fourth layer includes the SPPF module; The fifth layer includes the SSA module; The neck network includes a first down-sampling module, a CSP module, and a second down-sampling module which are connected in sequence.
5. The multi-task network road object detection method based on subtitle-aware pre-training according to claim 2 or 4, characterized in that: The step of inputting the preprocessed image data into an image encoder to obtain a first image feature specifically includes: The preprocessed image data is input into the first layer convolution module of the backbone network of the image encoder, and two-dimensional convolution, batch normalization and activation function operations are performed to obtain feature p1, where the dimension of feature p1 is 320×320×64×W; Input feature p1 into the second layer of the backbone network of the image encoder, perform three consecutive convolution modules and CSP modules, and obtain feature p2, where the dimension of feature p2 is 80×80×256×W; Input feature p2 into the third layer of the backbone network of the image encoder, and perform operations of the convolution module and the CIB module in sequence to obtain feature p3, where the dimension of feature p3 is 20×20×512×W; Input feature p3 into the fourth-layer SPPF module of the backbone network of the image encoder, perform maximum pooling operation, splicing operation and convolution module operation in sequence, and obtain feature p4, wherein the dimension of feature p4 is 20×20×512×W×R; Input feature p4 into the fifth layer SSA module of the backbone network of the image encoder, perform mask sparse attention module MSSA operation and convolution module operation in sequence, and obtain feature p5, wherein the dimension of feature p5 is 20×20×512×W×R; The feature p5 output by the fifth-layer SSA module of the backbone network is input into the first downsampling module of the neck network for downsampling operation, and the feature f3 is output. The feature f3 is concatenated with the feature p2 output by the second layer of the backbone network, and the concatenated feature is input into the CSP module of the neck network for feature extraction to obtain the feature f2. The feature f2 is input into the second downsampling module for downsampling operation, and the first image feature f1 is output.
6. The multi-task network road object detection method based on subtitle-aware pre-training according to claim 1 is characterized in that: The first image feature and the text token are input into a subtitle-aware decoder, and the text token is continuously output in an autoregressive manner to obtain a complete first text output, specifically including: Position encoding is performed on the first image feature, position encoding is performed on the text token sequence, vector concatenation is performed on the position-encoded first image feature and the text token sequence to obtain a visual-text feature vector, and position encoding is performed on the visual-text feature vector; The position-encoded visual-text feature vector is input into the subtitle-aware decoder, which continuously outputs text tokens in an autoregressive manner to obtain the complete first text output.
7. The multi-task network road object detection method based on subtitle-aware pre-training according to claim 1 is characterized in that: The subtitle-aware decoder is an improved Transformer-based decoder, and the improvement includes replacing the cross-attention module of the Transformer-based decoder with a sparse masked self-attention module MSAM.
8. The multi-task network road object detection method based on subtitle-aware pre-training according to claim 2 is characterized in that: The second training of the pre-trained image encoder and multi-task decoder head to obtain the final image encoder and multi-task decoder specifically includes: Acquire a second training data set, where the second training data set includes an autonomous driving scene image and its corresponding target detection annotation data, lane line detection annotation data, and segmentation annotation data of a drivable area; Preprocess and normalize the autonomous driving images in the second training dataset, and adjust the image size to 640x640x3; Input the preprocessed autonomous driving image to the first trained image encoder to obtain feature f3 output by the first downsampling module of the neck network of the image encoder, feature f2 output by the CSP module of the neck network, and feature f1 output by the second downsampling module of the neck network, wherein the dimension of feature f1 is 80x 80x256 x W, the dimension of feature f2 is 40x 40x512x W, and the dimension of feature f3 is 20x 20x 512x W x R; The feature f1 is input into the lane line detection head and the drivable area detection head, and the position of the lane line and the segmentation map of the drivable area are output respectively; Input features f2 and f3 into the object detection head and output the bounding box of the object detection; The output bounding box of the target detection, the position of the lane line, the segmentation map of the drivable area and the target detection annotation data, the lane line detection annotation data and the segmentation annotation data of the drivable area in the second training data set are jointly trained through the second loss function, and the parameters of the image encoder and the multi-task decoder head are updated through the back-propagation algorithm to obtain the final image encoder and multi-task decoder.
9. The multi-task network road object detection method based on subtitle-aware pre-training according to claim 8 is characterized in that: The target detection annotation data includes the position of the target in the autonomous driving scene image and its category label, the lane line detection annotation data includes the specific position and marking information of the lane line, and the drivable area segmentation annotation data represents the drivable area in the image.
10. The multi-task network road object detection method based on subtitle-aware pre-training according to claim 8, characterized in that: The second loss function is: L all =σ1l det +σ2l ll +σ3l da Among them, L all is the second loss function, σ1, σ2, σ3 are balance factors, l det is the target detection loss function, l ll is the lane detection loss function, l da Loss function for driving area segmentation; The target detection loss function is: l det =α1l class +α2l obj +α3l box Among them, α1, α2, α3 are balance factors, l class is the classification loss function, l obj is the target confidence loss function, l box is the bounding box regression loss function; The lane detection loss function, drivable area segmentation loss function, classification loss function, and target confidence loss function all use the cross entropy loss function l ce ; The cross entropy loss function l ce for: Where C is the number of categories, y i is the true label of the sample, p i is the model’s predicted probability for category i; The bounding box regression loss function is: Among them, b = (xb, yb, wb, hb) is the coordinate of the real bounding box, are the coordinates of the predicted bounding box, is the intersection-over-union ratio of the real box and the predicted box, The Euclidean distance between the center points of the two boxes, c is the diagonal length of the smallest enclosing box that encloses the true box and the predicted box, v is the aspect ratio consistency loss, which represents the difference in aspect ratio, and α is a tuning parameter used to control the impact of the aspect ratio loss.
Citation Information
Patent Citations
A Multi-Task Network Road Target Detection Method for Autonomous Vehicle Automated Driving
CN116665176B
Cited By
Brain tumor imaging diagnosis large model pre-training method, diagnosis method and system
CN121306517A