Target detection method in traffic scene based on image text multi-modal feature fusion

By combining the multimodal features of images and text, the DETR object detection network is improved, and the problem of insufficient detection accuracy of traditional methods in complex traffic scenarios is solved, achieving more efficient and accurate traffic scene object detection.

CN119942474APending Publication Date: 2025-05-06CHANGAN UNIV

Patent Information

Application Number
CN202510074522.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional single-modal image object detection methods are difficult to achieve accurate detection in complex traffic scenarios, especially ignoring the importance of text information.

Method used

By acquiring traffic scene images and preprocessing them, text data describing the content of the picture, combining the multimodal features of the image and text, the DETR object detection network is improved, a traffic scene object detection model is built, and it is trained to achieve detection.

Benefits of technology

It significantly improves the accuracy and efficiency of target detection in traffic scenarios, enhances the accuracy and robustness of detection, and is suitable for intelligent transportation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942474A_ABST
    Figure CN119942474A_ABST
Patent Text Reader

Abstract

The invention provides a traffic scene target detection method based on image text multi-modal feature fusion, and the method comprises the steps: obtaining a traffic scene image, carrying out the preprocessing of the traffic scene image, generating text data describing the content of the image according to the image after the preprocessing, and carrying out the detection of a target in a traffic scene. The method comprises the steps of preprocessing a traffic scene image, generating a traffic scene data set based on the preprocessed traffic scene image and corresponding text data, improving a DETR target detection network, constructing a traffic scene target detection model based on the improved DETR target detection network, training the traffic scene target detection model based on the traffic scene data set, and obtaining a traffic scene target detection result. And obtaining a final traffic scene target detection model, and inputting the to-be-detected image into the final traffic scene target detection model to obtain a detection result. The method is high in detection speed and robustness, can make up the defects of a conventional image detection method, can remarkably improve the detection accuracy and efficiency, and has a wide application prospect in an intelligent traffic system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation systems, and in particular to a method for detecting targets in traffic scenes based on image and text multimodal feature fusion. Background Art

[0002] With the development of intelligent transportation systems, vision-based traffic scene object detection technology has important applications in traffic management, autonomous driving and other fields. However, due to the complexity of traffic scenes, such as light changes, occlusions, and target diversity, traditional single-modality image target detection methods are difficult to meet the needs of accurate detection in complex scenes.

[0003] In the existing technology, convolutional neural networks (CNNs) are mainly used to extract features and detect targets in images. However, such methods often ignore text information, which is of great reference value for the detection and recognition of traffic targets. Therefore, how to combine the multimodal features of images and texts to improve the accuracy of target detection in traffic scenes has become an urgent problem to be solved. Summary of the invention

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A method for object detection in traffic scenes based on image and text multimodal feature fusion, comprising:

[0006] Acquire a traffic scene image and preprocess it, generate text data describing the image content based on the preprocessed image, and generate a traffic scene dataset based on the preprocessed traffic scene image and the corresponding text data;

[0007] Improve the DETR target detection network and build a traffic scene target detection model based on the improved DETR target detection network;

[0008] The traffic scene target detection model is trained based on the traffic scene data set to obtain the final traffic scene target detection model;

[0009] The image to be detected is input into the final traffic scene target detection model to obtain the detection result.

[0010] Preferably, a traffic scene image is obtained and preprocessed, and text data describing the image content is generated according to the preprocessed image, and a traffic scene dataset is generated based on the preprocessed traffic scene image and the corresponding text data, specifically:

[0011] Select some images from the COCO dataset according to the six categories of car, person, bus, bicycle, train, and motorcycle;

[0012] Perform denoising on the selected image;

[0013] Perform illumination equalization on the denoised image;

[0014] Perform data enhancement on the image after illumination equalization processing;

[0015] Annotate the data-augmented images, annotate each image with five sentences of English text describing the image content, and use them as text data;

[0016] Generate a traffic scene dataset based on the processed images and their corresponding text data.

[0017] Preferably, the selected image is subjected to denoising processing, specifically:

[0018] Get the selected image;

[0019] The selected images are denoised based on the pre-trained AFBNet denoising model.

[0020] Preferably, the denoised image is subjected to illumination equalization processing, specifically:

[0021] Get the denoised image;

[0022] Calculate the average brightness of the image;

[0023] Calculate the corresponding correction coefficient according to each pixel point of the image;

[0024] The brightness value corresponding to each pixel of the image is corrected based on the correction coefficient.

[0025] Preferably, the DETR target detection network is improved, specifically:

[0026] Based on the traditional DETR target detection network, a text feature extraction branch is introduced, and the text feature extraction branch is spliced ​​with the image feature extraction branch of the traditional DETR target detection network. After passing through the Transformer encoder-decoder group of the traditional DETR target detection network, the object bounding box in the image and the position distribution corresponding to the words or phrases in the text are finally output.

[0027] Preferably, the text feature extraction branch is a BETY model.

[0028] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0029] The present invention provides a method for target detection in a traffic scene based on image-text multimodal feature fusion, the method comprising: obtaining a traffic scene image, and preprocessing it, generating text data describing the content of the image according to the image after the preprocessing, generating a traffic scene data set based on the preprocessed traffic scene image and the corresponding text data, improving a DETR target detection network, and constructing a traffic scene target detection model based on the improved DETR target detection network, training the traffic scene target detection model based on the traffic scene data set, obtaining a final traffic scene target detection model, inputting the image to be detected into the final traffic scene target detection model, and obtaining a detection result. The present invention has high detection speed and robustness, can make up for the shortcomings of traditional image detection methods, can significantly improve the accuracy and efficiency of detection, and has broad application prospects in intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0031] Figure 2 This is a schematic diagram of the network structure of the denoising model AFBNet;

[0032] Figure 3 This is a schematic diagram of the noise estimation subnetwork structure;

[0033] Figure 4 Schematic diagram of improved skip connection structure;

[0034] Figure 5 This is a schematic diagram of the improved DETR target detection network;

[0035] Figure 6 This is a schematic diagram of the BERT model structure;

[0036] Figure 7 Schematic diagram of BERT’s input composition;

[0037] Figure 8 This is a schematic diagram of the specific implementation process of BERT;

[0038] Fig. 9 Simplified visualization of alignment loss for comparison;

[0039] Fig.10 Show schematic diagrams for some data sets;

[0040] Fig.11 The following is a schematic diagram for visualizing the comparison results. DETAILED DESCRIPTION

[0041] Figure 1 A flow chart of a method provided by an embodiment of the present invention, such as Figure 1As shown, the present invention provides a method for target detection in traffic scenes based on image-text multimodal feature fusion, comprising:

[0042] Step 100: Acquire a traffic scene image, preprocess it, generate text data describing the image content according to the preprocessed image, and generate a traffic scene dataset based on the preprocessed traffic scene image and the corresponding text data;

[0043] Step 200: improving the DETR target detection network, and building a traffic scene target detection model based on the improved DETR target detection network;

[0044] Step 300: training a traffic scene target detection model based on a traffic scene data set to obtain a final traffic scene target detection model;

[0045] Step 400: Input the image to be detected into the final traffic scene target detection model to obtain the detection result.

[0046] In step 100, a traffic scene image is obtained and preprocessed, and text data describing the image content is generated according to the preprocessed image. A traffic scene dataset is generated based on the preprocessed traffic scene image and the corresponding text data, specifically:

[0047] Step 101: Select some images from the COCO dataset according to the six categories of car, person, bus, bicycle, train, and motorcycle;

[0048] Step 102: performing denoising processing on the selected image;

[0049] Step 103: Performing illumination equalization processing on the denoised image;

[0050] Step 104: performing data enhancement processing on the image after the illumination equalization processing;

[0051] Step 105: annotate the data-enhanced images, annotate each image with five sentences of English text describing the image content, and use them as text data;

[0052] Step 106: Generate a traffic scene dataset based on the processed image and the corresponding text data.

[0053] In step 102, the selected image is subjected to denoising, specifically:

[0054] Get the selected image;

[0055] The selected image is denoised based on the pre-trained AFBNet denoising model. The AFBNet denoising model is introduced as follows:

[0056] The structural diagram of the AFBNet denoising model is as follows Figure 2 As shown in the figure, it consists of a noise estimation subnetwork and a denoising network based on attention mechanism and feature fusion.

[0057] The noise estimation subnetwork is as follows Figure 3 As shown in the figure, five full convolutional layers are used to estimate the noise level of the input noisy image. In each convolutional layer, the number of feature channels is set to 32, the convolution kernel size is 3×3, BN and pooling operations are not used, and padding operations are performed after each layer to ensure that the feature output size of each layer is the same. Finally, the estimated noise level map is output together with the input noisy image as the input of the denoising network.

[0058] The main body of the denoising network is a 16-layer U-Net. In the upsampling process of U-Net, a spatial channel attention mechanism is added before the jump connection in the upsampling process. By weighting the feature map in the channel and spatial dimensions, it is ensured that key information is not lost. Only attention-related features are merged through the attention mechanism, and complementary information features are extracted and fused, and the input of the jump connection is redefined. In addition, although both maximum pooling and average pooling have downsampled the data, maximum pooling pays more attention to feature selection, selects features with better classification recognition, and provides nonlinearity, similar to nms (non-maximum suppression). On the one hand, it can suppress noise, and on the other hand, it can improve the significance of the feature map in the region (screened maximum value). The present invention uses the Leaky ReLU activation function to replace the ReLU activation function, which makes the information of the negative axis not completely lost by adding a small positive slope to the negative semi-axis, ensuring that the weight of the neuron in the model training process will still be updated when the input is less than 0, solving the problem of neuron death. In this stage, the output of the noise estimation subnetwork is used as input, that is, the noisy image and the noise estimation level, and the denoising network ensures the denoising effect by denoising a specific noise level.

[0059] In order to retain more image feature details to improve the denoising effect, more useful features are focused on and more useless features are suppressed through spatial and channel attention at different scales. The features in the decoder and the features connected to the encoder are processed with attention, and then spliced ​​with upsampling. The feature map obtained after the attention mechanism will contain the importance information of different spatial positions, allowing the model to focus on certain target areas.

[0060] The present invention improves the skip connection part in the denoising network U-Net and adds the CBAM attention mechanism to the skip connection in the upsampling process. Specifically, the attention module is added before the skip connection, and the features in the decoder and the features connected from the encoder are processed with attention, and then spliced ​​with the upsampling. Attention is used to merge only the features related to attention, suppress useless features, extract complementary information and fuse them, and redefine the input of the skip connection. The redefined skip connection module is as follows: Figure 4 shown.

[0061] In step 103, the denoised image is subjected to illumination equalization processing, specifically:

[0062] Get the denoised image;

[0063] Calculate the average brightness μ of the image as:

[0064]

[0065] Where M is the total number of horizontal pixels in the image, N is the total number of vertical pixels in the image, I(x, y) is the input brightness value of the pixel (x, y), x is the horizontal coordinate of the pixel (x, y), and y is the vertical coordinate of the pixel (x, y);

[0066] According to each pixel point (x, y) of the image, the corresponding correction coefficient γ(x, y) is calculated as:

[0067]

[0068] Wherein, ε = (μ / k) is the normalization coefficient, and k is a natural number between 120 and 136;

[0069] The brightness value corresponding to each pixel of the image is corrected based on the correction coefficient according to the following formula:

[0070]

[0071] Where O(x, y) is the output brightness value of the pixel (x, y), c is a scale factor used to control the global brightness change and is a positive real number not greater than 1.

[0072] Step 200, improve the DETR target detection network, specifically:

[0073] The traditional DETR target detection network is an existing technology, so it will not be introduced in detail;

[0074] Based on the traditional DETR target detection network, a text feature extraction branch is introduced, and the text feature extraction branch is spliced ​​with the image feature extraction branch of the traditional DETR target detection network. After passing through the Transformer encoder-decoder group of the traditional DETR target detection network, the object bounding box in the image and the position distribution corresponding to the words or phrases in the text are finally output.

[0075] Detailed description of this section:

[0076] Since data of different modalities are complementary, the present invention introduces a text feature extraction branch based on DETR, aiming to improve the target detection performance by utilizing the feature fusion between different modalities. The overall network retains the image feature extraction branch in DETR. The backbone network extracts image features and processes them into feature sequences, which are then combined with two-dimensional position encoding into a plane vector to preserve spatial information. The text part uses a pre-trained BERT model to extract text features, linearly projects the features of the two modalities, maps them to the same feature space, and splices them. The feature processing part uses the Transformer encoder-decoder group in DETR, and finally outputs the object bounding box in the image and the position distribution corresponding to the words or phrases in the text;

[0077] The improved DETR target detection network structure is as follows Figure 5 As shown;

[0078] Introduction to the text feature extraction branch: The newly introduced text feature extraction branch uses the BERT model (Bidirectional Encoder Representations from Transformers). The structure of BERT is as follows: Figure 6 As shown in the figure. In the process of fusing text information (generally provided in the form of sentences) for multimodal target detection, the words in the text often contain specific information related to the location of objects in the image, which can effectively assist the network to more accurately locate and identify targets in the image. The text features are passed through BERT to generate a vector sequence of the same size as the input.

[0079] Before being input into BERT, the original text sequence needs to go through specific preprocessing steps. The first is token embeddings, which converts each word into a vector of fixed dimension. The input text is first segmented into a sequence of tokens by a tokenizer, and each word in the sentence is a representation vector. The representation sequence is converted into a one-dimensional vector by looking up a pre-trained embedding matrix, which provides a fixed-size vector representation for each representation. Next is segment embeddings, which is used to distinguish between sentences when BERT's input is multiple sentences. An additional embedding is added to each representation to indicate which sentence it belongs to. The last is position embeddings, which is equivalent to position encoding. This step ensures the order of the text sequence. The BERT model provides an information encoding representing the sequence order for the character / word at each position. The input of BERT is obtained by summing these three embeddings. The input composition of BERT is as follows Figure 7 shown.

[0080] The core of BERT is composed of multiple stacked Transformer encoder layers. Each encoder layer contains a self-attention module and a feedforward neural network. The role of the self-attention module is to allow the model to focus on representations at different positions when processing a sequence and assign attention weights to different representations, thereby obtaining feature relationships in the input sequence. The feedforward neural network further transforms the output of the self-attention mechanism to extract higher-level features. The specific implementation process of BERT is as follows Figure 8 shown.

[0081] The processed image and text feature vectors are concatenated in the sequence dimension to produce a single image and text feature sequence. This concatenation helps the model consider both image and text information and fuse them in a shared representation space, thereby achieving a joint understanding of image and text. In the encoder of DETR, enhanced extraction of image and text fusion features is performed. In the decoder part, the target query and the output of the encoder are also used as input. Finally, the output of the decoder is used to predict each target box in the image. Compared with other target detection algorithms, DETR omits the constraints of generating anchor boxes and the non-maximum suppression NMS post-processing steps, simplifying the entire target detection process. However, DETR requires a very long training cycle to converge, resulting in low operating efficiency, and the detection effect of small targets and occluded objects in the image needs to be improved.

[0082] The network of the present invention introduces a text feature extraction branch based on DETR, so in addition to the loss function of DETR, the contrastive alignment loss is introduced in the loss function part for image and text alignment. The contrastive alignment loss forces the alignment between the feature representation output by the decoder and the text representation output by the encoder. It reflects the similarity between the aligned target query and the representation. Its simple visualization Fig. 9 shown.

[0083] Introducing this loss ensures that the encoding of image objects and their corresponding (text) tags are closer than the encoding of unrelated text tags, because it is not only based on position information, but directly acts on the representation. More specifically, the contrast alignment loss for all objects is as follows:

[0084]

[0085] By symmetry, the contrastive loss for all representations is given by:

[0086]

[0087] The average of these two loss functions is taken as the contrastive alignment loss of the network, where is a series of i The representation vector to be aligned, is a set of targets to be aligned with a given representation vector t1, N is the maximum number of targets, L is the maximum number of representations, and t is a hyperparameter.

[0088] The total loss function of the network of the present invention is composed of L in DETR ba It is summed with the comparison alignment prediction loss.

[0089] The present invention was experimented and the experimental results were analyzed:

[0090] The dataset constructed by the present invention includes 6572 training sets and 1453 test sets, and the composition is as follows: Fig.10 shown.

[0091] The evaluation indicators used in the present invention are shown in Table 1. The higher the values ​​of these evaluation indicators, the better the algorithm performance.

[0092] Table 1 Multimodal target detection evaluation index table

[0093] Evaluation indicators Interpretation <![CDATA[AP 50 ]]> Average precision at IoU=0.5 <![CDATA[AP 75 ]]> Average precision at IoU=0.75 <![CDATA[AP s ]]> Average precision for small objects <![CDATA[AP m ]]> Average precision for medium-sized objects <![CDATA[AP l ]]> Average precision for large-sized objects mAP Take the average precision of all categories

[0094] In the experiment, in the image feature extraction part, ResNet101 is used as the backbone network to extract the required feature vectors; in the text feature extraction part, the pre-trained BERT is used as the text feature extractor. All experiments are built in Python3.8, Pytorch1.9.0, Cudal1.1 environment. The experiment is conducted on a cloud server equipped with GeForce RTX4090 and 24GB of video memory. The network training parameters are as follows:

[0095] Table 2 Network training parameters

[0096] parameter index Initial learning rate 0.01 Weight decay coefficient 0.0005 Optimization methods AdamW Batch size 64 Iterations 200 Training set Homemade dataset Test Set Homemade dataset

[0097] In order to prove the effectiveness of the model of the present invention, the present invention conducted experiments on the traffic scene dataset and compared it with the DETR model. The experimental results are shown in Table 3.

[0098] Table 3 Comparison results between the method of the present invention and DETR

[0099]

[0100] As can be seen from Table 3, the method proposed in this invention effectively integrates text information with the image object detection process, thereby enhancing the performance of image object detection. Specifically, the detection performance of large-scale objects has been greatly improved due to the introduction of text information. l From 55.0% to 62.3%; the detection performance of medium-sized objects also benefits from the fusion of text information, and its AP m From 45.9% to 49.5%; however, for small-scale targets, the performance improvement is relatively limited. This phenomenon may be due to the fact that in the process of extracting image feature vectors, the layer used to extract image feature vectors pays more attention to the information of large-scale and medium-scale targets, while the information extraction of small-scale targets is relatively small, resulting in limited improvement in small-scale target detection. Compared with the DETR algorithm, the mAP of the method of the present invention is improved by 3.8%. The experimental results show the rationality and effectiveness of the method of the present invention.

[0101] In order to further explore the role of text information in image target detection, this paper conducts a visual comparison of the performance of the DETR algorithm and the method proposed in this chapter on the data set. The results are as follows: Fig.11 As shown. Fig.11In the figure, the actual detection results of DETR are shown on the left, and the detection results of the method of the present invention are shown on the right. Through comparative analysis, it is found that the DETR algorithm has several problems in the detection process: on the one hand, some targets are not effectively identified, such as pedestrians on the street in the first image are not included in the detection range and a row of vehicles behind the person in the third image is also missed; on the other hand, some targets are misdetected, such as the train in the second image is mistakenly marked as a bus. In contrast, the method proposed in the present invention significantly improves the accuracy of detection by integrating text information, and the missed detection and wrong detection situations are effectively improved.

[0102] The present invention compares the improved DETR network with the single-stage target detection network SSD, the two-stage target detection network Faster-RCNN and YOLOv5, completes the training of each network model under the same GPU environment, and all networks are trained and tested on the data set. SSD and Faster-RCNN both use ResNet50 as the backbone network. The detailed experimental results are shown in Table 4.

[0103] Table 4 Evaluation results of different target detection networks

[0104]

[0105] As can be seen from the table, compared with the above network, the detection accuracy of the method of the present invention on large, medium and small objects is the best, with a mAP of 44.9%. This proves the effectiveness of the image-text feature fusion target detection model proposed in the present invention. It fully verifies that the use of complementary information between different data modalities can effectively enhance the detector's understanding of image content, thereby improving its ability to identify and locate targets in complex scenes.

Claims

1. A method for target detection in traffic scenes based on image and text multimodal feature fusion, characterized in that: include: Acquire a traffic scene image and preprocess it, generate text data describing the image content based on the preprocessed image, and generate a traffic scene dataset based on the preprocessed traffic scene image and the corresponding text data; Improve the DETR target detection network and build a traffic scene target detection model based on the improved DETR target detection network; The traffic scene target detection model is trained based on the traffic scene data set to obtain the final traffic scene target detection model; The image to be detected is input into the final traffic scene target detection model to obtain the detection result.

2. The target detection method in traffic scenes based on image-text multimodal feature fusion according to claim 1 is characterized in that: Obtain traffic scene images and preprocess them. Generate text data describing the image content based on the preprocessed images. Generate a traffic scene dataset based on the preprocessed traffic scene images and the corresponding text data. Specifically: Select some images from the COCO dataset according to the six categories of car, person, bus, bicycle, train, and motorcycle; Perform denoising on the selected image; Perform illumination equalization on the denoised image; Perform data enhancement on the image after illumination equalization processing; Annotate the data-augmented images, annotate each image with five sentences of English text describing the image content, and use them as text data; Generate a traffic scene dataset based on the processed images and their corresponding text data.

3. The target detection method in traffic scenes based on image-text multimodal feature fusion according to claim 2 is characterized in that: Perform denoising on the selected image, specifically: Get the selected image; The selected images are denoised based on the pre-trained AFBNet denoising model.

4. The target detection method in traffic scenes based on image-text multimodal feature fusion according to claim 2 is characterized in that: Perform illumination equalization on the denoised image, specifically: Get the denoised image; Calculate the average brightness of the image; Calculate the corresponding correction coefficient according to each pixel point of the image; The brightness value corresponding to each pixel of the image is corrected based on the correction coefficient.

5. The method for target detection in traffic scenes based on image-text multimodal feature fusion according to claim 1 is characterized in that: Improve the DETR target detection network, specifically: Based on the traditional DETR target detection network, a text feature extraction branch is introduced, and the text feature extraction branch is spliced ​​with the image feature extraction branch of the traditional DETR target detection network. After passing through the Transformer encoder-decoder group of the traditional DETR target detection network, the object bounding box in the image and the position distribution corresponding to the words or phrases in the text are finally output.

6. The method for target detection in traffic scenes based on image-text multimodal feature fusion according to claim 5 is characterized in that: The text feature extraction branch is a BETY model.

Citation Information

Patent Citations

  • Method and device for carrying out illumination equalization processing on scene image, computer device and computer storage medium

    CN110570384A

  • Traffic scene generation type image description method

    CN117173450A

  • Infrared small target detection method based on scene text information guidance

    CN118762364A

Cited By

  • Power distribution room scene training data augmentation method based on multi-modal image editing

    CN121527243A