Vehicle Target Detection Method Based on Multimodal Two-Stage Refinement of Binding Relationships

By using a two-stage refinement binding relationship method with multimodal data in vehicle target detection, the problem of poor detection effect in complex environments is solved, and higher detection performance and stability are achieved.

CN119625657BActive Publication Date: 2025-06-20江淮前沿技术协同创新中心 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411701664.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-20
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

The existing vehicle target detection technology has poor detection results under factors such as complex road background, bad weather, vehicle scale and distance changes, and vehicle occlusion, especially the serious missed detection of small targets, and the large gap between the data set and the real world traffic image, resulting in a decrease in detection accuracy and the unstable binding relationship between the query vector and the detection target.

Method used

A vehicle object detection method based on multimodal dual-stage refinement of binding relationship is adopted. By inputting RGB, infrared and grayscale images, multi-scale features are extracted, and aligned and weighted fusion is integrated in the depth dimension to generate a unified multimodal multi-scale feature map. The first stage is dynamic sampling and query update, and the second stage is to optimize the binding relationship by refining the binding relationship module to improve detection stability.

Benefits of technology

It effectively improves the performance and generalization capabilities of vehicle target detection, and can accurately detect vehicles in a variety of harsh and complex environments, reduces small target missed detection, and improves the stability and accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625657B_ABST
    Figure CN119625657B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle target detection method based on multi-modal two-stage refinement of binding relationships. This method uses multi-modal input of RGB, grayscale, and infrared images, extracts multi-scale features through independent networks, and fuses between multi-scale features of different modalities to construct a multi-modal multi-scale feature map. In the first stage, a preliminary binding of the detection target and the query is performed, and sampling points are determined through the interaction between the query and the multi-modal multi-scale feature map. Based on this position information, features are sampled from the multi-modal multi-scale feature map, and the sampled features are mapped to the channel dimension of the query. Then, the query is updated through a cross-attention mechanism, and historical queries are introduced to enhance the effect of the current query. In the second stage, score calculation is performed based on the preliminary binding relationship, and a threshold is set to distinguish stable and unstable binding relationships. For targets with scores higher than the threshold, the weight is increased to strengthen the binding; while for targets with scores lower than the threshold, the binding relationship is optimized by rematching the query and the target. Finally, the prediction head decodes the queries that have been strengthened and adjusted, and outputs the category and bounding box of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to vehicle target detection and image recognition technologies, and particularly to a vehicle target detection method based on multi-modal and two-stage refined binding relationships. Background Art

[0002] Vehicle target detection is a core research direction in the fields of traffic engineering and computer vision. It has important research significance and value in intelligent driving, pedestrian detection, and the establishment of an intelligent transportation system, and is even a prerequisite for an intelligent city. Vehicle target detection aims to identify vehicle information participating in traffic in traffic images and determine the position and size of vehicles in the images. Vehicle target detection can detect different types of vehicles on complex roads and perform annotations. According to vehicle information, data statistics can be carried out to generate a traffic flow statistics report. Based on this report, the traffic department can optimize traffic signals, reduce road congestion, and improve road safety.

[0003] Vehicle target detection is a highly challenging task with many difficulties. For example, the complex background of urban roads, the change of lighting conditions, the influence of bad weather (rain, snow, wind, frost), different scales and distances of vehicles, vehicle occlusion, etc. These adverse factors will all interfere with the detection effect. In recent years, researchers at home and abroad have been committed to improving the performance and generalization ability of vehicle target detection. Through various methods such as multi-modal input, multi-scale features, and attention mechanisms, the different dimensions and global context information of images have been effectively utilized, and traffic target vehicles have been detected based on these rich features.

[0004] In summary, the existing vehicle target detection has the following problems:

[0005] 1) The target detection algorithm will be affected by the different appearances and sizes of vehicles on the road, especially the missed detection of small targets is extremely serious.

[0006] 2) The factors of bad weather environment and the complex traffic surrounding environment background that cause ordinary surveillance images to be unclear will affect the detection effect.

[0007] 3) There is a gap between the existing vehicle target detection datasets and real-world traffic images, resulting in a significant reduction in the accuracy of detection results.

[0008] 4) The existing models based on query vectors all adopt the strategy of overall training of query vectors, and there will be a problem that the binding relationship between the query vector and the detection target is unstable. Summary of the Invention

[0009] Object of the Invention: The object of the present invention is to solve the deficiencies existing in the prior art and provide a vehicle target detection method based on multi-modal and two-stage refined binding relationships.

[0010] Technical solution: A vehicle target detection method based on multi-modal two-stage refinement of binding relationships according to the present invention includes the following steps:

[0011] Step (1): Input a set of multi-modal image data, including RGB images, infrared images, and grayscale images, and use a feature extractor to perform feature extraction to generate multi-scale feature maps corresponding to each modality;

[0012] Step (2): For the different-scale feature maps of each modality obtained in step (1), first align the feature maps of each modality in the depth dimension to ensure that all modalities have a unified feature dimension at each scale; subsequently, directly weight and fuse the feature information of the three modalities at each scale to fully exploit the collaborative relationship between multi-modalities, enhance and complement the feature information of different modalities; finally, generate a unified multi-scale feature map with multi-modal characteristics, namely the multi-modal multi-scale feature map F fused ;

[0013] The weights can be configured manually according to the scenario or through training hyperparameters; different detection environments are given different fusion weights for modal features to improve the generalization ability of the model;

[0014] Step (3): Perform the preliminary binding of the first-stage query and the detection target, that is, based on the multi-modal multi-scale feature map F fused in step (2) fused perform dynamic sampling, and the obtained sampled feature F sampled is used to update the query q update ; then enhance the current query EMA through the historical queries in the query cache c ; The specific method is:

[0015] Step (3.1): First, perform dynamic sampling, perform cross-attention on the query and the constructed multi-modal multi-scale feature map to determine the initial position; map the initial position to the feature maps of different scales, select the features near the mapped initial position on each scale feature map, and weight and fuse the features obtained at different scales; the formula is as follows:

[0016]

[0017] P = argmax(A);

[0018] P final = P + offset(P);

[0019] F sampled = sample(F fused , P final );

[0020] In the above formula, A represents the query q and F fusedThe weight of attention, d is the scaling factor, P represents selecting the most relevant position as the initial position using argmax() after interaction, and offset(P) is mapping the initial position to generate offset positions at different scale features, P final represents the final sampling position, F sampled is the sampled feature after the sampling operation sample();

[0021] Step (3.2), for the update operation of the query, perform cross-attention calculation on the sampled feature and the query, and update the relevant features into the query. The formula is as follows:

[0022]

[0023] In the above formula, q update represents the query after attention update, Q q and V q represent Q and V for generating attention of q, is K for generating attention of the sampled feature;

[0024] Step (3.3), for the enhancement operation of the historical query in the query cache, perform EMA weighting on the historical queries according to similarity; set the weight of the current query to γ, and the total weight of the historical queries to 1 - γ. The formula is as follows:

[0025]

[0026] EMA c is the enhanced query, m represents the number of historical queries, H i is the i-th historical query;

[0027] Step (4), perform the second-stage fine-tuning, that is, calculate the score of the binding relationship through the refinement binding relationship module, determine the stability of the preliminary binding in step (3), and perform different processing strategies according to the stability determination result;

[0028] First, set the threshold S of stability. If the score is greater than S, then strengthen the binding relationship and give these stable targets greater weights; if the score is less than S, then perform dynamic matching. Instability indicates that the current matching relationship is not ideal;

[0029] On the premise of retaining stable targets, rematch the remaining targets to explore more reliable and stable binding relationships;

[0030] Step (5), send the output obtained in step (4) into the prediction head to obtain the class class and the bounding box result bounding box.

[0031] Further, the feature extractor in the step (1) includes ResNet-50, EfficientNet, and CNN, which respectively extract multi-scale features from RGB images, infrared images, and grayscale images to obtain multi-scale feature maps of different modalities.

[0032] Further, the construction formula of the multi-modal multi-scale feature map F fused is as follows:

[0033] F aligned = Linear(F input , H target , W target );

[0034]

[0035] In the above formula, Linear represents performing a linear transformation, F aligned is the aligned feature, A, B, and C represent three different modalities, F input is the input modal feature, H target is the dimension of the height of the transformed feature, W target is the dimension of the width of the transformed feature, σ represents the activation function, and α, β, and ω represent weights.

[0036] Further, the refined binding relationship module in the step (4) includes two branches: stable relationship reinforcement and unstable relationship dynamic binding. The specific working process is as follows:

[0037] Step (4.1): Calculate the score for the enhanced query. Set the stability threshold S, and the formula is as follows:

[0038]

[0039] η, δ, respectively represent the classification confidence, the bounding box confidence C IoU and the weight of the matching confidence C match between the query and the detected target. The matching confidence can reflect the binding accuracy between the query and the detected target, and the weight is set to 0.5; the weight η of the classification confidence is 0.3, and the weight δ of the IoU confidence is 0.2;

[0040] Step (4.2): Stability judgment: Set the stability threshold S. For each stability score T, judge the stability of the target according to the stability threshold T: If T ≥ S, it is judged as stable; if T < S, it is judged as unstable;

[0041] Step (4.3), Stability Optimization and Unstable Dynamic Matching: Different training strategies are selected for stable and unstable binding relationships. For stable bindings, the weights of the classification function and the bounding box function are increased, while for unstable bindings, the weight of the penalty term is increased to promote the exploration of stable and reliable binding relationships. The formula is as follows:

[0042]

[0043] Here, θ, λ, and ξ represent the classification loss function bounding box loss function and penalty term weights, respectively.

[0044] Beneficial Effects: The present invention integrates multiple modal data inputs, effectively improving the applicable scenarios and being able to handle various harsh and complex weather environments. The two-stage refinement training process of the present invention conducts a preliminary binding of coarse-grained detection targets and queries in the first stage and performs a finer-grained training based on the first stage. By distinguishing the stability of the binding relationship and performing targeted optimization and iteration, the detection effect and stability of the network are further improved.

[0045] Compared with the prior art, the present invention has the following advantages:

[0046] (1) According to the realistic complex scenarios, the present invention extracts multi-scale features from multi-modal inputs and constructs multi-modal multi-scale features. Based on the multi-scale features, dynamic sampling effectively captures the information of different targets, and while effectively improving the detection effect, it also maintains a relatively small amount of computation.

[0047] (2) The query vector enhancement module constructed in the present invention realizes the dynamic update and optimization of query features by retaining the effective features missing in the current query vector from the historical query vectors. This module can not only effectively make up for the information missing in the query vector, making the update operation a positive feedback behavior, thereby further enhancing the expression ability of the query features.

[0048] (3) The present invention adopts a two-stage training strategy, and conducts coarse-grained and fine-grained training in a progressive manner to optimize the binding relationship between the query vector and the detection target. It effectively solves the problem of unstable matching targets and improves the detection performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the overall classification flowchart of the present invention;

[0050] Figure 2 is the schematic diagram of the network model of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0052] As Figure 1 shown, the vehicle target detection method based on multi-modal two-stage refinement of binding relationships of the present invention includes the following steps:

[0053] Step (1), input a set of multi-modal image data, including RGB images, infrared images, and grayscale images, and use a feature extractor to perform feature extraction to generate multi-scale feature maps corresponding to the modalities;

[0054] Step (2), for the different-scale feature maps of each modality obtained in step (1), align them in the depth dimension of the modality feature maps to ensure that all modalities have a unified feature dimension at each scale; subsequently, directly weight and fuse the feature information of the three modalities at each scale, fully explore the synergistic relationship between multi-modalities, and enhance and complement the feature information of different modalities; finally, generate a unified multi-scale feature map with multi-modal characteristics, that is, the multi-modal multi-scale feature map F fused ;

[0055] Step (3), perform the preliminary binding of the first-stage query and the detection target, that is, based on the multi-modal multi-scale feature map F fused after fusion in step (2), perform dynamic sampling, and the obtained sampled feature F sampled is used to update the query q update ; then enhance the current query EMA through the historical queries in the query cache c ; the specific method is:

[0056] Step (3.1), first perform dynamic sampling, and perform cross-attention on the query and the constructed multi-modal multi-scale feature map to determine the initial position; map the initial position to the feature maps of different scales, select the features near the mapped initial position on each scale feature map, and weight and fuse the features obtained at different scales; the formula is as follows:

[0057]

[0058] P = argmax(A);

[0059] P final = P + offset(P);

[0060] F sampled = sample(F fused , P final );

[0061] In the above formula, A represents the query q and F fusedThe weight of attention, d is the scaling factor, P represents selecting the most relevant position as the initial position using argmax() after interaction, offset(P) is mapping the initial position to generate offset positions at different scale features, and P final represents the final sampling position, and F sampled is the sampled feature after the sampling operation sample();

[0062] Step (3.2), for the update operation of the query, perform cross-attention calculation on the sampled feature and the query, and update the relevant features into the query. The formula is as follows:

[0063]

[0064] In the above formula, q update represents the query after attention update, Q q and V q represent Q and V for generating attention of q, and K is for generating attention of the sampled feature;

[0065] Step (3.3), for the enhancement operation of the historical query in the query cache, perform EMA weighting on the historical queries according to similarity; set the weight of the current query as γ, and the total weight of the historical queries as 1 - γ. The formula is as follows:

[0066]

[0067] EMV c is the enhanced query, m represents the number of historical queries, and H i is the i-th historical query;

[0068] Step (4), perform the second-stage fine-tuning, that is, calculate the score of the binding relationship through the refinement binding relationship module, determine the stability of the preliminary binding in step (3), and perform different processing strategies according to the stability determination result;

[0069] First, set the threshold S of stability. If the score is greater than S, then strengthen the binding relationship and give these stable targets greater weights; if the score is less than S, then perform dynamic matching, indicating that the current matching relationship is not ideal;

[0070] Step (5), send the output obtained in step (4) into the prediction head to obtain the class and the bounding box result.

[0071] The feature extractor in step (1) of this embodiment includes ResNet-50, EfficientNet, and CNN, which respectively extract multi-scale features of RGB images, infrared images, and grayscale images to obtain multi-scale feature maps of different modalities.

[0072] The multi-modal multi-scale feature map F in step (2) of this embodiment fused is constructed as follows:

[0073] F aligned = Linear(F input , H target , W target );

[0074]

[0075] In the above formula, Linear represents a linear transformation, F aligned is the aligned feature, and A, B, and C represent different modalities; F input is the input modal feature, H target is the dimension of the height of the transformed feature, W target is the dimension of the width of the transformed feature, σ represents the activation function, and α, β, and ω represent weights.

[0076] In step (4) of this embodiment, the refined binding relationship module includes two branches: stable relationship enhancement and unstable relationship dynamic binding. The specific working process is as follows:

[0077] Step (4.1): For the enhanced query, calculate the score and set the stability threshold S. The formula is as follows:

[0078]

[0079] η, δ, respectively represent the classification confidence, the bounding box confidence C IoU and the weight of the confidence C matc of the query and the detection target matching;

[0080] Step (4.2): Stability judgment: Set the stability threshold S. For each stability score T, judge the stability of the target according to the stability threshold T: If T ≥ S, it is judged as stable; if T < S, it is judged as unstable;

[0081] Step (4.3): Stability optimization and unstable dynamic matching: Different training strategies are selected for the stable and unstable binding relationships. The weights of the classification function and the bounding box function are increased for stable bindings, and the weight of the penalty term is increased for unstable bindings. The formula is as follows:

[0082]

[0083] Here, θ, λ, and ξ respectively represent the classification loss function the bounding box loss function and the penalty term weights.

[0084] To verify the rationality and effectiveness of the technical solution of the present invention, in this embodiment, the FLIR ADAS dataset is selected for experiments, and AP (Average Precision) is used as the objective evaluation index for the detection results. This embodiment is implemented based on the deep learning framework Pytorch and uses a graphics processing unit (GPU) to accelerate the operation, with 20GB of memory and an Nvidia GeForce GTX2080Ti graphics card.

[0085] The FLIR ADAS dataset provides high-quality pairs of RGB images and infrared images. The dataset mainly contains two categories: pedestrians and vehicles, covering hundreds of thousands of annotated data, and is suitable for object detection tasks in low-light and other complex environments.

[0086] Table 1 Detection Results of the FLIR ADAS Dataset

[0087]

[0088] As shown in Table 1, the experimental results of the present invention perform excellently on the FLIR ADAS dataset. After experimental verification, the object detection model of the present invention is significantly superior to existing similar models in vehicle detection tasks.

Claims

1. A vehicle target detection method based on multi-modal two-stage refinement binding relationship, characterized in that: It includes the following steps: Step (1): Input a set of multimodal image data, including RGB images, infrared images, and grayscale images, and use a feature extractor to perform feature extraction to generate multi-scale feature maps of corresponding modalities. Step (2): For the multi-scale feature maps of each modality obtained in step (1), first align the feature maps of each modality in the depth dimension; subsequently, directly weight and fuse the feature information of the three modalities at each scale. Finally, a unified multi-scale feature map with multi-modal characteristics is generated, namely the multi-modal multi-scale feature map F fused ; Step (3) performs preliminary binding of the first-stage query and detection targets, i.e., based on the multi-modal multi-scale feature map F fused in step (2) fused Perform dynamic sampling and obtain the sampling feature F sampled For update query q upbate ; Then enhance the current query EMA by querying the historical queries in the cache c ; The specific method is: Step (3.1): First, perform dynamic sampling, perform cross-attention on the queried and constructed multimodal multi-scale feature maps to determine the initial position; map the initial position to the feature maps of different scales, select the features near the mapped initial position on each scale feature map, and weight and fuse the features obtained at different scales. The formula is as follows: P = argmax(A); P final =P+offset(P); F samp l ed =sample(F fused ,P final ); In the above formula, A represents query q and F fused The attention weight, Q q is the matrix vector obtained by matrix transformation of query q through random linear layer, K fused is a multi-modal multi-scale feature map F fused The matrix vector is obtained by matrix transformation through random linear layer. T represents vector matrix transpose, d is the scaling factor, P represents the most relevant position selected as the initial position using argmax() after interaction, and offset(P) is the offset position generated by mapping the initial position to features of different scales. final represents the final sampling position, F sampled It is the sampling feature after the sampling operation sample(); Step (3.2): For the update operation of the query, perform cross-attention calculation on the sampled features and the query, and update the relevant features into the query. The formula is as follows: In the above formula, q update represents the query after attention update, Q q and V q Q and V represent the attention generated by q, is the K of sampling feature generation attention; in the attention calculation, Q and V are Q in the formula q and V q , indicating that the query q needs to be transformed into a matrix through a random linear layer, and K is the The sampling features also need to be transformed by random linear layers; Step (3.3): For the enhancement operation of the historical query in the query cache, perform EMA weighting on the historical queries according to similarity; set the weight of the current query to γ, and the sum of the weights of the historical queries to 1 - γ. The formula is as follows: EMA c is the enhanced query, m represents the number of historical queries, H i is the i-th historical query; Step (4): Perform fine-tuning in the second stage, that is, calculate the score of the binding relationship through the refined binding relationship module, determine the stability of the preliminary binding in step (3), and perform different processing strategies according to the stability judgment result. First, set the stability threshold S. If the score is greater than or equal to S, strengthen the binding relationship, identify the target calculated by the stability score and higher than or equal to the set threshold as a stable target, and assign a larger weight to this type of target in subsequent weight calculations; when the stability score is lower than the set threshold S, trigger the dynamic matching mechanism. Instability indicates that the current matching relationship is not ideal. Step (5): Send the output of step (4) into the prediction head to obtain the class label and location information.

2. The vehicle target detection method based on multi-modal two-stage refinement binding relationship according to claim 1 is characterized in that: The feature extractor in step (1) includes ResNet-50, EfficientNet, and CNN, which respectively extract multi-scale features of RGB images, infrared images, and grayscale images to obtain multi-scale feature maps of different modalities.

3. The vehicle target detection method based on multi-modal two-stage refinement binding relationship according to claim 1 is characterized in that: The multi-modal multi-scale feature map F in step (2) fused The construction formula is as follows: F aligned =Linear(F input ,H target ,W target ); In the above formula, Linear means linear transformation, F aligned is the aligned feature, A, B and C represent three different modes, F input is the modal feature of the input, H target is the dimension of the high feature after transformation, W target is the dimension of the feature width after conversion, σ represents the activation function, and α, β, and ω represent weights.

4. The vehicle target detection method based on multi-modal two-stage refinement binding relationship according to claim 1 is characterized in that: The refined binding relationship module in step (4) includes two branches: stable relationship enhancement and unstable relationship dynamic binding. The specific working process is as follows: Step (4.1): Calculate the score for the enhanced query, and set the stability threshold S. The formula is as follows: η represents the weight of the classification confidence; δ represents the bounding box confidence C IoU The weight of Represents the query and detection target matching confidence C match The weight of Step (4.2): Stability judgment: Set the stability threshold S. For each stability score T, judge the stability of the target according to the stability threshold S: If T ≥ S, it is judged as stable; if T < S, it is judged as unstable. Step (4.3): Stability optimization and unstable dynamic matching: Select different training strategies for the stability and instability of the binding relationship. Increase the weights of the classification function and the bounding box function for stable bindings, and increase the weight of the penalty term for unstable ones. The formula is as follows: Here, θ, λ and ξ represent the classification loss functions respectively. Bounding Box Loss Function and penalty items The weight of .

Citation Information

Patent Citations

  • Visual question and answer method based on image target features and multilayer attention mechanism

    CN110287814A

  • Distributed logistics transportation scheduling method, device, equipment and storage medium

    CN111382160A