Adaptive fusion complex traffic environment detection method
The ISMR-YOLO model introduces an infrared detection channel and a multi-level reversible auxiliary supervision mechanism into the YOLOv1m baseline architecture, which solves the shortcomings of single-modal detection, realizes the adaptive fusion and information transmission of multi-modal features, and improves the detection accuracy and robustness of multiple categories of traffic participants in complex traffic environments.
Patent Information
- Application Number
- CN202510483558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, single-modal traffic environment detection methods have difficulty in stably obtaining sufficient target feature information in complex and changeable dynamic traffic environments, resulting in insufficient detection accuracy and robustness, and there are problems of modal contribution imbalance and information accumulation error when multimodal fusion is used.
The ISMR-YOLO model is adopted. By introducing an infrared detection channel into the baseline architecture of the YOLOv1m model, combined with an interactive feature enhancement module, a feature adaptive weight fusion module and a multi-level reversible auxiliary supervision mechanism, adaptive fusion and information transmission of multimodal features are achieved, and the modal feature weights are dynamically adjusted to alleviate the information bottleneck problem.
It significantly improves the accuracy and robustness of multi-category traffic participant detection, adapts to complex urban traffic scenarios, optimizes multimodal detection performance, and improves the generalization ability and robustness of the model in complex environments.
Smart Images

Figure CN120635836A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic environment detection, and in particular relates to an adaptive fusion complex traffic environment detection method. Background Art
[0002] With the continuous advancement of global urbanization, the number and diversity of traffic participants (pedestrians, motor vehicles, non-motor vehicles, etc.) are increasing, and in real-world traffic environments, they often present high dynamics and multi-category mixing, which makes traffic safety issues more prominent. In such dynamic traffic scenarios, achieving accurate, efficient, and stable detection of traffic participants is of great significance to the safe operation of intelligent connected vehicles and public intelligent traffic monitoring systems. However, the challenge of reliable detection of traffic participants comes not only from the characteristics of the categories themselves, but more from the various complex scenarios formed by the combination of traffic participants and the traffic environment. Specifically, complex lighting conditions, dynamic occlusion of targets, and small target size will all have a great impact on the efficiency and accuracy of detection. In particular, in single-modality detection methods (visible light images or infrared images), while visible light image detection can provide rich target feature textures and detailed information in well-lit traffic environments, its performance degrades significantly in low-light or nighttime environments, and even completely loses performance in traffic environments with extreme lighting conditions such as darkness, strong light, fog, haze, or rain and snow. Infrared image detection can obtain information by capturing the target's energy and thermal radiation, and performs well in extreme conditions such as darkness or strong light. However, due to its high signal-to-noise ratio, low resolution, and lack of texture detail, it is often difficult to accurately identify the category of traffic targets with unclear appearance features. Therefore, single-modality detection greatly limits the improvement of target accuracy and robustness in complex traffic environments.
[0003] In recent years, convolutional neural networks (CNNs) have been widely used in image classification and object detection tasks due to their powerful feature learning capabilities and adaptability. CNN-based object detection algorithms can be roughly divided into two categories: single-stage and two-stage algorithms. Among them, the single-stage detection algorithm YOLO (You only look once), first proposed by REDOM et al. in 2016, unifies object classification and localization through regression thinking, achieving end-to-end detection from input image to obtaining the object bounding box position and category. Compared to the typical two-stage algorithm R-CNN series, the YOLO series achieves a better balance between detection accuracy and speed, and has therefore been applied by many scholars to traffic and pedestrian detection.
[0004] Figure 1A typical example of complementary features between visible light and infrared image pairs. The first row of image pairs was captured in a traffic scene exposed to strong light and glare. The visible light image on the left, affected by the strong light and some glare, provides almost no feature information of the target in the red frame. The infrared image on the right can provide corresponding appearance contour information. The second row of image pairs was captured in a well-lit daytime traffic scene. In well-lit conditions, visible light images can provide rich details and texture information about traffic participants, making it easy to identify the target in the red frame as a car. However, the target in the infrared image is difficult to classify due to the complex background and lack of significant thermal radiation differences.
[0005] However, although the above studies have made progress in single-modality detection research, they are also limited by the characteristics of the single modality itself. It is difficult to stably obtain sufficient target feature information in a complex and changing dynamic traffic environment to meet the high accuracy, robustness and stability requirements required for traffic target detection. Figure 1 Demonstrating the unique complementary nature of visible light and infrared images has made it possible to achieve further breakthroughs in the performance of traffic detection models (making it possible to further improve target detection in complex traffic environments), leading to a surge in research into multimodal fusion techniques. Wang et al. proposed a cross-scale iterative attention adversarial fusion network, CrossFuse, which uses attention weights to interact with features of different scales from different modalities, achieving a unique description of visible details.
[0006] The numerous methods and experiments described above demonstrate that the fusion of visible and infrared images is the optimal approach for detecting traffic participants in complex urban traffic environments. However, several key challenges remain to be addressed in multimodal fusion traffic detection. First, the single nature of the research objectives. Most studies focus on a single class of traffic targets, while simultaneous detection of multiple traffic participants, particularly in complex urban traffic scenarios, has not been fully explored. Second, the inherent flaws in the information source processing mechanism. Specifically, the input visible and infrared images typically cannot simultaneously present the most comprehensive target feature information for each. This results in insufficient input of traffic participant feature information from the outset. Furthermore, most CNN-based methods use fixed convolutional kernel weights or standard attention mechanisms, making it difficult to dynamically adjust the importance of multimodal features. This, in turn, can lead to an imbalance in modal contributions during information fusion. Finally, the cumulative error in global information transfer. Deep networks are prone to information bottlenecks during the layer-by-layer feature extraction process. The accumulated deviation in key modal information directly impacts the integrity of the fused features and the accuracy of the model's detection results. Summary of the Invention
[0007] The purpose of the present invention is to provide an adaptive fusion method for complex traffic environment detection, which mainly solves the problem that the existing fixed convolution kernel weights or ordinary attention mechanism are difficult to achieve dynamic adjustment of the importance of multi-modal features, which leads to an imbalance in the contribution of modalities when the traffic environment information is fused.
[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0009] An adaptive fusion complex traffic environment detection method is implemented based on the ISMR-YOLO model; the ISMR-YOLO model includes the YOLOv10m model as a baseline model, an infrared detection channel built additionally in the backbone architecture of the baseline model, an interactive feature enhancement module and a feature adaptive weight fusion module respectively integrated into the three layers of the backbone network dual-channel feature extraction, and a multi-level reversible auxiliary supervision mechanism module integrated into the backbone architecture and input architecture of the baseline model; wherein,
[0010] The infrared detection channel and the existing visible light detection channel together form a dual-channel feature extraction mechanism to support the simultaneous input and processing of infrared and visible light modalities;
[0011] The interactive feature enhancement module and the feature adaptive weight fusion module work together to complete the interactive enhancement and adaptive weight fusion of multimodal features;
[0012] The multi-level reversible auxiliary supervision mechanism module uses reversible gradient auxiliary branches and multi-level auxiliary information integration mechanism to build a reliable auxiliary path for global information transmission.
[0013] Furthermore, in the present invention, the interactive feature enhancement module realizes interactive enhancement between multimodal features through modal collaborative channel enhancement branch and global perception space feature enhancement branch; wherein,
[0014] The modal collaborative channel enhancement branch is used to highlight the important parts of the feature information between modalities by assigning channel weights to the target features of each modality, while effectively reducing the impact of redundant information, thereby achieving the purpose of feature complementarity and enhancement. The specific implementation steps are as follows:
[0015] S11 performs 1×1 convolution operations on the RGB and infrared modal target features of the input traffic status image to reduce the feature dimension and extract channel features:
[0016] W RGB =σ1(F1(X RGB )),
[0017] W IR =σ2(F2(X IR ))
[0018] Among them, X RGB ∈R C×H×W is the RGB feature, X IR ∈R C×H×W are the features of the IR modal target, F1(·) and F 12 (·) is a 1×1 convolution kernel, which is used to compress the channel dimension and extract the global information of the modal features; σ1 and σ2 are Sigmoid activation functions, which are used to normalize the extracted channel features to ensure that the weight range is between [0, 1];
[0019] S12, normalize the weights of the two modalities through the Softmax function to obtain the cross-modal channel weights:
[0020]
[0021] in, and is the normalized channel weight, which indicates the importance of RGB and IR modalities to their respective channel information. The role of Softmax is to ensure that the sum of the weights between the two modalities is 1, emphasizing the complementary relationship between the information of different modalities.
[0022] S13, based on the assigned modal weights, weighted optimization is performed on the features of the other modality, thereby achieving feature complementarity between the modalities:
[0023]
[0024]
[0025] Among them, X' RGB and X' IR Represents the interactively enhanced RGB and IR modal features; represents element-wise weighted operations;
[0026] S14, the interactively enhanced features are added to their own weights, and both sides of the multimodal model obtain the final modality-enhanced features:
[0027] A RGB =X' RGB +W RGB ,
[0028] A IR =X' IR +W IR .
[0029] Furthermore, in the present invention, the global perception spatial feature enhancement branch is used to capture the spatial distribution relationship between modalities and perform interactive enhancement to further enhance the global expression capability of multimodal feature information; the specific implementation steps are as follows:
[0030] S21, use the global pooling operation to perform global spatial interaction on each modal feature to generate compensation features:
[0031]
[0032] Among them, q and v are learnable weight parameters used to balance the contribution between RGB and IR modalities; α4 and α5 are activation functions used to normalize the enhanced features;
[0033] S22, using the compensated modal features, calculate the interactive enhancement features:
[0034]
[0035]
[0036] in, is the interactive enhanced spatial feature with RGB as the modality; F3(·) is a further 1×1 convolution operation used to refine the enhanced features.
[0037] Furthermore, in the present invention, the feature adaptive weight fusion module is used to achieve efficient fusion of multimodal features, and the specific implementation steps are as follows:
[0038] S31, the feature adaptive weight fusion module uses the interactive feature enhancement module to enhance the IR features after interaction and RGB features As input, let its dimensions be C×H×W again; then the modal weights are extracted again through spatial convolution and the weight matrix is rebuilt:
[0039]
[0040]
[0041] Among them, W IR and W RGB The weight matrices of IR and RGB modal features are established respectively, reflecting the importance of each modality in the spatial dimension; Conv IR and Conv RGB It is a spatial convolution operation used to extract the saliency information of modal features again;
[0042] S32, generates dynamically adjusted modal weights through the Softmax function:
[0043]
[0044]
[0045] Among them, norm(W IR ,W RGB ) is the modal weight measurement factor, which is used to balance the weights of infrared and RGB modalities; W1 and W2 are the adjusted modal weight matrices, which represent the importance ratios of infrared and RGB modal features, respectively;
[0046] S33, using the dynamically generated weight matrices W1 and W2, the IR features and RGB features Perform weighted fusion to generate new fusion features:
[0047]
[0048] Where: X is the new feature matrix after fusion, which contains the dynamic interaction and enhancement information of the infrared and visible light modalities; The weight W1 is used to highlight the important appearance feature information of the infrared modality in low-light scenes while suppressing background noise. It indicates that the weight W2 is used to highlight the texture and color information of the RGB modality under sufficient lighting conditions.
[0049] Furthermore, in the present invention, the multi-level reversible auxiliary supervision mechanism consists of three parts: a main branch, a reversible gradient auxiliary branch, and a multi-level auxiliary supervision information module; wherein,
[0050] The main branch is used to implement the calculation of the reasoning stage;
[0051] The reversible gradient auxiliary branch is used to solve the information bottleneck problem caused by the deepening of the neural network. It generates reliable gradients for the loss function by maintaining the integrity of information in the feedforward process.
[0052] The multi-level auxiliary supervision information module targets the deep supervision module by introducing multi-level gradient integration in the architecture of multiple prediction branches to alleviate the problem of insufficient coordination between shallow features and deep features.
[0053] Furthermore, in the present invention, the specific steps for implementing the reversible gradient auxiliary branch are as follows:
[0054] S41, given the key feature information in the input data X, it will be gradually diluted or lost during layer-by-layer feature extraction and spatial transformation, which can be expressed as:
[0055] I(X,X)≥I(X,f θ (X))≥I(X,g φ (f θ(X))),
[0056] Where I represents the interaction information, f and g are transformation functions, θ and φ are the parameters of f and g respectively;
[0057] S42, when the transformation function r has an inverse function v, the data can be transmitted without loss of information during the feedforward process:
[0058] X=v ζ (r ψ (X)),
[0059] I(X,X)=I(X,r ψ (X))=I(X,v ζ (r ψ (X))),
[0060] where ψ and ζ are the parameters of the functions r and v, respectively.
[0061] Furthermore, in the present invention, the multi-level auxiliary supervision information module is implemented as follows: inserting multi-level auxiliary supervision information to perform sparse feature filtering and integration on the gradient information of different prediction heads:
[0062] X=v ζ (r ψ (X)·M),
[0063] Where M is a dynamic binary mask used to eliminate noise features and retain key information.
[0064] Compared with the prior art, the present invention has the following beneficial effects:
[0065] (1) The new multimodal fusion detection model of ISMR-YOLO in the present invention fully integrates the complementary characteristics of visible light and infrared modalities by designing a feature enhancement fusion module and improving the backbone architecture on the baseline model, while eliminating error accumulation to ensure the uniformity of global target feature information, thereby adapting to the task of detecting multiple categories of traffic participants in complex urban traffic scenarios.
[0066] (2) The multimodal feature enhancement and fusion method of the present invention is mainly composed of an interactive feature enhancement module (IFE) and a feature adaptive weight fusion module (SAWF). The IFE module strengthens the expressive power of intra-modal features by combining the two mechanisms of modal collaborative channel enhancement and global perception spatial feature enhancement, while realizing deep interaction of inter-modal information; the SAWF module dynamically adjusts the weight ratio of modal features to give full play to the contour advantages of the infrared modality and the texture advantages of the visible light modality, while effectively suppressing background noise and improving the accuracy and robustness of the fusion features to better adapt to the multimodal detection needs in complex scenarios.
[0067] (3) The present invention solves the common information bottleneck problem in deep networks by constructing a multi-level reversible auxiliary supervision mechanism (MRASM) and introducing reversible gradient auxiliary branches and a multi-level auxiliary information integration mechanism, alleviates the negative impact of feature information error accumulation, and further improves the generalization ability and robustness of the model in complex traffic scenarios.
[0068] (4) Experimental results on three public multimodal datasets (FLIR, LLVIP, and MSRS) show that the proposed algorithm model significantly surpasses the baseline model in terms of average mean average precision (mAP), false positive rate (FPR), and log-averaged miss rate (MR-2), and achieves the best overall performance compared to mainstream models. Ablation experiments also verify the effectiveness and applicability of the proposed module, and visual detection experiments objectively demonstrate its excellent detection performance and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 This is a typical example of the complementary characteristics of visible light and infrared images.
[0070] Figure 2 This is a network structure diagram of the ISMR-YOLO model of the present invention.
[0071] Figure 3 This is the internal structure diagram of the interactive feature enhancement module (IFE) in the present invention.
[0072] Figure 4 This is the internal structure diagram of the feature adaptive weight fusion module (SAWF) in the present invention.
[0073] Figure 5This is a structural diagram of the prior art using path aggregation in an embodiment of the present invention.
[0074] Figure 6 This is a structural diagram of a reversible architecture adopted in the prior art according to an embodiment of the present invention.
[0075] Figure 7 This is a structural diagram of the prior art using deep supervision in an embodiment of the present invention.
[0076] Figure 8 This is a structural diagram of the multi-level reversible auxiliary supervision mechanism adopted in the present invention. DETAILED DESCRIPTION
[0077] The present invention will be further described below with reference to the accompanying drawings and examples. The embodiments of the present invention include but are not limited to the following examples.
[0078] Example
[0079] like Figure 2 As shown, the present invention discloses an adaptive fusion complex traffic environment detection method, which is implemented based on the ISMR-YOLO model and aims to meet the challenges of detecting multiple categories of traffic participants in various complex traffic environments in cities. The ISMR-YOLO model includes the YOLOv10m model as a baseline model. YOLOv10m is one of the most advanced target detection models in the current YOLO series proposed by Wan et al. in 2024. It achieves a significant balance optimization between detection accuracy and speed by introducing convolutional layer design, and is very suitable for improvement as a baseline model. Structurally, YOLOv10m continues the classic design framework of the YOLO series, and includes four main modules: input layer (Input), backbone feature extraction layer (Backbone), neck feature fusion layer (Neck) and detection layer (Head). It is worth noting that the partial self-attention module (PSA) proposed by YOLOv10m has an outstanding ability to limit itself to the low-resolution stage while optimizing the global computational complexity. This makes the YOLOv10m model stand out from previous models of the same series and is also the main baseline architecture of the model in this invention.
[0080] The ISMR-YOLO model proposed in this invention further improves the comprehensive performance of multi-modal and multi-category traffic target detection by making multiple improvements to the baseline model. First, in order to achieve efficient processing and fusion of multimodal data, ISMR-YOLO builds an additional infrared detection channel in the Backbone architecture of the baseline model, which together with the original visible light detection channel constitutes a dual-channel feature extraction mechanism to support the synchronous input and processing of infrared and visible light modalities. Secondly, in order to strengthen the feature interaction and fusion between modalities, the interactive feature enhancement module (IFE) and the self-adaptive weighted fusion module (SAWF) are respectively integrated into the three levels of the dual-channel feature extraction of the backbone network. The two modules work together to complete the interactive enhancement and adaptive weight fusion of multimodal features, greatly improving the expressive power and robustness of the fused features. Finally, ISMR-YOLO incorporates a multi-level reversible auxiliary supervision mechanism (MRASM) into its overall architecture. By designing reversible gradient auxiliary branches and a multi-level auxiliary information integration mechanism, it constructs a reliable auxiliary path for global information transmission, effectively alleviating the information bottleneck and error accumulation problems in deep networks.
[0081] Due to the limitations of visible light and infrared imaging systems, multimodal targets of multiple categories of traffic participants often cannot simultaneously present the most complete feature appearance and details. In some special traffic scenarios, one of the modes may even completely lose its information contribution, which causes the detection model to lose target information at the beginning of data fusion. To address this phenomenon, this embodiment proposes an IFE module to solve it. Figure 3 As shown in Figure 2, the IFE module achieves interactive enhancement between multimodal features through the Modality-Specific Channel Enhancement (MSCE) branch and the Global-Aware Spatial Feature Enhancement (GASF) branch. The IFE module is designed to enhance the complementarity between visible light and infrared image modalities, suppress intermodal noise, and achieve deep interaction of intermodal information. It also provides higher-quality feature information expression input for the adaptive weight fusion of multimodal features.
[0082] In this embodiment, the modal collaborative channel enhancement branch is used to highlight the important parts of the inter-modal feature information by assigning channel weights to the target features of each modality, while effectively reducing the impact of redundant information, thereby achieving the purpose of feature complementarity and enhancement. The specific implementation steps are as follows:
[0083] Initial extraction of feature weights: Perform 1×1 convolution operations on the features of the RGB and infrared modal targets of the input traffic status image to reduce the feature dimension and extract channel features:
[0084] W RGB =σ1(F1(X RGB )),
[0085] W IR =σ2(F2(X IR ))
[0086] Among them, X RGB ∈R C×H×W is the RGB feature, X IR ∈R C×H×W are the characteristics of the IR modal target, F1(·) and F 12 (·) is a 1×1 convolution kernel used to compress the channel dimension and extract global information of the modal features. σ1 and σ2 are sigmoid activation functions used to normalize the extracted channel features to ensure that the weight range is between [0, 1]. This step can extract the most important channel information in each modality and provide a basis for the next step of modal weight allocation.
[0087] Normalization and distribution of modal weights: The weights of the two modalities are normalized using the Softmax function to obtain the cross-modal channel weights:
[0088]
[0089] in, and is the normalized channel weight, which indicates the importance ratio of RGB and IR modalities to their respective channel information. The role of Softmax is to ensure that the sum of the weights between the two modalities is 1, emphasizing the complementary relationship between the information of different modalities. The normalized weight distribution makes the importance ratio of the two modalities at the channel level clearer, avoiding the problem of information imbalance during feature fusion, and effectively reducing the negative impact of redundant information.
[0090] Feature interaction and enhancement: Based on the assigned modality weights, the features of the other modality are weighted and optimized to achieve feature complementarity between the modalities:
[0091]
[0092]
[0093] Among them, X' RGB and X' IR Represents the interactively enhanced RGB and IR modal features; Represents an element-level weighted operation. The interactively enhanced features are added to their own weights, and both sides of the multimodal model obtain the final modality-enhanced features:
[0094] A RGB =X' RGB +W RGB ,
[0095] A IR =X' IR +W IR .
[0096] Through feature interaction and enhancement operations, the feature information of RGB and IR modalities are effectively complemented and enhanced, laying the foundation for rich and accurate target feature detail information for the subsequent global perception spatial feature enhancement process.
[0097] In this embodiment, the global perception spatial feature enhancement branch is used to capture the spatial distribution relationship between modalities and perform interactive enhancement to further enhance the global expression capability of multimodal feature information. The specific implementation steps are as follows:
[0098] Feature space compensation: Use global pooling operations to perform global spatial interactions on each modality feature to generate compensation features:
[0099]
[0100] Among them, q and v are learnable weight parameters used to balance the contributions between RGB and IR modalities; α4 and α5 are activation functions used to normalize the enhanced features; the global spatial interaction operation can capture the significant areas of each modality feature in the spatial distribution, providing accurate information representation for subsequent feature fusion.
[0101] Intermodal spatial feature interaction: Using the compensated modal features, we can calculate the interaction enhancement features:
[0102]
[0103]
[0104] in, The IFE module effectively enhances the interaction and enhancement of the two modal features during the feature extraction process. This provides more complete and reliable feature information for the subsequent adaptive fusion of multimodal features, thereby improving model detection accuracy.
[0105] Traffic participants often present complex and diverse scenes in their urban space operations. In multimodal fusion tasks, the feature representation capabilities of different modalities will change constantly due to changes in traffic scenarios. Traditional convolution usually uses fixed convolution kernel weights, so it is impossible to dynamically adjust the weight distribution according to the characteristics of traffic participants in the input modality. In order to further improve the robustness and stability of the baseline model in the multimodal feature fusion stage, the present invention designs a Self-Adaptive Weighted Fusion (SAWF) module, such as Figure 4 As shown in Figure 1, it is used to achieve efficient fusion of multimodal features. The SAWF module aims to adaptively optimize the fusion quality of infrared and RGB modal features by dynamically readjusting the weight distribution of different modal features during the fusion process, thereby fully leveraging the contour advantages of the infrared modality in low-light or occluded scenes and the detail and texture advantages of the RGB modality in well-lit conditions. The specific implementation principles of the SAWF module are as follows:
[0106] Generation of weight matrix: Feature adaptive weight fusion module uses interactive feature enhancement module to enhance the IR features after interaction and RGB features As input, let its dimensions be C×H×W again; then the modal weights are extracted again through spatial convolution and the weight matrix is rebuilt:
[0107]
[0108]
[0109] Among them, W IR and W RGB The weight matrices of IR and RGB modal features are established respectively, reflecting the importance of each modality in the spatial dimension; Conv IR and Conv RGB It is a spatial convolution operation used to extract the saliency information of modal features again;
[0110] Generate dynamically adjusted modal weights through the Softmax function:
[0111]
[0112]
[0113] Among them, norm(W IR ,W RGB ) is the modal weight measurement factor, which is used to balance the weights of infrared and RGB modalities; W1 and W2 are the adjusted modal weight matrices, which represent the importance ratios of infrared and RGB modal features, respectively; through the above calculations, the dynamically generated weight matrix can timely reflect the feature priorities of the two modalities in different traffic scenarios.
[0114] Dynamic feature fusion: IR features are fused using dynamically generated weight matrices W1 and W2. and RGB features Perform weighted fusion to generate new fusion features:
[0115]
[0116] Where: X is the new feature matrix after fusion, which contains the dynamic interaction and enhancement information of the infrared and visible light modalities; The weight W1 is used to highlight the important appearance feature information of the infrared modality in low-light scenes while suppressing background noise. The weight W2 emphasizes the texture and color information of the RGB modality under well-lit conditions. The fusion feature X of the two achieves adaptive dynamic balance and effective complementary fusion in the multimodal feature space.
[0117] By incorporating the adaptive fusion strategy of the SAWF module, not only is the quality of feature fusion significantly optimized, providing a more expressive and robust feature representation for high-precision multimodal traffic participant target detection, but it also provides high-quality feature information input for subsequent target detection tasks, thereby effectively improving the accuracy, robustness and stability of the baseline model in detecting multimodal traffic participants in various complex urban traffic scenarios.
[0118] In the process of transmitting multimodal information about traffic participants in complex urban traffic environments, deep neural networks need to perform layer-by-layer feature extraction and spatial transformation on the input modal features. However, this processing method causes important information in the input data to be gradually diluted or lost during the feedforward process, thereby causing an "information bottleneck" problem. At the same time, as the number of network layers increases, the semantic information in the features may suffer irreversible loss, resulting in incomplete or biased gradients during the model update process. This unreliable gradient will significantly weaken the model's ability to detect traffic participants in complex traffic scenarios, especially under conditions of changing lighting, target occlusion, and complex backgrounds, resulting in missed detection of key targets.
[0119] like Figures 5 to 8As shown in the figure, in order to solve the above-mentioned problem of unreliable gradient caused by information bottleneck and feature loss, the present invention proposes a multi-level reversible auxiliary supervision mechanism (MRASM). MRASM mainly consists of three parts: the main branch, the reversible gradient auxiliary branch and the multi-level auxiliary supervision information module. Figure 8 It can be seen that MRASM only relies on the main branch for calculations during the inference phase, so it does not introduce any additional inference costs. The design of the reversible gradient auxiliary branch aims to solve the information bottleneck problem caused by the deepening of the neural network. By maintaining the integrity of the information in the feedforward process, a reliable gradient is generated for the loss function, thereby effectively improving the stability and transmission quality of the network. The multi-level auxiliary supervision information module improves the common error accumulation problem of the deep supervision mechanism. By introducing multi-level gradient integration in the architecture of multiple prediction branches, it effectively alleviates the problem of insufficient coordination between shallow features and deep features, and is suitable for the optimization needs of the lightweight baseline model YOLOv10m. Next, this embodiment will introduce in detail the design and implementation of the two core components, the reversible gradient auxiliary branch and the multi-level auxiliary supervision information module.
[0120] The specific steps for implementing the reversible gradient auxiliary branch are as follows:
[0121] S41, given the key feature information in the input data X, it will be gradually diluted or lost during layer-by-layer feature extraction and spatial transformation, which can be expressed as:
[0122] I(X,X)≥I(X,f θ (X))≥I(X,g φ (f θ (X))),
[0123] Where I represents the interaction information, f and g are transformation functions, θ and φ are the parameters of f and g respectively;
[0124] S42, when the transformation function r has an inverse function v, the data can be transmitted without loss of information during the feedforward process:
[0125] X=v ζ (r ψ (X)),
[0126] I(X,X)=I(X,r ψ (X))=I(X,v ζ (r ψ (X))),
[0127] where ψ and ζ are the parameters of the functions r and v, respectively.
[0128] The design of the reversible gradient auxiliary branch significantly improves the reliability and integrity of target feature information during gradient propagation, preventing gradient deviations caused by gradual dilution or loss of information while avoiding the high inference costs of traditional reversible architectures. Furthermore, this branch can drive the backbone network to learn more effective feature representations by supplementing semantic associations.
[0129] In this embodiment, the multi-level auxiliary supervision information module is implemented as follows: inserting multi-level auxiliary supervision information to perform sparse feature filtering and integration on the gradient information of different prediction heads:
[0130] X=v ζ (r ψ (X)·M),
[0131] Where M is a dynamic binary mask used to eliminate noise features and retain key information.
[0132] Multi-level auxiliary supervision information, through a gradient aggregation mechanism, ensures the coordinated optimization of shallow and deep features, significantly enhancing the consistency and integrity of the model's multimodal feature representation. This is particularly effective in complex scenarios such as changing lighting, dynamic objects, and sparse backgrounds, effectively improving the detection performance of multiple types of traffic participants after multimodal fusion. Furthermore, this module is also compatible with detection models similar to the baseline model, further enhancing the model's robustness by bridging the semantic gap between shallow and deep features.
[0133] The multi-level reversible auxiliary supervision mechanism effectively adapts to the needs of multimodal fusion detection in complex urban traffic environments through the main branch, reversible gradient auxiliary branch, and multi-level auxiliary supervision information module. Among them, conventional calculations rely on the main branch and do not introduce any additional inference cost for the model; the reversible gradient auxiliary branch generates reliable gradients by recovering important features lost due to information bottlenecks, thereby enhancing the collaborative expression ability of visible light and infrared image features; the multi-level auxiliary supervision information module alleviates the error accumulation problem in deep supervision by integrating shallow and deep gradients. The design of MRASM effectively improves the accuracy and robustness of the baseline model for multimodal fusion detection in complex urban traffic environments.
[0134] Through the above design, the complementary characteristics of visible light and infrared modalities are fully integrated, and error accumulation is eliminated to ensure the uniformity of global target feature information, thereby adapting to the task of detecting multiple categories of traffic participants in complex urban traffic scenarios.
[0135] The above embodiment is only one of the preferred implementation methods of the present invention and should not be used to limit the scope of protection of the present invention. Any changes or modifications that have no substantive meaning made to the main design concept and spirit of the present invention, as long as the technical problems solved are still consistent with the present invention, should be included in the scope of protection of the present invention.
Claims
1. An adaptive fusion complex traffic environment detection method, characterized by: The detection method is implemented based on the ISMR-YOLO model; the ISMR-YOLO model includes the YOLOv10m model as a baseline model, an infrared detection channel additionally built into the backbone architecture of the baseline model, an interactive feature enhancement module and a feature adaptive weight fusion module respectively integrated into the three levels of the backbone network dual-channel feature extraction, and a multi-level reversible auxiliary supervision mechanism module integrated into the backbone architecture and input architecture of the baseline model; wherein, The infrared detection channel and the existing visible light detection channel together form a dual-channel feature extraction mechanism to support the simultaneous input and processing of infrared and visible light modalities; The interactive feature enhancement module and the feature adaptive weight fusion module work together to complete the interactive enhancement and adaptive weight fusion of multimodal features; The multi-level reversible auxiliary supervision mechanism module uses reversible gradient auxiliary branches and multi-level auxiliary information integration mechanism to build a reliable auxiliary path for global information transmission.
2. The adaptive fusion complex traffic environment detection method according to claim 1, characterized in that: The interactive feature enhancement module realizes interactive enhancement between multimodal features through modal collaborative channel enhancement branch and global perception space feature enhancement branch; wherein, The modal collaborative channel enhancement branch is used to highlight the important parts of the feature information between modalities by assigning channel weights to the target features of each modality, while effectively reducing the impact of redundant information, thereby achieving the purpose of feature complementarity and enhancement. The specific implementation steps are as follows: S11 performs 1×1 convolution operations on the RGB and infrared modal target features of the input traffic status image to reduce the feature dimension and extract channel features: W RGB =σ1(F1(X RGB )), W IR =σ2(F2(X IR )) Among them, X RGB ∈R C×H×W is the RGB feature, X IR ∈R C×H×W are the characteristics of the IR modal target, F1(·) and F 12 (·) is a 1×1 convolution kernel, which is used to compress the channel dimension and extract the global information of the modal features; σ1 and σ2 are Sigmoid activation functions, which are used to normalize the extracted channel features to ensure that the weight range is between [0, 1]; S12, normalize the weights of the two modalities through the Softmax function to obtain the cross-modal channel weights: in, and is the normalized channel weight, which indicates the importance of RGB and IR modalities to their respective channel information. The role of Softmax is to ensure that the sum of the weights between the two modalities is 1, emphasizing the complementary relationship between the information of different modalities. S13, based on the assigned modal weights, weighted optimization is performed on the features of the other modality, thereby achieving feature complementarity between the modalities: Among them, X' RGB and X I ' R Represents the interactively enhanced RGB and IR modal features; represents element-wise weighted operations; S14, the interactively enhanced features are added to their own weights, and both sides of the multimodal model obtain the final modality-enhanced features: A RGB =X' RGB +W RGB , A IR =X I ' R +W IR 。 3. The adaptive fusion complex traffic environment detection method according to claim 2, characterized in that: The global perception spatial feature enhancement branch is used to capture the spatial distribution relationship between modalities and perform interactive enhancement to further enhance the global expression capability of multimodal feature information. The specific implementation steps are as follows: S21, use the global pooling operation to perform global spatial interaction on each modal feature to generate compensation features: Among them, q and v are learnable weight parameters used to balance the contribution between RGB and IR modalities; α4 and α5 are activation functions used to normalize the enhanced features; S22, using the compensated modal features, calculate the interactive enhancement features: in, is the interactive enhanced spatial feature with RGB as the modality; F3(·) is a further 1×1 convolution operation used to refine the enhanced features.
4. The adaptive fusion complex traffic environment detection method according to claim 3 is characterized in that: The feature adaptive weight fusion module is used to achieve efficient fusion of multimodal features. The specific implementation steps are as follows: S31, the feature adaptive weight fusion module uses the interactive feature enhancement module to enhance the IR features after interaction and RGB features As input, let its dimensions be C×H×W again; then the modal weights are extracted again through spatial convolution and the weight matrix is rebuilt: Among them, W IR and W RGB The weight matrices of IR and RGB modal features are established respectively, reflecting the importance of each modality in the spatial dimension; Conv IR and Conv RGB It is a spatial convolution operation used to extract the saliency information of modal features again; S32, generates dynamically adjusted modal weights through the Softmax function: Among them, norm(W IR ,W RGB ) is the modal weight measurement factor, which is used to balance the weights of infrared and RGB modalities; W1 and W2 are the adjusted modal weight matrices, which represent the importance ratios of infrared and RGB modal features, respectively; S33, using the dynamically generated weight matrices W1 and W2, the IR features and RGB features Perform weighted fusion to generate new fusion features: Where: X is the new feature matrix after fusion, which contains the dynamic interaction and enhancement information of the infrared and visible light modalities; The weight W1 is used to highlight the important appearance feature information of the infrared modality in low-light scenes while suppressing background noise. It indicates that the weight W2 is used to highlight the texture and color information of the RGB modality under sufficient lighting conditions.
5. The adaptive fusion complex traffic environment detection method according to claim 4, characterized in that: The multi-level reversible auxiliary supervision mechanism consists of three parts: the main branch, the reversible gradient auxiliary branch and the multi-level auxiliary supervision information module; wherein, The main branch is used to implement the calculation of the reasoning stage; The reversible gradient auxiliary branch is used to solve the information bottleneck problem caused by the deepening of the neural network. It generates reliable gradients for the loss function by maintaining the integrity of information in the feedforward process. The multi-level auxiliary supervision information module targets the deep supervision module by introducing multi-level gradient integration in the architecture of multiple prediction branches to alleviate the problem of insufficient coordination between shallow features and deep features.
6. The adaptive fusion complex traffic environment detection method according to claim 5, characterized in that: The specific steps of implementing the reversible gradient auxiliary branch are as follows: S41, given the key feature information in the input data X, it will be gradually diluted or lost during layer-by-layer feature extraction and spatial transformation, which can be expressed as: I(X,X)≥I(X,f θ (X))≥I(X,g φ (f θ (X))), Where I represents the interaction information, f and g are transformation functions, θ and φ are the parameters of f and g respectively; S42, when the transformation function r has an inverse function v, the data can be transmitted without loss of information during the feedforward process: X=u ζ (r ψ (X)), I(X,X)=I(X,r ψ (X))=I(X,v ζ (r ψ (X))), where ψ and ζ are the parameters of the functions r and v, respectively.
7. The adaptive fusion complex traffic environment detection method according to claim 6, characterized in that: The multi-level auxiliary supervision information module is implemented as follows: inserting multi-level auxiliary supervision information to perform sparse feature filtering and integration on the gradient information of different prediction heads: X=v ζ (r ψ (X)·M), Where M is a dynamic binary mask used to eliminate noise features and retain key information.
Citation Information
Cited By
Target detection method, system, device and medium
CN121170266A
Urban night traffic monitoring method based on multi-modal fusion and deep learning
CN122049830A