A lightweight heterogeneous collaborative target detection method for fog driving conditions
Patent Information
- Application Number
- CN202512053873.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-12-31
AI Technical Summary
然而,这些方法均存在明显局限性
1)在雾、逆光等复杂气象条件下,本发明通过雾化增强、几何校正与自适应感受野协同,稳定输出高质量检测结果,显著降低小目标漏检与误检率。
Smart Images

Figure CN121884305B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent driving technology, specifically to a lightweight heterogeneous cooperative target detection method for driving in foggy conditions. Background Technology
[0002] In the current era of rapid evolution in intelligent transportation, perception capabilities for low visibility and complex weather environments are becoming a crucial prerequisite for the large-scale deployment of autonomous driving systems. Adverse environments such as fog, rain, backlighting, and nighttime conditions can drastically reduce image contrast and obscure signals from distant, small targets, thus amplifying the detector's sensitivity to domain shifts, illumination changes, and geometric distortions. Simultaneously, real-world road scenarios exhibit highly uneven target scale distribution, frequent occlusion, and a high proportion of long-tailed categories, making conventional models prone to decreased recall and increased false positives in cross-scale fusion and boundary detail preservation. Furthermore, the limited computing power and power consumption at edge devices force perception algorithms to achieve a new balance between real-time performance, accuracy, and model size. Existing fog-related target detection technologies mainly follow three approaches: image preprocessing-based detection methods, YOLO series models based on CNNs, and end-to-end detection models based on Transformers. However, all these methods have significant limitations. While preprocessing methods such as dehazing algorithms can improve image quality, they suffer from severe computational latency and are prone to feature distortion, failing to meet real-time detection requirements. YOLO-based detection models perform poorly in foggy conditions. Their backbone networks are sensitive to feature degradation in fog, and their feature pyramids struggle to effectively fuse multi-scale features. Furthermore, overly lightweight detection heads weaken the model's representational capabilities. While Transformer models possess powerful feature extraction capabilities, their high computational cost and long training cycles make them difficult to deploy in resource-constrained automotive systems.
[0003] In summary, the existing technologies have the following shortcomings: First, feature extraction is insufficient. Traditional convolution is sensitive to fog noise and feature degradation, making it difficult to effectively extract key features. Second, feature fusion efficiency is low. Existing feature pyramid networks struggle to accurately assess the contribution of each feature under fog conditions, resulting in poor fusion performance. Third, detection head redundancy exists. Although the YOLO series detection heads are becoming increasingly lightweight, excessive simplification weakens the model's ability to discriminate features in foggy scenes. Fourth, preprocessing dependency exists. Existing methods often employ a defogging process followed by detection, introducing additional computational burden and failing to meet real-time requirements. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a lightweight heterogeneous cooperative target detection method for driving in foggy conditions, comprising the following steps: Step S1: Collect road images in foggy weather using a vehicle-mounted dual-mode camera and an adjustable LED fill light device. Perform adaptive exposure control and image preprocessing based on an atmospheric scattering model to obtain a standardized input image. Step S2: Input the standardized input image into the parallel multi-scale feature extraction network, and extract multi-level feature maps through the context-guided downsampling module and the parallel grouped multi-scale convolution module; Step S3: Input the multi-level feature map into the bidirectional weighted cross-scale semantic fusion network, and perform cross-scale feature fusion through the fog perception upsampling module and the fog adaptive feature fusion module to obtain the fused feature map; Step S4: Input the fused feature map into the lightweight grouping decoupled detection head, output the target category confidence and bounding box parameters through the classification branch and regression branch respectively, and output the detection result after non-maximum suppression.
[0005] Preferably, the adaptive exposure control in step S1 includes: Atmospheric light based on dark channel statistical estimation of preview frames The maximum response in the highlighted area is used to determine this. Contrast-robust global transmittance index Characterizes the level of visibility; Target grayscale mean With saturation upper limit To constrain this, we will jointly optimize the exposure time. With LED intensity ; When the global transmittance index Below the threshold At that time, prioritize increasing LED intensity. To avoid motion blur.
[0006] Preferably, step S1 further includes controllable atomization enhancement, specifically: With clear, fog-free images With estimated depth As input, the scattering coefficient after perturbation With atmospheric light Synthetic enhanced fog image : ; in, , , and These are the random disturbance quantities uniformly sampled within a preset interval. The baseline scattering coefficients are estimated based on real fog images. It is a baseline atmospheric luminosity estimated based on real fog images; Using soft constraint regularization terms Limiting the enhancement strength, the soft constraint regularization term It includes the weighted sum of the contrast measure and the total variation measure.
[0007] Preferably, step S1 further includes dual exposure fusion, specifically: Acquire short-exposure images in the neighborhood at the same time. With long exposure images ; Calculate the short exposure images respectively With long exposure images Pixel-level fusion weights and : ; ; The input brightness is generated by weighted fusion to extend the dynamic range. : ; In the formula, This represents the desired target gray level; This represents the weight decay coefficient that controls the weights around the target gray level. Indicates long exposure time The image captured below is in pixels Brightness at that location; Indicates short exposure time The image captured below is in pixels Brightness at that location; It is the allowed oversaturation threshold for a pixel. It is the allowed undersaturation threshold for a pixel; This is an indicator function that takes the value 1 when the condition within the parentheses is met. It is a very small constant greater than 0.
[0008] Preferably, step S1 further includes visible light and near-infrared dual-mode guided fusion, specifically: Output visible light channel With near-infrared channel ; Based on transmittance estimation Calculate the fusion weights using the Sigmoid function. : ; Output combined luminance : ; In the formula, Represents the Sigmoid function; It is the slope coefficient of the Sigmoid function with respect to transmittance; It is the transmittance threshold.
[0009] Preferably, the context-guided downsampling module in step S2 includes: The input feature map is downsampled using a convolution with a stride of 2; Complementary contextual features are extracted by parallel convolutional branches with local depthwise convolutional branches and dilated depthwise convolutional branches. The outputs of the two branches are concatenated along the channel dimension and then processed by batch normalization and activation function. The number of channels is compressed using 1×1 convolution, and the channels are recalibrated using a global channel attention unit.
[0010] Preferably, the parallel grouped multi-scale convolution module in step S2 includes: The input feature map is divided into two parallel paths after being convolved by a 1×1 convolution. The first path directly proceeds to the splicing operation, while the second path is processed through N GMC units; The GMC unit is sequentially subjected to 3×3 standard convolution, 5×5 grouped convolution and 7×7 grouped convolution, and channel segmentation is performed after each level of convolution; the features of three different scales are concatenated in the channel dimension and channel reshaping is performed by 1×1 convolution. Finally, the outputs of the two paths are concatenated using a jump connection.
[0011] Preferably, the fog sensing upsampling module in step S3 includes: Get the Layer input feature map Corresponding transmittance diagram ; Based on local features and fog intensity with structural anisotropy coefficient Learnable recombination weights within the neighborhood are calculated using a weight generator. ; according to Perform upsampling; Wherein, the fog intensity Depend on The anisotropy coefficients of the structure are obtained by logarithmic clipping mapping. Defined by the ratio of principal eigenvalues of the structure tensor.
[0012] Preferably, step S3 further includes a fog attenuation-transmission joint alignment and fusion strategy, specifically: For pyramid layer sets Features of each layer At pixel position Define the intrinsic visibility weight. The intrinsic visibility weight Determined by a weighted combination of transmittance and gradient magnitude; Solving from each layer to the anchor layer Alignment operator Minimize the sum of the weighted alignment error and the regularization term; Normalized fusion weights for each layer are calculated based on visibility weights and alignment similarity. And generate fused feature maps : ; ; In the formula, Indicates the level of scoring; This is a pixel-by-pixel similarity measurement function; Characteristic diagram representing the anchor layer; This indicates element-wise multiplication; This is the global average pooling operator.
[0013] Preferably, step S4 further includes a risk assessment, specifically: Estimate the relative distance for each detected target. With relative velocity ; Calculate collision time ,in To prevent positive constants from being divided by zero; Calculate the overall risk score The comprehensive risk score Based on the collision time TTC and the spatial relationship between the target and the road passable area The weighted combination of the confidence scores is obtained by normalization using the Sigmoid function; When the comprehensive risk score A collision warning is output when the threshold is exceeded.
[0014] Compared with the prior art, the beneficial effects of the present invention include at least the following: 1) Under complex weather conditions such as fog and backlight, this invention achieves stable output of high-quality detection results through the synergy of fogging enhancement, geometric correction and adaptive receptive field, significantly reducing the false detection rate and missed detection rate of small targets.
[0015] 2) This invention ensures real-time operation and low power consumption through a lightweight heterogeneous design and decoupled detection head, which facilitates rapid deployment and large-scale operation and maintenance of edge devices on vehicles or roadsides, thereby improving the overall safety and reliability of the system.
[0016] 3) Significantly improves the accuracy and recall rate of fog perception without relying on defogging before detection, and achieves vehicle-grade real-time performance with extremely low parameter and computational requirements, reducing hardware costs and energy consumption.
[0017] 4) Dynamic receptive field and grouped convolution enhance the robustness of small target detection, reduce false positives and false negatives, and improve the reliability and interpretability of advanced driver assistance system warnings. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall process of an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0020] To avoid redundant interpretations of symbols, this specification adopts the following unified notation conventions as shown in Table 1. The complete definition will not be repeated when a related symbol is used for the first time thereafter.
[0021] Table 1 like Figure 1 As shown, this embodiment of the invention provides a lightweight heterogeneous cooperative target detection method for driving in foggy conditions, including the following steps: Step S1: Acquire road images under foggy conditions In this embodiment of the invention, road images are continuously captured under specific lighting conditions using a data acquisition and fogging enhancement module with a vehicle-mounted forward-looking camera and adjustable LED supplementary lighting.
[0022] 1.1 Image Acquisition Module The image acquisition module is designed for foggy and low-visibility road scenarios. With a hardware framework of visible light and near-infrared dual-mode forward-looking camera and adjustable LED supplementary lighting, it acquires road image sequences with unified geometric and photometric specifications. Without increasing the latency of the inference end, it outputs physically consistent, high-quality input that can be used for subsequent feature extraction and cross-scale fusion, as well as metadata related to the imaging process.
[0023] Module in system clock Below at a fixed frame rate With exposure duration Acquire raw signals, The frame number (sampling time index) is used to represent the acquisition time of a single frame. Simultaneously emit an intensity of [value] in each frame. Controllable LED pulses are used to reduce the signal-to-noise ratio degradation when fog-induced backscattering dominates, and the line delay introduced by the rolling shutter is timestamped, setting the effective exposure midpoint time of the j-th line to be... ,in The time series is used as a temporal constraint in subsequent geometric correction to eliminate geometric distortions caused by inter-line motion and ensure cross-frame pose alignment, thereby forming a stable spatially aligned input before entering the geometry and receptive field adaptive feature extraction module.
[0024] 1.2 Atmospheric Scattering Imaging Model The imaging model employs a combination of atmospheric scattering assumptions and camera radiometry. For any pixel location... With band The radiance is The entrance pupil light flux is related to the optical transfer function and the sensor quantum efficiency. After accumulating points, then in the exposure Internally generated charge: ; in It is a pixel During the exposure time The amount of charge accumulated internally, This represents the combination of readout noise and photon noise. This represents the variance of the noise term; Fog scattering employs a physically consistent atmospheric scattering model. The intensity of the atomized radiation reaching the camera is: ; ; In the formula, the first term on the right-hand side represents direct light, and the second term represents atmospheric light caused by fog. The camera reads the following... .
[0025] 1.3 Adaptive Exposure and Fill Light Control To improve visibility and dynamic range utilization under low-contrast conditions, the module estimates fog parameters with extremely low overhead and solves the joint settings of exposure and fill light online before acquisition.
[0026] Based on preview frames By combining the statistics of the dark channel with the maximum response in the bright space region, an atmospheric light estimation is constructed: ; in These are the highlighted areas extracted from the preview frame, typically corresponding to the sky, distant fog background, etc., used to estimate atmospheric light. To show the pixel positions within the highlighted area set of the preview frame, Represents pixels Normalized brightness at that location In order to be in The operator is used to find the pixel with the highest brightness.
[0027] Visibility levels are approximated by a contrast-robust global transmittance index: ; in This refers to the minimum neighborhood statistics after minimum channel or guided filtering. It is a simple physical mapping of the dark channel. This represents the median value of the pixel-level transmittance across the entire image, providing a more robust visibility characterization against noise and local anomalies.
[0028] Let the average gray level of the target be And saturation upper limit Therefore, the joint selection of exposure and LED intensity can be transformed into constrained optimization: ; in Indicates the contrast utilization index; For output image The variance; Describe the expected value of the grayscale of the output image; The desired average brightness; It is the maximum number of pixels allowed by the camera. Representing constraints This applies to almost all pixels in the image.
[0029] When solving online linear approximation Linearizing with the small noise assumption, we get: ; in It is the camera response. First-order linear approximation coefficients under low noise conditions; It is the equivalent gain function that combines system gain, normalization factor, etc. The equivalent constant noise term under the small noise approximation. The desired conditions constrained under this approximation. Given a closed form: ; in The optimal exposure time is obtained under the condition of considering only brightness constraints and approximating a linear camera response. It is the expectation for the entire image, used to represent the average incident brightness per unit exposure under given fog conditions, and is expressed as the boost of the near-infrared channel by the LED. Correcting the signal-to-noise ratio in distant dark areas, where The signal boost provided by near-infrared LED supplemental lighting This is the equivalent gain coefficient of the light source and sensor system.
[0030] United Land Limited to Inside, among which and These are the minimum and maximum exposure times allowed by the system, used to avoid underexposure and motion blur; LED intensity is limited to... Inside, This is the maximum intensity that an LED can output. The calculated global transmittance index... Below the threshold Then prioritize increasing Strategies to improve visibility to avoid The motion blur caused by excessive size can be formalized as follows: ; in This is the proportionality coefficient. It is the optimal LED intensity control value estimated based on the current visibility. It is a positive part operator, guaranteeing that when The LED intensity will not be increased at that time.
[0031] 1.4 Controllable Atomization Enhancement (Training Sample Generation) To accommodate the domain diversity and physical consistency of subsequent networks, a controllable fog enhancement channel is integrated on the acquisition side to expand the training distribution and generate samples consistent with real fog-induced degradation. Its core is based on the identified... Random perturbations are performed around the center within a safe range.
[0032] make This represents the brightness at pixel x of a clear road image acquired and normalized under fog-free conditions (i.e., the aforementioned fog-free scene radiance). (digital image form) and estimated depth Enhancement pairs are constructed from sparse lasers or monocular depth priors: ; in , , , Used to control the intensity of disturbances; The baseline scattering coefficients are estimated based on real fog images; It is a baseline atmospheric luminosity estimated based on real fog images; It is a random perturbation of the scattering coefficient; It is a random disturbance to atmospheric light; This indicates sampling with equal probability over a given interval; It is the upper bound of the scattering coefficient perturbation amplitude, which controls the intensity of fog concentration enhancement; It is the upper limit of the atmospheric light disturbance amplitude, which controls the intensity of background brightness enhancement; This indicates the enhanced synthetic fog image at the pixel level. The brightness at that location is used for training data augmentation; In pixels Estimated scene depth at the location; Used to measure atmospheric brightness after disturbance in the synthesis of fog.
[0033] To prevent over-enhancement from causing distribution drift, soft constraint regularization is introduced: ; in For contrast measurement; For total variation; As weight; Digital brightness of the raw foggy image directly read from the image sensor before fog enhancement; It is a soft constraint regularization term designed to enhance strength and suppress distribution drift. It is minimized during training. To limit the intensity of enhancement while maintaining mission consistency, among which This represents task loss, such as the joint loss of classification and regression in a detection network, used to ensure detection performance.
[0034] In the formula Indicates the entire image The calculated RMS contrast ratio (scalar) is specifically implemented using the following formula: ; in It can be any input image. It is the set of all pixel locations in the image. It is the total number of pixels. It sums the values at all pixel positions x within the image domain. It is an image In pixels The brightness value at that location, It is the mean. It is an image The RMS contrast ratio is represented by the square root of the variance, which characterizes the overall contrast level.
[0035] 1.5 Double Exposure Fusion (for backlighting and glare scenes) To address localized saturation caused by backlighting and glare, the module employs a seamless strategy of dual-exposure micro-shift fusion, in addition to exposure selection. (Setting the same...) Two images with different exposures were captured within the same neighborhood. Image and This is then used to generate a high dynamic range extended range input brightness through weighted fusion. The weights employ a combined approach of saturation suppression and noise suppression: ; ; ; in Control the weight decay around the target gray level; For indicator functions; Indicates short exposure time The image captured below is in pixels The brightness of the area is adjusted to preserve details in the highlighted areas; Indicates long exposure time The image captured below is in pixels The brightness of the area is adjusted to improve the signal-to-noise ratio in the darker areas; and Corresponding to short exposure images With long exposure images Pixel-level fusion weights; This represents the desired target gray level; This is the allowed oversaturation threshold for a pixel; pixels exceeding this value are considered overexposed. This is the permissible undersaturation threshold for a pixel; pixels below this value are considered severely underexposed. and It is an indicator function that takes the value 1 when the condition inside the parentheses is met, and 0 otherwise; It is the dynamic range extended input brightness obtained after weighted fusion of dual exposures, which is used as the input for subsequent modules; It is a very small constant greater than 0, used to prevent the denominator from being... and This results in division by zero during local degradation.
[0036] To avoid mismatches caused by motion, timestamps are used before merging. With optical flow right Perform subpixel alignment. The assumption can be maintained through brightness. The least squares approximation is obtained, where j is the image row index. It is the spatial gradient of image brightness. It is the derivative of image brightness with respect to time.
[0037] 1.6 Visible and Near-Infrared Dual-Mode Guided Fusion To enhance fog penetration and provide heterogeneous information available to the backend without significantly increasing bandwidth, the module outputs a visible light channel. With near-infrared channel Meanwhile, guided fusion based on transmittance estimation from the acquisition side is also provided: ; ; in It is the Sigmoid function that affects transmittance. The slope coefficient; It is the transmittance threshold; In pixels The visible light channel output brightness at that location; In pixels The output brightness of the near-infrared channel at that location; It is the combined output brightness of the two channels after guided fusion on the acquisition side, which is then used by the backend detection and recognition module. The fusion weighting coefficients are obtained based on local transmittance estimation. A larger value indicates that the pixel relies more on the visible light channel, while a smaller value indicates that it relies more on the near-infrared channel.
[0038] Transmittance estimation The initial value of the dark channel is obtained by smoothing it with a fast steering filter. Dark channel intensity estimation is achieved through... The calculations yielded the results. This fusion provides a more robust observation of fog-affected areas at the perception layer while maintaining realistic textures in clear skies and near-field areas, ensuring consistency with the input task. The resulting data, after normalization and adaptive brightness correction, is fed into the geometry and receptive field adaptive feature extraction module.
[0039] Step S2: Perform feature extraction using a parallel multi-scale feature extraction network. In this embodiment, a geometry and receptive field adaptive feature extraction module is used to compensate for the loss of detail and scale changes caused by imaging conditions; a GMCC-Net backbone composed of parallel multi-scale convolution and depthwise separable convolution is used, and dynamic kernel generation and guided context recalibration are integrated to automatically allocate appropriate equivalent receptive fields to different regions, thereby obtaining standardized and information-rich feature representations.
[0040] 2.1 ContextGuidedBlockDown (CGD) Module First, a context-guided module for downsampling is used for downsampling. During downsampling, local and surrounding contexts are fused and global gating is used to adaptively retain the most useful channels, so as to reduce the resolution while minimizing the loss of effective information.
[0041] Given the input feature map ,in Represents batch size. Represents the number of input channels. and It is the spatial size of the input feature map, first passed through the convolution kernel. stride void ratio Downsampling is performed on the convolution: ; in It refers to the convolution kernel weights and their overall shape. , This is the number of input channels. These are the local indices of the convolution kernel in the height and width directions. Indicates the first The bias of each output channel, and , The output feature map is at the 1st The spatial location under each channel The value at that location, and These are the coordinate indices of the output feature map in the height and width directions. It is the kernel radius, used to shift local indices to a position equal to or less than the kernel radius. The neighborhood centered on, express The feature value at position , in the input feature map, is the th Each channel and spatial location has three possible values, ranging from 0 to 2.
[0042] The output feature map is: ; , ; Then through local branches and surrounding branches The modeling process separately addresses the immediate context and the broader context. These two branches employ parallel methods—local depthwise convolution and dilated depthwise convolution—to extract complementary context. ; ; in and This indicates the local branch and surrounding branches in the output channel. Spatial coordinates The output pixel value; and This represents the weights of the depthwise convolutional kernels for the local branch and surrounding branches, and , ; and The corresponding branch is in the channel Bias terms on; This indicates that the input feature map is in the channel Spatial coordinates The pixel value at that location.
[0043] Then, concatenation, normalization, and activation functions are used to concatenate and fuse the output features: ; ; in Indicating in channel dimension The above is a splicing operation, so that the number of channels increases from Become ; This represents the output features of BN and activation; and It is based on the mean and variance of the channels; and This represents the learnable scaling and translation parameters of BN; It is a numerically stable term to prevent the denominator from being zero or too small; ReLU represents the element-wise activation function. This indicates that in the channel dimension and The concatenated feature tensor.
[0044] Then use a 1×1 convolution to reduce the number of channels from 2 Compress to The formula is shown below: ; in This indicates that after a 1×1 convolution, the feature map shows the... Each channel at pixel position The value; These are the weights of the convolution kernel; The input feature map is at the th Each channel pixel position The value at; It is the first The bias of each output channel.
[0045] Then the number of channels is... Feature maps are input to the global channel attention unit , The average pooling layer is used to extract global statistics for each channel, and then a two-layer MLP is used to generate channel weights. These weights are used to scale the fused features channel by channel, thereby highlighting useful contextual information and suppressing redundant channels. ; ; ; ; in This represents the feature map obtained after 1×1 convolution. In the Each channel, spatial coordinates The pixel value at that location can be obtained by performing global average pooling on each channel across the entire spatial location. The value of , where and These represent the height and width of the output feature map, respectively. For the first A global feature scalar for each channel. (This refers to the global feature scalar of all channels.) Composition vector Input a two-layer MLP: the first layer obtains the intermediate vector after channel compression. The second layer yields the channel weight vector. .vector The Each component These are the weighting coefficients for the corresponding channels, used to... Channel-by-channel scaling is performed to obtain the recalibrated output features. .
[0046] These are the characteristic responses obtained through an MLP network, used to control the generation and reuse of channels. and These are the weights and biases of the first fully connected layer in an MLP. and These are the weights and biases of the second fully connected layer in an MLP. It is the ReLU activation function. It is a channel-gated weight vector, with one weight for each channel. , Then it means Output characteristics after recalibration.
[0047] 2.2 Parallel Grouped Multi-Scale Convolutional Module (CSP-Par-GMC) The feature map output after downsampling by the CGD module is then input into the CSP-Par-GMC module. Let the given input feature map be... Where B represents the batch size and C represents the number of channels. First, the input information undergoes 1×1 convolution and split operations in the CSP-Par-GMC module, forming two parallel paths. One path retains the original information and directly enters the concat operation, while the other path is processed through N GMC units.
[0048] In the GMC unit, the input feature map first undergoes a standard 3×3 convolution. This step extracts basic spatial features from the input features while maintaining the number of channels, providing a basic feature representation for subsequent multi-scale analysis. Then... Equal division along the channel dimension ,in It continues to be passed to the next level for more granular analysis. Retain the feature information at the current level, and wait for final fusion. After 5×5 grouped convolution, by using a larger convolution kernel (5×5) to increase the receptive field, we obtain... .
[0049] Grouped convolutions reduce the number of parameters, with each group processing only a single channel, enabling independent spatial analysis between channels. Again, regarding... Channel segmentation is performed to obtain ,in Delivered to the most granular level of analysis. Preserve medium-scale feature information. Then perform a third level of fine feature extraction. After 7×7 grouped convolution, we obtain Finally, the features at three different scales are concatenated at the channel level: ; in This is a fused feature map, representing the result of concatenating feature information from different levels. This fusion method allows each level to contribute information with varying levels of abstraction and receptive field. The fused features... The 1×1 convolution enables information exchange and reorganization between channels, ensuring direct gradient propagation, alleviating the gradient vanishing problem in deep network training, and preserving the original input information to prevent feature degradation.
[0050] ; in The final feature map output is obtained by performing a 1×1 convolution on the fused features. Indicates use Convolution kernel on input feature map Performing convolution operations can effectively compress channels and reconstruct features. This is the original input feature map, used for skip connections to avoid information loss in deep networks. Skip connections help preserve original information and prevent gradient vanishing or excessive feature degradation.
[0051] In terms of stage arrangement, GMCC-Net adopts a strategy of progressive downsampling, multi-scale parallel enhancement, and lightweight replacement of high-level layers: the first two levels use CGD to complete context-constrained downsampling, maximizing the preservation of edges and contours weakened in foggy conditions; the middle layers stack CSP-Par-GMC modules to expand feature diversity and complete cross-scale integration; the high-level layers use DSConv to replace conventional convolution, significantly reducing the amount of computation and parameters without losing accuracy, so that the backbone remains real-time friendly to foggy conditions.
[0052] Step S3: Perform feature fusion using a bidirectional weighted cross-scale semantic fusion network. To address the chain reaction of input degradation leading to difficulties in cross-scale alignment and distortion of small targets under low visibility conditions in foggy weather, a fog perception optimization system is constructed that integrates all these elements. First, an adaptive mechanism based on physical transmittance is designed, transforming resampling and dynamic convolution from simple module stacking into interpretable fog-driven operators F-DSCB and FA-C3k2-DKMB. Then, a novel fog attenuation-transmission joint alignment fusion strategy replaces the conventional pyramid paradigm, simultaneously constraining information flow from both target visibility and cross-layer consistency perspectives. Finally, fog-specific training and inference algorithms are employed to deeply couple the atmospheric scattering model with the detection learning target, forming a physical-data dual closed loop from input to detection output.
[0053] 3.1 Fog-sensing learnable upsampling module (F-DSCB) First, a fog-aware DSCB (F-DSCB) is constructed at the learnable upsampling end. Let the... The layer input feature map is Corresponding pixel position The fog field transmittance diagram is as follows The transmittance is obtained by rapid estimation based on the atmospheric scattering model on the acquisition side and smoothed by guided filtering to ensure physical consistency and real-time performance.
[0054] Upsampling to The recombination weights are not directly generated by the texture, but rather by... Local features within the neighborhood and Joint decision: ; in, These are the weight coefficients for local learning, which define the contribution of different pixels to feature fusion within a local region; Indicated by Centered, preset window The discrete offset vector within the neighborhood, the softmax function for all offsets within that neighborhood The weights on the scale are normalized so that their summation in the neighborhood is 1; It is the upper limit threshold of fog intensity; It is a measure of fog intensity obtained by transmittance mapping; In the feature map above A local neighborhood (supporting region) centered on a pre-defined size. The structural anisotropy coefficient, defined by the principal eigenvalue ratio of the structural tensor, is used to identify striped edge structures and suppress the blur diffusion in fog-induced contrast collapse regions. It is a weight generator composed of lightweight convolutional branches and gated MLP, which combines local patch features and fog intensity. With structural anisotropy The encoding is a set of unnormalized weights.
[0055] The final upsampling result is: ; in The F-DSCB module is in position The upsampled output features, i.e. the reconstructed features of fog perception, are obtained by recombining the features in the neighborhood according to weights. It is the original feature map In position The eigenvector at that location.
[0056] Apply fog-sensing channel swapping in the channel dimension Its number of groups Depend on Adaptive scaling enhances cross-group information flow in low-visibility scenarios to compensate for high-frequency loss. This design utilizes... and A fog-driven weight redistribution is achieved, which significantly improves the boundary sharpness and cross-scale consistency after upsampling in fog and small target scenes.
[0057] 3.2 Fog Adaptive Feature Fusion Module (FA-C3k2-DKMB) We employed the FA-C3k2-DKMB fog adaptive feature fusion module, specifically designed for foggy conditions, and devised a novel fog-sensing dynamic kernel-weighted FA-DKMB that combines transmittance and orientation selection. Let the bottleneck input feature map be... The three depthwise convolution branches each have a square kernel. transverse band-shaped nucleus With longitudinal band nucleus The corresponding output is: ; ; ; in This represents channel-wise depthwise convolution; Core size; It is the output feature map after depthwise convolution; It is the input feature map.
[0058] To characterize the different degrees of erosion of structures by fog in all directions, a branch weight pair is constructed. : ; in yes The global mean; The variance of the gradients in the x and y directions; Taken from the ratio of eigenvalues of the structure tensor; It's local entropy. Branch probability. The convexity is satisfied, and the output is .
[0059] when rise and When significant (large), the model tends to enhance the strip-shaped kernel branches to restore the directionally consistent structure preferentially smoothed by the fog; when Gao Er When the value is moderate, the square kernel is increased to stabilize local texture details. The weight generation of this mechanism is explicitly bound to fog field statistics and structural anisotropy, so that the adaptive receptive field is reorganized and applied to physically interpretable weight transfer. It is used alternately with each fusion node in the bidirectional path to ensure that the scale reconfiguration of fog perception is completed before entering the next round of fusion.
[0060] 3.3 Fog Attenuation-Transmission Joint Alignment Fusion Strategy (AT-Fusion) To fundamentally solve the problems of cross-layer misalignment and contribution assessment distortion, a strategy of fog attenuation and transport co-aligned fusion (AT-Fusion) is proposed. Its goal is to rewrite the objective function of cross-scale information convergence with visibility consistency as a priori.
[0061] Suppose a pyramid layer set For each feature map in the pyramid, each layer of features At pixel position The intrinsic visibility weight is defined as follows: That is, the first Layer in position Pixel-level weights on, where: ; Here, is the visibility weighting function used to balance the contributions of transmittance and gradient magnitude, where These are indicators obtained from transmittance or fog intensity. The weighting used to amplify the position when transmittance is low (heavy fog) For the gradient magnitude at the corresponding position, Characterizes local edge strength. and To adjust the non-negative weighting coefficients of the two relative effects, An exponential hyperparameter for controlling sensitivity to low transmittance regions.
[0062] AT-Fusion no longer simply adds the uplink and downlink paths, but aligns them before merging: ; in, For the set of layers of the feature pyramid, Indicates the first Layer feature map, The anchor layer feature map selected based on the visibility criterion; For the image domain; For the first Layer at pixel position The inherent visibility weight of the location; Indicates from the first Layer to anchor layer The alignment-resampling operator acts on the feature map. The result above is denoted as For example, it can be caused by a small deformation field. Defined as The regularization term, which imposes smoothing and magnitude constraints on all alignment operator parameters, can be, for example, taken as a function of the gradient and norm of each deformation field. The sum of the integrals; This is a hyperparameter used to adjust the relative weights of data items and regularization items.
[0063] To achieve cross-layer convergence that involves alignment followed by fusion in both uplink and downlink paths, this invention obtains the alignment operator... Then, the contribution weights of each layer are determined jointly based on visibility and similarity, specifically: for the pyramid layer set In the anchor layer The fused feature map at the location is defined as: ; in To make the first The result after aligning the layer feature map to the anchor layer coordinate system. For the first Layer normalized fusion weights. Layer scoring. Determined by both visibility weight and alignment similarity: ; For the aforementioned The intrinsic visibility weight map of the layer at each pixel location. For pixel-wise similarity measurement functions (such as normalized inner product or correlation coefficient). This indicates element-wise multiplication. This is a global average pooling operator that averages the weight-similarity map across the entire image in the spatial dimension, thus obtaining a scalar. This is considered as the overall contribution of this layer. Through the aforementioned softmax normalization, the weighted levels contribute to the final fused features. The impact is greater.
[0064] This fusion strategy, with visibility consistency at its core, avoids the erroneous amplification of noise contributions from severely degraded layers in foggy conditions, and is compatible with the existing bidirectional path interface of EBLFPN, resulting in low engineering replacement costs.
[0065] 3.4 Training and Inference Algorithms Specific to Foggy Weather Based on the above structure, a fog-specific training and inference algorithm is designed to ensure that the physical parameters and detection loss are consistent constraints.
[0066] (1) Transmittance-Guided Feature Compensation (TG-FC): For the first... Layer features in window Internal computation Define the compensation coefficient Output ,in and These respectively indicate the use of on the current channel. or The mean and standard deviation calculated in this way are used to analyze the feature map. Normalization For learnable parameters, Preventing division by zero. This step improves the effective dynamic range and gradient resolution in low-visibility areas without explicit dehazing.
[0067] (2) Visibility-Weighted Detection Loss (VW-DL): For the classification branch, the following is applied: ; Focal Loss For classification loss, represents the loss term used to calculate class prediction in object detection tasks. The total number of samples, The visibility weight is used to weight the nth sample in the detection loss. It's Focal Loss, used to solve the class imbalance problem. To predict probabilities, For real labels, This is an adjustment parameter in Focal Loss, used to control the ratio between positive and negative samples; it corrects the discrete probability terms in the Distributed Bounding Fold (DFL) representation of the regression branch with equal weights. ; in The target probability of the nth sample at the kth offset. It is the probability predicted by the network at the k-th offset, and is included in the IoU or EIoU loss. Modulate difficult samples; thus, during the training period, explicitly focus on low-visibility samples without disrupting the overall balance of negative and positive samples.
[0068] (3) Haze-Consistency Regularizer (HCR): Constrains cross-channel consistency induced by the atmospheric scattering model: ; in For the visible domain approximation of network reconstruction; Estimate the shallow reflectance components; It is the loss value of the fog field consistency regularization term; It is the value of the visible domain image reconstructed by the network at pixel x, which can be regarded as an approximation of the foggy (or observed image) generated by the network, enhancing the consistency between the shallow layer and physical observation; It is a general number; The atmospheric light intensity parameters in the imaging model are directly estimated online from the preview frames by the acquisition side. It maintains consistency with the front-end physical imaging model, thereby ensuring an end-to-end closed loop for the entire system.
[0069] During the inference phase, the lightweight IoU-NMS and EGHead decoupled detection head outputs structured results, while retaining the TTC and risk score calculation interfaces to facilitate triggering warnings and decisions on the ADAS side. It is important to emphasize that the fog perception modifications in the feature domain of F-DSCB and FA-DKMB, along with the visibility consistency principle of AT-Fusion, maintain semantic and implementation consistency with existing system processes. Therefore, they can be directly embedded into cross-scale fusion, object detection inference, and risk assessment stages without altering the external interfaces of the upstream acquisition and geometric correction modules.
[0070] Step S4: Perform target classification and localization using a lightweight grouping decoupling detection head. For feature fusion, a lightweight grouped decoupled detection head (EGHead) is used to complete multi-scale detection. Stable training is achieved by decoupling the classification and regression branches and reducing the cost of group convolution. The bounding box regression adopts a distributed representation, and the post-processing uses Intersection over Union (IoU) and Non-maximum Suppression (NMS) to output structured results, and combines them with business requirements to generate risk assessment indicators, such as collision warnings.
[0071] 4.1 EGHead Structure The structure of EGHead for each layer of features F : ; in , For group convolution stacking, It represents the discrete distribution of each edge, where F represents the input feature map from a certain layer of the feature pyramid. This represents the probability of a target existing at each location in this layer, determined by the classification branch. It is obtained by activating Sigmoid.
[0072] 4.2 Distributed Boundary Regression Distributed bounding box regression discretizes each boundary offset across K bins and learns their probability distribution. The superscript K indicates that the length of this vector is K, meaning it has K components. During inference, the expected value is used as the continuous offset: ; ; and with distribution focal loss Supervision soft tags, It is the number of bins into which the boundary offset is discretized; It is a length of The probability distribution vector, where each dimension corresponds to the probability of a bin. , It is a vector The Middle The component indicates that the offset falls within the first component. The probability of each bin; To obtain the desired continuous boundary offset by applying a discrete distribution; then through... The probability distribution of the boundary offset in each bin is obtained.
[0073] 4.3 Risk Assessment The risk assessment estimates the relative distance for each target. relative speed And calculate TTC: ; And a comprehensive risk score normalized to [0,1] is given: ; in This indicates the spatial relationship with the road's passable area / lane lines; the score represents the detection confidence level. It is an estimate of the relative distance to the target; It is an estimate of the target's relative speed to the vehicle; It is Time-To-Contact, which represents the estimated time required for a collision with the target to occur while maintaining the current relative velocity at a constant level. Take from the denominator and The larger value, It is a very small positive constant used to avoid division by zero when the speed is very small or zero. The weights are adjustable based on business requirements. Upon triggering a threshold, it outputs metrics such as collision warnings / braking suggestions.
[0074] When the method of this invention is used, the overall workflow is as follows.
[0075] The overall workflow of this invention embodiment is as follows: Step 1: Image Acquisition – Under low visibility conditions, a visible light / near-infrared dual-mode camera with adjustable LED fill light is used to synchronously acquire images of the road ahead, lock the exposure and timestamp, and obtain a raw sequence with stable contrast.
[0076] Step 2: Input preprocessing – Perform noise reduction, white balance and adaptive contrast enhancement (such as CLAHE), unify resolution and color space; perform controllable fogging enhancement on some samples based on atmospheric scattering model to expand data distribution.
[0077] Step 3: Geometric Correction – Using Spatial Transformation Network (STN) and high-order polynomial distortion model to correct imaging pose and lens distortion, achieving viewpoint alignment and geometric standardization.
[0078] Step 4: Feature Extraction – The adaptive feature extraction module inputs the corrected image into the parallel grouped multi-scale depth separable convolutional backbone (GMCC-Net). Guided context recalibration is used to suppress fog noise and enhance boundary and far-range details in the position-channel dimension. Then, the feature diversity is expanded and cross-scale integration is completed through cross-stage partially connected parallel grouped multi-scale convolution (CSP-ParGMC).
[0079] Step 5: Cross-scale fusion—The cross-scale semantic fusion module performs weighted summation fusion on a top-down / bottom-up bidirectional path, and adopts learnable upsampling and channel shuffling (DSCB) operation to maintain boundary and semantic alignment; C3k2-DKMB dynamic multi-scale bottleneck is introduced at the fusion node to reshape the receptive field.
[0080] Step 6: Object Detection Inference – This step uses a lightweight grouped decoupled detection head (EGHead) to complete multi-scale detection on features at various scales. The classification branch outputs the class confidence, and the regression branch uses distributed bounding box representation (DFL) to obtain continuous bounding box parameters.
[0081] Step 7: Candidate box filtering - Sort by score and perform IoU-NMS non-maximum suppression to remove redundant boxes and output structured results (class, score, bbox).
[0082] Step 8: Risk Assessment – The target detection and risk assessment module combines relative distance and approach speed estimation to calculate the time to collision (TTC) and the overall risk score. If the threshold is exceeded, a warning / braking suggestion is generated.
[0083] Step 9: Results Recording and Adaptive Optimization – Save key intermediate quantities and evaluation results, perform hard case re-feeding and threshold adaptive adjustment based on false detection / missed detection samples, and continuously calibrate detection and early warning performance.
[0084] In summary, the method proposed in this invention takes a four-level closed loop as its core: physically consistent data acquisition and fogging enhancement, geometric and receptive field adaptive feature extraction, bidirectional weighted cross-scale semantic fusion, and decoupled detection and risk assessment, forming an end-to-end technical path from low visibility perception to executable early warning. The system acquires raw images under vehicle-mounted forward-looking camera and controlled supplemental lighting conditions, and expands the sample domain through controlled fogging and diversified enhancement based on an atmospheric scattering model. After geometric correction, a GMCC-Net backbone composed of parallel multi-scale and depthwise separable convolutions is used, combined with dynamic kernel generation and context-guided recalibration, to allocate equivalent receptive fields to different regions as needed without significantly increasing computational load, thereby robustly restoring the detail loss caused by fog-induced contrast reduction. Subsequently, bidirectional weighted EBLFPN and the learnable upsampling and channel shuffling module DSCB enhance cross-layer alignment and key region representation, significantly improving the distinguishability of distant and small-sized targets. At the detection end, the lightweight grouped decoupled detection head EGHead and distributed bounding box regression DFL work together with IoU-NMS to efficiently output structured results, and integrate quantitative indicators such as relative distance and approach speed to calculate TTC and comprehensive risk score, forming a deployable warning interface for ADAS.
[0085] Compared with existing methods, this invention has comprehensive advantages in low-computing-power deployment, real-time performance, robustness in foggy weather, and small target detection rate. It can be flexibly adapted to vehicle-side / roadside integrated scenarios and edge computing platforms, facilitating large-scale deployment and standardized operation and maintenance. It provides reliable, scalable, and traceable technical support for intelligent driving perception, road monitoring, and proactive safety decision-making under low visibility conditions.
[0086] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; only preferred embodiments of the present invention are illustrated. The descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. As long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0087] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this invention should be determined by the appended claims.
Claims
1. A lightweight heterogeneous co-operative target detection method for fog driving conditions, characterized in that, Includes the following steps: Step S1: Collect road images in foggy weather using a vehicle-mounted dual-mode camera and an adjustable LED fill light device. Perform adaptive exposure control and image preprocessing based on an atmospheric scattering model to obtain a standardized input image. Step S2: Input the standardized input image into the parallel multi-scale feature extraction network, and extract multi-level feature maps through the context-guided downsampling module and the parallel grouped multi-scale convolution module; Step S3: Input the multi-level feature map into the bidirectional weighted cross-scale semantic fusion network, and perform cross-scale feature fusion through the fog perception upsampling module and the fog adaptive feature fusion module to obtain the fused feature map; Step S4: Input the fused feature map into the lightweight grouped decoupled detection head, output the target category confidence and bounding box parameters through the classification branch and regression branch respectively, and output the detection result after non-maximum suppression; The bidirectional weighted cross-scale semantic fusion network performs weighted addition fusion in both top-down and bottom-up bidirectional paths, and the weighted addition fusion is performed through a fog attenuation-transmission joint alignment fusion strategy; The fog sensing upsampling module includes: acquiring a first layer input feature map corresponding transmittance map ; Based on local features, fog intensity With structural anisotropy coefficient , the learnable reorganization weights within the neighborhood are calculated by a weight generator ; According to up-sampling; Wherein, the fog intensity Depend on The anisotropy coefficients of the structure are obtained by logarithmic clipping mapping. Defined by the ratio of principal eigenvalues of the structure tensor; The fog attenuation-transmission joint alignment fusion strategy includes: For pyramid layer sets Features of each layer At pixel position Define the intrinsic visibility weight. The intrinsic visibility weight Determined by a weighted combination of transmittance and gradient magnitude; Solving from each layer to the anchor layer Alignment operator Minimize the sum of the weighted alignment error and the regularization term; Normalized fusion weights for each layer are calculated based on visibility weights and alignment similarity. And generate fused feature maps : ; ; In the formula, Indicates the level of scoring; This is a pixel-by-pixel similarity measurement function; Characteristic diagram representing the anchor layer; This indicates element-wise multiplication; This is the global average pooling operator.
2. The method according to claim 1, characterized in that, The adaptive exposure control in step S1 includes: Atmospheric light based on dark channel statistical estimation of preview frames The maximum response in the highlighted area is used to determine this. Contrast-robust global transmittance index Characterizes the level of visibility; Target grayscale mean With saturation upper limit To constrain this, we will jointly optimize the exposure time. With LED intensity ; When the global transmittance index Below the threshold At that time, prioritize increasing LED intensity. To avoid motion blur.
3. The method according to claim 1, characterized in that, Step S1 further includes controllable atomization enhancement, specifically: With clear, fog-free images With estimated depth As input, the scattering coefficient after perturbation With atmospheric light Synthetic enhanced fog image : ; in, , , and These are the random disturbance quantities uniformly sampled within a preset interval. The baseline scattering coefficients are estimated based on real fog images. It is a baseline atmospheric luminosity estimated based on real fog images; Using soft constraint regularization terms Limiting the enhancement strength, the soft constraint regularization term It includes the weighted sum of the contrast measure and the total variation measure.
4. The method according to claim 1, characterized in that, Step S1 further includes dual exposure fusion, specifically: Acquire short-exposure images in the neighborhood at the same time. With long exposure images ; Calculate the short exposure images respectively With long exposure images Pixel-level fusion weights and : ; ; The input brightness is generated by weighted fusion to extend the dynamic range. : ; In the formula, This represents the desired target gray level; This represents the weight decay coefficient that controls the weights around the target gray level. Indicates long exposure time The image captured below is in pixels Brightness at that location; Indicates short exposure time The image captured below is in pixels Brightness at that location; It is the allowed oversaturation threshold for a pixel. It is the allowed undersaturation threshold for a pixel; This is an indicator function that takes the value 1 when the condition within the parentheses is met. It is a very small constant greater than 0.
5. The method according to claim 1, characterized in that, Step S1 further includes visible light and near-infrared dual-mode guided fusion, specifically: Output visible light channel With near-infrared channel ; Based on transmittance estimation Calculate the fusion weights using the Sigmoid function. : ; Output combined luminance : ; In the formula, Represents the Sigmoid function; It is the slope coefficient of the Sigmoid function with respect to transmittance; It is the transmittance threshold.
6. The method according to claim 1, characterized in that, The context-guided downsampling module in step S2 includes: The input feature map is downsampled using a convolution with a stride of 2; Complementary contextual features are extracted by parallel convolutional branches with local depthwise convolutional branches and dilated depthwise convolutional branches. The outputs of the two branches are concatenated along the channel dimension and then processed by batch normalization and activation function. The number of channels is compressed using 1×1 convolution, and the channels are recalibrated using a global channel attention unit.
7. The method according to claim 1, characterized in that, The parallel grouped multi-scale convolution module in step S2 includes: The input feature map is divided into two parallel paths after being convolved by a 1×1 convolution. The first path directly proceeds to the splicing operation, while the second path is processed through N GMC units; The GMC unit is sequentially subjected to 3×3 standard convolution, 5×5 grouped convolution and 7×7 grouped convolution, and channel segmentation is performed after each level of convolution; the features of three different scales are concatenated in the channel dimension and channel reshaping is performed by 1×1 convolution. Finally, the outputs of the two paths are concatenated using a jump connection.
8. The method according to claim 1, characterized in that, Step S4 further includes risk assessment, specifically: Estimate the relative distance for each detected target. With relative velocity ; Calculate collision time ,in To prevent positive constants from being divided by zero; Calculate the overall risk score The comprehensive risk score Based on the collision time TTC and the spatial relationship between the target and the road passable area The weighted combination of the confidence scores is obtained by normalization using the Sigmoid function; When the comprehensive risk score A collision warning is output when the preset threshold is exceeded.
Citation Information
Patent Citations
Lightweight unmanned aerial vehicle image detection method fusing context and multi-scale features
CN120431484A
Foggy day real-time target detection method and system based on Fourier transform image defogging
CN121121078A