Unified degraded scene target detection method based on target clues

By constructing a two-stage object detection model based on target clues, using clear images to extract object clues and perform denoising processing, combined with the diffusion model to perform feature mapping in hidden space, the problem of low object detection accuracy in degraded scenarios is solved, and higher detection performance and robustness are achieved.

CN120279245AActive Publication Date: 2025-07-08UNIV OF SCI & TECH BEIJING
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510306221.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-08
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing object detection algorithm has low detection accuracy in degraded scenarios and cannot meet the needs. Especially under complex conditions such as low light, haze, rain and snow, motion blur, the image quality decreases lead to blurring or loss of target features, affecting detection performance.

Method used

A two-stage object detection model based on target clues is adopted. By constructing detection branches, enhanced branches and diffusion denoising modules, the object clues are extracted and denoised using clear images, and feature maps are combined with the diffusion model in hidden space to achieve object detection in degraded scenarios.

Benefits of technology

It improves the target detection accuracy in degraded scenarios, reduces dependence on external image enhancement, optimizes computing redundancy, improves detection performance and robustness, and adapts to interference from multiple complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279245A_ABST
    Figure CN120279245A_ABST
Patent Text Reader

Abstract

The invention discloses a unified degraded scene target detection method based on a target clue, and belongs to the technical field of target detection, and the method comprises the steps: constructing a two-stage target detection model based on restoration of clear features; wherein the model comprises a detection branch, an enhancement branch and a diffusion denoising module; the enhancement branch is used for extracting feature information of the input image as an object clue; the diffusion de-noising module is used for de-noising the object clue to obtain a target clue; the detection branch is used for realizing degraded scene target detection in combination with a target clue; training the model; and utilizing the trained model to realize degraded scene target detection. According to the technical scheme, a feasible solution normal form is provided for solving the problems that degraded scene target information is seriously insufficient and various types of complex environment interference feature information exists, and a new direction and thought are provided for further researching a degraded scene detection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a unified degraded scene object detection method based on object clues. Background Art

[0002] Object detection refers to the technology of automatically identifying and locating objects in an image, and its task is to detect the regions and positions of target objects in a given image. In recent years, object detection algorithms based on deep learning have made great progress and branched into two important directions, namely, methods that only use convolutional neural networks (CNNs) and methods based on Transformers. Among them, the networks that only use CNNs can be further divided into two-stage detection frameworks (including a preprocessing step for generating object proposals) and single-stage detection frameworks.

[0003] Two-stage object detection algorithms are a type of object detection method based on the Region Proposal Network (RPN). They are divided into two steps: generating candidate target regions (Proposals) and processing the candidate regions to predict the categories and bounding boxes of the targets. In the first step, the RPN uses a sliding window method to extract regions on the image. Each proposed region is defined as a rectangle containing a sample object, and these regions are mapped to a smaller fixed-size feature map through a CNN, enabling shared computing. In the second step, the detection model uses these proposed regions as inputs and performs classification and regression to determine the bounding boxes and categories of the predicted targets. The main advantage of this method is relatively high detection accuracy, but the computational complexity is relatively high.

[0004] One-stage object detection algorithms do not require the RPN stage and directly generate the class probabilities and position coordinate values of objects. They are object detection algorithms that directly output the detection results through a single forward pass, so they have a faster detection speed. YOLO regards object detection as a regression problem of spatially separated bounding boxes and related class probabilities. The entire detection is a single network that predicts the bounding boxes and class probabilities from the complete image in one evaluation and can be directly optimized end-to-end in terms of detection performance.

[0005] Object detection algorithms based on Transformers are a research direction that has received much attention in recent years. Their main advantage is that they can use the Transformer model to replace the traditional RCNN module. Therefore, there is no need to use general object detection processes such as anchor boxes and non-maximum suppression (NMS) methods. They can directly encode the global image and regard all object detection box predictions as a standard unordered set prediction problem, and use the self-attention mechanism to capture the dependencies between objects.

[0006] Although general object detection networks perform excellently in standard scenarios, directly applying these networks in degraded scenarios often faces many challenges. Degraded scenarios include complex conditions such as low light, haze, rain, snow, motion blur, occlusion, and low resolution. These factors will significantly reduce the image quality, resulting in blurred or lost object features, thus affecting the detection performance. Problems such as decreased image quality, unclear object features, decreased multi-scale detection performance, and insufficient model robustness will all lead to a significant reduction in detection accuracy.

[0007] Object detection in degraded scenarios is an important research direction in the field of computer vision, aiming to solve the problem of how to accurately identify and locate objects under the conditions of decreased image quality, complex environment, or unclear object features. Existing methods have tried to solve the above problems from various aspects, but most methods are still enhancing images at the pixel level, introducing a large amount of computational redundancy and not making it fully beneficial for detection.

[0008] (1) Direct training method: Directly use degraded images to retrain the object detector. Its advantages lie in end-to-end optimization and avoiding cascading errors, but the disadvantages are significant. First, it has a high dependence on data and requires a large amount of labeled data in degraded scenarios. In actual scenarios, there are various types of degradation, and the cost of data acquisition and annotation is extremely high. Second, its generalization ability is limited. The model trained for a specific degradation type is difficult to adapt to other degraded scenarios (such as the coexistence of haze and low light). In addition, due to the low quality of the original images in degraded scenarios, it may be difficult to extract effective features directly for detection, resulting in performance bottlenecks.

[0009] (2) Cascade training method: Cascade image enhancement and detection in stages. First, enhance the degraded images and then input them into the pre-trained detector. The advantages of this method lie in modular design and flexible replacement, but its disadvantages are more prominent. First, enhancement and detection are carried out in stages, resulting in serious computational redundancy and difficult to meet the real-time requirements. Second, the performance of the enhancement algorithm directly affects the detection results. If the enhancement fails (such as excessive denoising leading to the loss of objects), the detector is difficult to remedy, resulting in error accumulation. In addition, image enhancement is oriented towards visual quality (such as PSNR, SSIM), while detection requires preserving key object features, and the two goals may conflict, further limiting the detection performance.

[0010] (3) Joint training method: Enhancing and detecting modules through end-to-end joint training is theoretically the most promising, but its drawbacks cannot be ignored. First, the design complexity is high, requiring careful design of the network structure and loss function, and the training is difficult. Second, the joint model usually has a large number of parameters, high requirements for data and computing resources, and it is difficult to deploy in resource-constrained scenarios. In addition, the coupling between enhancement and detection may lead to difficult analysis of the intermediate process and reduce the interpretability of the model. Although joint training can theoretically achieve the global optimal solution, its implementation complexity and high requirements for data and computing power limit its wide promotion in practical applications.

[0011] In summary, the existing object detection algorithms have low detection accuracy in degraded scenarios and cannot meet the requirements. Summary of the Invention

[0012] The present invention provides a unified degraded scenario object detection method based on object clues to solve the technical problem that the existing object detection algorithms have low detection accuracy in degraded scenarios and cannot meet the requirements.

[0013] To solve the above technical problems, the present invention provides the following technical solutions:

[0014] On the one hand, the present invention provides a unified degraded scenario object detection method based on object clues, including:

[0015] Construct a two-stage object detection model based on restored clear features; wherein, the model includes a detection branch, an enhancement branch, and a diffusion denoising module; the enhancement branch is used to extract the feature information of the input image as object clues; the diffusion denoising module is used to denoise the object clues to obtain target clues; the detection branch is used to combine the target clues to achieve object detection in degraded scenarios.

[0016] Train the two-stage object detection model based on restored clear features.

[0017] Use the trained two-stage object detection model based on restored clear features to achieve object detection in degraded scenarios.

[0018] Furthermore, the two-stage object detection model based on restored clear features adopts a two-stage decoupled training mechanism.

[0019] In the first-stage training, train the detection branch and the enhancement branch; wherein, the enhancement branch takes a clear image as input, extracts the feature information of the clear image as object clues, and inputs the object clues into the detection branch; the detection branch takes the degraded image and the object clues output by the enhancement branch as input, performs object detection on the degraded image, and optimizes the first-stage training by minimizing the detection loss.

[0020] In the second-stage training, freeze the enhancement branch, train the diffusion denoising module, and fine-tune the detection branch; wherein, the diffusion denoising module takes the object cues output by the enhancement branch as input, and its output is input to the detection branch. The consistency loss is used as intermediate supervision for training, while maintaining the detection loss constraint to achieve the joint training of the diffusion denoising module and the enhancement branch.

[0021] Further, the diffusion denoising module performs denoising processing on the object cues output by the enhancement branch by constructing a denoising process based on a diffusion model to obtain target cues.

[0022] Further, the inference process of the two-stage object detection model based on restored clear features includes:

[0023] Taking the degraded image as input, mapping the degraded image to the pseudo-clear domain to obtain an enhanced image;

[0024] Input the enhanced image into the enhancement branch to obtain the object cues Z corresponding to the enhanced image e ;

[0025] Input Z e into the diffusion denoising module. Use Z e as the conditional input to the diffusion model to control the diffusion model to gradually remove the degradation effect during iterative denoising, and obtain the feature estimation of the clear image as the target cues;

[0026] Input the target cues and the degraded image into the detection branch to achieve object detection of the degraded image.

[0027] Further, the detection branch includes: a detection backbone network, a feature fusion module, and a detection head; wherein,

[0028] The detection backbone network gradually performs spatial downsampling on the input degraded image through multiple stages of stacked convolutional layers to obtain a multi-scale feature map as the feature information of the degraded image;

[0029] The feature fusion module fuses the feature information of the degraded image with the target cues to obtain fused features;

[0030] The detection head based on the fused features realizes object localization and classification through the anchor box mechanism or key point prediction.

[0031] Further, the process by which the feature fusion module fuses the feature information of the degraded image with the target cues to obtain fused features includes: after the feature map F n is output in the nth stage of the detection backbone network, separating F n along the channels into F α and Fβ , where F α is spliced with the target clue along the channel, and then a fusion convolution operation is performed using the C3 module in the YOLOv5 detector to adjust the number of channels to obtain enhanced features; the enhanced features are spliced with F β again to obtain the final fused features, which are used as the input for the next stage in the detection backbone network; where n is a preset value.

[0032] Furthermore, the detection backbone network deploys feature fusion modules at three levels, namely P3, P4, and P5, of the feature pyramid.

[0033] Furthermore, during the processing at the corresponding levels, the target clue will perform bilinear interpolation through the corresponding downsampling rate.

[0034] Furthermore, the enhancement branch extracts the object clue corresponding to the input image by stacking convolutional layers.

[0035] On the other hand, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.

[0036] On another hand, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.

[0037] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0038] The technical solution of the present invention adopts a two-stage collaborative training mechanism, decouples the detection task from the image restoration training, and optimizes the conflict between image restoration and object detection objectives in traditional end-to-end training. And a hierarchical feature fusion module is proposed to adaptively compensate for the degraded features using the target clue at the P3-P5 levels of the feature pyramid. Moreover, the target clue extraction is modeled as a conditional diffusion process in the latent space, getting rid of the dependence on external image enhancement, and realizing an end-to-end mapping from the noise space to the clear latent space. Thus, it provides a feasible solution paradigm for solving the serious lack of target information in degraded scenes and the problem of interference with feature information in various types of complex environments, provides a new direction and idea for further research on degraded scene detection algorithms, and has great application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a flowchart of the unified degraded scene object detection method based on target clues provided by an embodiment of the present invention;

[0041] Figure 2 It is a framework diagram of a two-stage object detection model based on restored clear features provided by an embodiment of the present invention;

[0042] Figure 3 It is a system block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0043] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0044] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the word "exemplarily" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0045] The first embodiment

[0046] This embodiment provides a unified degraded scene object detection method based on target clues, which adopts a two-stage decoupled training mechanism: in the first stage, a target clue latent space is constructed using clear image supervision, and in the second stage, a degradation-robust latent space mapping is achieved through a conditional diffusion model. And a target clue-driven feature compensation is constructed: dynamically regulating the cross-level fusion of the feature pyramid (P3-P5) to achieve adaptive degradation compensation with shallow detail retention and deep semantic consistency. This method can be implemented by an electronic device, which can be a terminal or a server. The execution process of this method is as Figure 1 shown, including the following steps:

[0047] S1. Construct a two-stage object detection model based on restored clear features; wherein, the model includes a detection branch, an enhancement branch and a diffusion denoising module; the enhancement branch is used to extract the feature information of the input image as object clues; the diffusion denoising module is used to perform denoising processing on the object clues to obtain target clues; the detection branch is used to combine the target clues to achieve object detection in the degraded scene;

[0048] S2. Train the two-stage object detection model based on restored clear features;

[0049] S3. Use the trained two-stage object detection model to achieve object detection in the degraded scene.

[0050] It should be noted that the current object detection algorithms based on the degraded scene generally adopt a serial processing paradigm of "enhancement first and then detection", that is, first preprocess the low-quality input through pixel-level image enhancement technology, and then use the enhanced result as the only input of the detection network. This paradigm has the following bottlenecks: First, the enhancement network and the detection task are designed separately. The enhancement network is usually designed independently outside the detection task, and there is a lack of a multi-task collaborative optimization mechanism between the two. Although the PSNR / SSIM indicators of the enhanced image are significantly improved, its feature distribution deviates from the detector's preference, and the high-level semantic features required for the detection task are not explicitly modeled, making it difficult for the detection network to effectively utilize the fine-grained texture information extracted in the enhancement stage; the mainstream enhancement models rely too much on specific degradation physical priors, resulting in a sharp decline in performance in complex scenes where heterogeneous degradation factors (such as haze + motion blur + noise) coexist; in addition, even some joint optimization schemes are still limited to the pixel-level enhancement framework and fail to solve the problems of redundant calculation and ineffective feature transfer. These dual limitations of feature adaptability and scene generalization ability severely restrict the practical value of the algorithm in the real open environment.

[0051] To address the above problems, this embodiment proposes a two-stage object detection model based on restored clear features, which adopts a two-stage collaborative training mechanism. In the first stage, during the training process, we add a clear feature generation branch. Through the significant detail extraction module, clear features Z c are extracted from the clear image I c , and these features contain information that is difficult to detect or even missing in the degraded image. Subsequently, using the feature embedding module and downsampling operation, we flexibly inject Z c into the three-scale feature maps of the object detection backbone network and fully fuse it with the original degraded features, thereby improving the detection performance. In the second stage of training, we remove the model's dependence on external information by restoring the feature Z c , and at the same time avoid the computational redundancy brought by pixel-level image restoration. Since the process of recovering the clear image information Z d from the degraded image I c is extremely similar to the process of gradually denoising the latent space diffusion model (LDM) at the feature layer level, we use a conditional diffusion model to estimate the clear feature Z d when only the degraded image I c is input. Specifically, we also use the significant detail extraction module to extract the feature Z d from I d, and use it as a condition to input into the diffusion model, controlling the diffusion model to gradually remove the degradation effect during iterative denoising, restore the detailed information, and finally obtain a high-quality clear feature estimation. Then, like in the first stage, inject the restored features into the object detection backbone network for joint fine-tuning, so as to obtain a two-stage detection network based on the restored features. In addition, our one-stage network is combined with the underwater enhancement network to provide another complete solution.

[0052] The overall architecture of this two-stage object detection model based on restored clear features is as Figure 2 shown, including a detection branch, an enhancement branch, and a diffusion denoising module. The enhancement branch takes the clear image as input and extracts clear object cues through stacked convolutional layers. The detection branch uses a feature fusion module (FFM) to integrate the object cues with the original detection backbone to improve object detection in visually degraded scenes. To directly extract object cues from the degraded image, we utilize the powerful data estimation ability of the diffusion model, using the degraded object cues as conditions, and by replacing Z c with to achieve a unified inference pipeline that does not rely on external enhancement methods or degradation-specific priors. The working process of the model can be summarized as follows:

[0053] Given the degraded image I d as input, our goal is to detect the target object in I d . The basic object detection is completed by the object detection branch, which includes a detection backbone network and a detection head. To compensate for the degradation effect of I d compared to the clear image I c , we propose to extract the cues that are helpful for object detection, called object cues Z, through the image enhancement branch and the diffusion module, and use I c as a guide during the training phase. Specifically, as Figure 1 shown, the image enhancement branch takes the clear image I c (or the enhanced image I e ) as input and outputs Z c (or Z e ). To obtain the object cues without relying on I c (or I e ), our diffusion model module takes I d as input, and by constructing a denoising process based on the diffusion model, directly estimates Z d from Z c . Finally, the estimated object cues (i.e., ) are integrated into the object detection branch through the feature fusion module to enhance the object detection performance in visually degraded scenes. Next, the core technologies of the model will be introduced in detail module by module.

[0054] 1. Detection Branch

[0055] This module proposes an adaptive feature enhancement method for degradation scenarios. Through a feature reconstruction mechanism guided by multi-scale object cues, it effectively improves the robustness of object detection under complex imaging conditions. Here, it should be noted that traditional object detection frameworks usually adopt a cascaded architecture of "backbone network - detection head". The backbone network gradually performs spatial downsampling through stacked convolutional layers in N stages and finally outputs multi-scale feature maps. The detection head then realizes object localization and classification based on these feature maps through the anchor box mechanism or key point prediction. Since the "backbone network - detection head" of different detection frameworks has different designs, in implementation, we use Faster R-CNN with MobileNetV3-large as the backbone network as the baseline. We propose a feature enhancement paradigm guided by object cues. This paradigm does not rely on preprocessing for image quality improvement, but instead builds an object cue matrix in the latent space and establishes a degradation-robust feature compensation mechanism. The acquisition of the object cue matrix will be detailed in the next two branches. Specifically, after the feature map F is output at the nth stage of the backbone network n a feature fusion module (FFM) guided by object cues is introduced. This module adopts a two-branch processing mechanism: First, F n is separated along the channels into F α and F β , where F α is concatenated with Z along the channels, and then the C3 module commonly used in the YOLOv5 detector is used as the main fusion convolution operation and the number of channels is adjusted to obtain enhanced features. Finally, the enhanced features are re-concatenated with F β to obtain the final fused features, which are used as the input for the next stage in the backbone network. To ensure the consistent enhancement of multi-scale features, this method deploys the FFM module at three key levels of the feature pyramid: P3 (8× downsampling), P4 (16× downsampling), and P5 (32× downsampling). When processing at the P n layer, the object cues will be bilinearly interpolated through the corresponding downsampling rate to form scale-aware feature enhancement. This hierarchical design enables the network to: retain object detail features in the shallow layer; maintain semantic consistency in the deep layer; and achieve feature compensation through cross-scale information flow. It should be noted that due to the general design, it can be migrated to any framework based on convolutional network design.

[0056] 2. Enhancement Branch:

[0057] In the training stage, clear images can be directly obtained as supervision signals. At this time, feature decoupling can be performed through a significant detail extraction module (SDE) to obtain a high-confidence object cue matrix Z c, which encodes sensitive features such as target edges and textures. In the inference stage, considering the constraint that clear images are not available in real scenarios, we design a degradation-robust enhancement-detection closed-loop architecture. We use image enhancement methods as a preprocessing step to map degraded images to the pseudo-clear domain to obtain enhanced image I e , and use it to replace the clear image as the input of SDE to complete the entire inference process. Since there is no method that can enhance the image quality in any scenario, the choice of this model is usually determined by specific degradation scenarios, such as DEANet for foggy scenarios, GLARE for low-light scenarios, etc. The SDE module is designed as Figure 2 shown, which is stacked by several convolutional kernels of different sizes and different strides, and finally uses a 1×1 convolution to compress the channels to 1. However, this method has very serious limitations. During inference, it must rely on external inputs, making it bound to image enhancement methods.

[0058] 3. Diffusion denoising module:

[0059] To build an end-to-end unified inference framework that does not rely on external image enhancement and achieve adaptive extraction of target cues in degradation scenarios, we propose a cross-domain modeling method based on latent space feature restoration. The target cue extraction scheme based on the pixel domain is limited by the physical imaging prior differences of different degradation types in the pixel domain, making it difficult to establish a universal degradation-clear mapping relationship in the pixel domain and directly perform effective target cue extraction. In view of this, we turn to the implicit representation space and model the restoration of target cues as a conditional denoising process to establish a mapping relationship from the latent space of clear images to the space of degraded images.

[0060] Forward process: The forward process of the diffusion model is a process of gradually adding noise. The input is progressively perturbed through a parameterized Markov chain, that is, Z represented by the latent space of clear images c is mapped to the noise space by gradually adding Gaussian noise. During the process of gradually adding noise, affected by the noise, different types of degradations will become indistinguishable, which is more conducive to the establishment of a unified framework.

[0061] Backward process: The backward process of the diffusion model is a process of gradually removing noise. The mapping from the noise space to the clear latent space is achieved through progressive denoising. This process starts with random Gaussian noise and, under the guidance of the degraded target cue Z d , gradually strips off the noise interference through iterative refinement, and finally reconstructs the latent space representation of the clear image to achieve the mapping from the noise space to the clear latent space. Specifically, in the framework of the conditional diffusion model, the reverse denoising process uses randomly sampled Gaussian noise as the initial state. In each iteration, the degraded target cue Z generated by SDE dAs a conditional signal, it is embedded into the U-Net backbone network through the attention mechanism to guide the network to accurately separate the noise components and essential features. After T preset progressive optimizations, the reconstructed result aligned with the latent space distribution of the clear image is finally output. Benefiting from the efficient and compact representation characteristics of the latent space and the intrinsic similarity between the conditional matrix and the target representation, it only takes 5 iterations to converge to the optimal solution. In addition, the denoising network adopts a streamlined U-Net architecture, using fewer intermediate layer channels and stacked residual blocks, making our diffusion model lighter.

[0062] 4. Training and Inference of the Model:

[0063] During the training process, it is difficult to directly obtain target clues from the degraded image without intermediate guidance. Therefore, we adopt a two-stage training strategy to extract target clues from the degraded image. In the first stage, the detection branch and the enhancement branch are mainly trained to realize the target clues as the feature maps of the learning latent space representation driven by the object detection task, that is, to extract clear target clues Z from the clear image. c . Taking the degraded image and the clear image as inputs, the detection backbone and the detection head are jointly trained. The detection branch receives Z c as the object clue for detection, and optimizes the training of the first stage by minimizing the detection loss. In the second stage of training, the diffusion model is mainly trained and the detection model is fine-tuned. Specifically, only the degraded image is used as the input, and the degraded target clue is extracted as the conditional input of the diffusion model. The enhancement branch is frozen, and the output Z c is used as the guidance. The consistency loss is used as the intermediate supervision to train the diffusion model, while maintaining the detection loss constraint to maintain the compatibility of the feature space and the detection task.

[0064] During the inference process, there are two modes: 1). The detection branch uses the extracted from the enhanced image as an additional input, which corresponds to the above-mentioned enhancement branch. It should be noted that since the enhancement model and the detection branch are not jointly fine-tuned, this method is considered a cascaded method, called D4Det-Sep. 2). The detection branch uses the restored by DM as an additional input, which corresponds to the above-mentioned denoising process of the diffusion model and is our final unified method, called D4Det.

[0065] In summary, this embodiment provides a unified degraded scene object detection method based on target cues. The core lies in: Two-stage collaborative training mechanism: decouples the detection task from image restoration training, optimizing the conflict between image restoration and object detection objectives in traditional end-to-end training. Multi-scale target cue fusion architecture: proposes a hierarchical feature fusion module that adaptively compensates for degraded features using target cues at the P3-P5 levels of the feature pyramid. Latent space conditional denoising paradigm: models target cue extraction as a latent space conditional diffusion process, getting rid of the dependence on external image enhancement and achieving an end-to-end mapping from the noise space to the clear latent space. This unified degraded scene object detection method based on target cues has the following advantages compared with the prior art:

[0066] (1) Compared with the enhancement structure method based on physical priors: Existing methods usually design dedicated networks for a single degraded scene. For example, in foggy days, they rely on the atmospheric scattering model, and for low illuminance, they are based on the Retinex theory. Although they can achieve excellent results in specific scenarios, their strong scene dependence severely restricts the cross-scene generalization ability. Therefore, this paper proposes a target cue-driven feature space decoupling framework that maps the image to an implicit feature space, adaptively learns cross-scene shared degradation representations, and avoids relying on explicit physical assumptions. At the same time, it forces the network to only retain cross-scene general salient features while suppressing task-irrelevant redundant information and strengthening the expression of general details.

[0067] (2) Compared with other pixel-level image enhancement methods: Pixel-level enhancement methods focus on optimizing the low-level visual quality in the pixel domain, such as contrast enhancement and noise suppression, but ignore the collaborative enhancement of high-level semantic information and detector-sensitive features. Therefore, we construct a mapping from the degraded image to the implicit feature space and model feature enhancement as a progressive denoising process in the latent space, that is, starting from the degraded features and random noise as the initial state, and iteratively approaching the ideal feature distribution friendly to detection through reverse diffusion. This mechanism dynamically decouples the degradation interference and the target essential features through the noise prediction network, making the enhanced feature map not only visually reasonable but also retain the information beneficial for detection.

[0068] Next, the effectiveness of the method in this embodiment is verified.

[0069] The comparison results of the performance of this method with other state-of-the-art methods on the foggy dataset VOC-F and the RTTS dataset are shown in Table I. We choose Faster-RCNN with a MobilenetV3-large backbone as the baseline. Baseline* in the table indicates that the training data are clear images. Compared with other cascaded methods under the premise of using the same enhancement method, D4Det-Sep is significantly better than the corresponding cascaded methods. On the synthetic dataset VOC-F, when using DEANet as the enhancement network, while the cascaded method causes a 1.02% decrease in the map metric, this method improves by 2.21%. D4Det has achieved further improvement compared to D4Det-Sep. At the same time, on the real foggy dataset RTTS, its metric has increased by 6.67% compared to the baseline and by 4.13% compared to the best-performing JE-YOLO among other combined methods. This shows that without relying on an external image enhancement network and only using degraded images as input, a diffusion model can be used to refine better object cues. The experimental results verify the effectiveness of the method we proposed.

[0070] Table Ⅰ

[0071] Performance comparison with existing methods on foggy datasets

[0072]

[0073] The performance of this method and other state-of-the-art methods on the low-light VOC-D and ExDark datasets is shown in Table II. It can be seen from Table II that D4Det-Sep optimized by object cues is better than all other direct and cascaded methods, with significant detection gains on the ExDark dataset. Our D4Det also shows the highest mAP performance on synthetic and real datasets.

[0074] Table II

[0075] Performance comparison with existing methods on low-light datasets

[0076]

[0077] Furthermore, we extended the situation to a more general case, that is, there is no clear image corresponding to the degraded image during training. Therefore, we preprocess the degraded images using an enhancement method and use the enhanced images to complete the entire training process instead of clear images. As shown in Table III, on the underwater datasets URPC and DUO, our D4Det still demonstrates performance superior to other methods.

[0078] Table III

[0079] Performance comparison with existing methods on underwater datasets

[0080]

[0081] Extensive experiments of the solution proposed by the present invention in foggy days, low illuminance and underwater show that the method of the present invention is superior to the performance of other existing methods. The algorithm proposed by the present invention provides a feasible solution paradigm for solving the serious shortage of target information in degraded scenes and the problem of interference feature information in various types of complex environments. Since it does not involve any prior physical models, in addition to the above scenarios, this algorithm can be extended to vehicle detection in degraded scenes such as rainy and snowy days, providing a new direction and idea for further research on degraded scene detection algorithms.

[0082] Second Embodiment

[0083] This embodiment provides an electronic device, as Figure 3 shown, the electronic device includes: a processor and a memory; wherein, the processor and the memory can be connected through a communication bus; at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the first embodiment above. In addition, the electronic device may further include a transceiver, the processor and the transceiver can be connected through a communication bus, and the transceiver is used for communicating with other devices.

[0084] Next, in combination with Figure 3 specifically introduce each component of the electronic device:

[0085] Among them, the processor is the control center of the electronic device. The electronic device may include multiple processors, and each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can be a single processor or a collective term for multiple processing elements. For example, the processor can be one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0086] In a specific implementation, as an embodiment, the processor may include one or more CPUs. For example Figure 3 CPU0 and CPU1 shown in []. Of course, this is only an exemplary illustration.

[0087] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation method can refer to the above method embodiments and will not be elaborated here.

[0088] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit ( Figure 3 not shown) of the electronic device. The embodiments of the present invention do not make specific limitations thereto.

[0089] The transceiver may include a receiver and a transmitter ( Figure 3 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit ( Figure 3 not shown) of the electronic device. The embodiments of the present invention do not make specific limitations thereto.

[0090] In addition, it should be noted that Figure 3 the structure of the electronic device shown in does not constitute a limitation to the device. The actual device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above may refer to the technical effects described in the first embodiment above, so they will not be elaborated here.

[0091] Third Embodiment

[0092] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment above. Among them, the computer-readable storage medium may be ROM, random access memory, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.

[0093] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention can take the form of all or part of a hardware embodiment, all or part of a software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented using software, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0094] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0095] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions in Figure 1 one process or multiple processes and / or blocks Figure 1the functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide for implementing in the process Figure 1 one process or more processes and / or boxes Figure 1 the steps of the functions specified in one box or more boxes.

[0096] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element. In addition, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B may be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship, which can be specifically understood with reference to the context. "At least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural items (items). For example, at least one of a, b or c may mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.

[0097] In addition, it can be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not mean the sequence of execution order. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0098] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0099] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. One can select some or all of the units according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0100] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0101] Finally, it should be noted that the above description is only the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those of ordinary skill in the art, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle of the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A unified degraded scene object detection method based on target clues, characterized in that, Including: Construct a two-stage object detection model based on restored clear features; wherein, the model includes a detection branch, an enhancement branch, and a diffusion denoising module; the enhancement branch is used to extract the feature information of the input image as object clues; the diffusion denoising module is used to denoise the object clues to obtain target clues; the detection branch is used to combine the target clues to achieve object detection in a degraded scene; Train the two-stage object detection model based on restored clear features; Use the trained two-stage object detection model based on restored clear features to achieve object detection in a degraded scene.

2. The unified degradation scenario object detection method based on target clues according to claim 1, characterized in that The two-stage object detection model based on restored clear features adopts a two-stage decoupled training mechanism; In the first-stage training, train the detection branch and the enhancement branch; wherein, the enhancement branch takes a clear image as input, extracts the feature information of the clear image as object clues, and inputs the object clues into the detection branch; the detection branch takes the degraded image and the object clues output by the enhancement branch as input, performs object detection on the degraded image, and optimizes the first-stage training by minimizing the detection loss; In the second-stage training, freeze the enhancement branch, train the diffusion denoising module and fine-tune the detection branch; wherein, the diffusion denoising module takes the object clues output by the enhancement branch as input, and its output is input into the detection branch, and the consistency loss is used as intermediate supervision for training, while maintaining the detection loss constraint to achieve the joint training of the diffusion denoising module and the enhancement branch.

3. The unified degraded scenario object detection method based on target clues according to claim 1, characterized in that The diffusion denoising module denoises the object clues output by the enhancement branch by constructing a denoising process based on a diffusion model to obtain target clues.

4. The unified degradation scenario object detection method based on target clues according to claim 3, wherein The inference process of the two-stage object detection model based on restored clear features includes: Taking a degraded image as input, mapping the degraded image to a pseudo-clear domain to obtain an enhanced image; Input the enhanced image into the enhancement branch to obtain the object clue Z corresponding to the enhanced image e ; Input Z e into the diffusion denoising module, and use Z e as the conditional input to the diffusion model, controlling the diffusion model to gradually remove the degradation effect during iterative denoising, and obtaining the feature estimation of the clear image as the target clue; Inputting the target clues and the degraded image into the detection branch to achieve object detection of the degraded image.

5. The unified degradation scenario object detection method based on target clues according to claim 1, characterized in that, The detection branch includes: a detection backbone network, a feature fusion module, and a detection head; wherein, The detection backbone network gradually performs spatial downsampling on the input degraded image through multiple stages of stacked convolutional layers to obtain a multi-scale feature map as the feature information of the degraded image; The feature fusion module fuses the feature information of the degraded image with the target clues to obtain a fused feature; The detection head based on the fused feature realizes object localization and classification through an anchor box mechanism or key point prediction.

6. The unified degradation scenario object detection method based on target clues according to claim 5, wherein The feature fusion module fuses the feature information of the degraded image with the target clue. The process of obtaining the fused feature includes: after the feature map F is output at the n-th stage of the detection backbone network n is obtained, F n is separated along the channels into F α and F β , where F α is concatenated with the target clue along the channels, and then the C3 module in the YOLOv5 detector is used for fused convolution operation and the number of channels is adjusted to obtain the enhanced feature; the enhanced feature is re-concatenated with F β to obtain the final fused feature, which is used as the input for the next stage in the detection backbone network; where n is a preset value.

7. The unified degraded scenario object detection method based on target clues according to claim 6, wherein, The detection backbone network deploys feature fusion modules at three levels, P3, P4, and P5, of the feature pyramid.

8. The unified degraded scenario object detection method based on target clues according to claim 7, characterized in that During the processing of the corresponding levels, the target clues will be bilinearly interpolated through the corresponding downsampling rate.

9. The unified degradation scenario object detection method based on target clues according to claim 1, wherein, The enhancement branch extracts the corresponding object clues of the input image through stacked convolutional layers.

Citation Information

Patent Citations

  • Multi-branch target detection method based on traffic scene

    CN110059554A

  • Marine ship target detection method based on deep learning in sea fog environment

    CN115909064A

  • Underwater image enhancement method based on de-noising diffusion probability model

    CN116883259A

  • Underwater low-illumination image enhancement method based on conditional diffusion model

    CN117911302A

  • Driving target detection method and device in rainy and foggy days, electronic equipment and medium

    CN117953462A