Interference scene target detection method based on channel enhancement and related device
By introducing a fast Fourier transform-based channel enhancement network FCENet to the object detection network, the channel enhancement processing of images is solved, and the traditional object detection network's performance degradation under contrast and low light conditions is achieved, and stronger robustness and detection capabilities are achieved.
Patent Information
- Application Number
- CN202510111289.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
AI Technical Summary
The traditional object detection network Faster R-CNN shows a problem of performance degradation when facing various interference scenarios, especially in contrast and low night light conditions.
The channel enhancement network FCENet based on the Fast Fourier transform is used to perform channel enhancement processing on the detected images, recover the contrast-damaged images, and input the enhanced images to the Faster R-CNN model for object detection.
The ability of the object detection network to learn and represent deep features of the image is significantly improved, and the robustness of the object under contrast changes and low light conditions is enhanced, ensuring that the target can be effectively detected while only using clear images for network training.
Smart Images

Figure CN119942084A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image target detection, and in particular relates to an interference scene target detection method based on channel enhancement and a related device. Background Art
[0002] At present, the target detection method based on deep learning performs well in normal scenes, but its performance is seriously reduced when facing various interference scenes. Among them, the traditional target detection network Faster R-CNN performs target detection on images interfered by contrast and low light at night, and the performance is reduced. The results are shown in the attached figure. Figure 1 As shown in the figure, Contrast interference refers to the following: the brightness difference between the target area and the background area becomes blurred or unclear, making it difficult to accurately extract the boundary or details of the target; Figure 1 It can be seen that in the attached Figure 1 (a) and Figure 1 (b) In two night scene images, the traditional object detection network Faster R-CNN mistakenly detects the bicycle as a person, a train, and a chair respectively; Figure 1 (c) and Figure 1 (d) In the two images affected by contrast, the traditional object detection network Faster R-CNN mistakenly detects a bus as a TV and a car as an airplane. The above serious performance degradation reduces the reliability of object detection technology and greatly limits the scope of application of object detection technology in life. Summary of the invention
[0003] In view of the technical problems existing in the prior art, the present invention provides a method and related devices for detecting targets in interference scenarios based on channel enhancement to solve the technical problem that the performance of the traditional target detection network Faster R-CNN degrades when facing various interference scenarios.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is:
[0005] The present invention provides a method for detecting interference scene targets based on channel enhancement, comprising:
[0006] The image to be detected is input into a channel enhancement network FCENet based on fast Fourier transform, channel enhancement processing is performed, and an enhanced image is output; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected;
[0007] The enhanced image is input into the Faster R-CNN model to perform target detection, and the target detection result is output.
[0008] Furthermore, the fast Fourier transform-based channel enhancement network FCENet includes a global average pooling operation module, a fast Fourier transform module, a spatial domain information extraction module and an information fusion module;
[0009] The global average pooling operation module is used to perform a global average pooling operation on the image to be detected to obtain a global average pooling result;
[0010] The fast Fourier transform module is used to perform a two-dimensional fast Fourier transform on the global average pooling result to obtain frequency information after the two-dimensional Fourier transform; and obtain amplitude spectrum information and phase spectrum information according to the frequency information after the two-dimensional Fourier transform;
[0011] The spatial domain information extraction module is used to perform a convolution operation on the global average pooling result to obtain a first tensor and a second tensor;
[0012] The information fusion module is used to perform a Hadamard product operation on the amplitude spectrum information and the first tensor to obtain enhanced amplitude spectrum information; perform a Hadamard product operation on the phase spectrum information and the second tensor to obtain enhanced phase spectrum information; fuse the enhanced amplitude spectrum information with the enhanced phase spectrum information to obtain a fused tensor; and multiply the fused tensor by the image to be detected to obtain an enhanced image.
[0013] Furthermore, the amplitude spectrum information is specifically:
[0014]
[0015] Among them, x a is the amplitude spectrum information; x g is the global average pooling result; FFT 2D (x g ) is the global average pooling result x g Frequency information after two-dimensional fast Fourier transform;
[0016] The phase spectrum information is specifically:
[0017] x p =arctan(FFT 2D (x g ))
[0018] Among them, x p is the phase spectrum information.
[0019] Further, a convolution operation is performed on the global average pooling result to obtain a first tensor and a second tensor, as follows:
[0020] Inputting the global average pooling result into the upper convolution branch and the lower convolution branch of the spatial domain information extraction module respectively, performing convolution operation, and obtaining a first tensor and a second tensor;
[0021] The first tensor is specifically:
[0022]
[0023] in, is the first tensor;
[0024] The second tensor is specifically:
[0025]
[0026] in, is the second tensor.
[0027] Furthermore, the enhanced amplitude spectrum information is specifically:
[0028]
[0029] in, is the enhanced amplitude spectrum information; is the first tensor; x a is the amplitude spectrum information;
[0030] The enhanced phase spectrum information is specifically:
[0031]
[0032] in, is the enhanced phase spectrum information; is the second tensor; x p is the phase spectrum information.
[0033] Furthermore, the process of multiplying the fused tensor with the image to be detected to obtain an enhanced image is specifically: multiplying the fused tensor with the image to be detected through a broadcast multiplication operation to obtain an enhanced image.
[0034] The present invention also provides an interference scene target detection system based on channel enhancement, comprising:
[0035] An image enhancement module is used to input the image to be detected into a channel enhancement network FCENet based on fast Fourier transform, perform channel enhancement processing, and output an enhanced image; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected;
[0036] The target detection module is used to input the enhanced image into the Faster R-CNN model to perform target detection and output the target detection result.
[0037] The present invention also provides an interference scene target detection device based on channel enhancement, comprising:
[0038] a processor suitable for executing a computer program;
[0039] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the interference scene target detection method based on channel enhancement is executed.
[0040] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the interference scene target detection method based on channel enhancement is implemented.
[0041] The present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the interference scene target detection method based on channel enhancement is implemented.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] The present invention provides a target detection method for interference scenes based on channel enhancement, which uses a channel enhancement network FCENet based on fast Fourier transform to perform channel enhancement processing on a to-be-detected image to restore the image with impaired contrast; then, the enhanced image is subjected to target detection through a Faster R-CNN model, thereby improving the detection network's ability to learn and represent deep features of the image, and being able to meet the requirement of using only clear images for network training, while significantly enhancing its robustness under contrast changes and low light conditions.
[0044] Furthermore, in the fast Fourier transform-based channel enhancement network FCENet, by combining fast Fourier transform and adaptive learning, damaged images can be restored and enhanced, which is particularly suitable for processing contrast-interfered images and improving image clarity and contrast. Specifically, the global average pooling operation is used to reduce the size of the input image, laying the foundation for the lightweight network. The fast Fourier transform is used to transfer the global fusion features from the spatial domain to the frequency domain, and the convolution operation is used to enhance the global features of the spatial domain. Subsequently, the enhanced spatial domain features are fused with the frequency domain information to obtain new frequency domain information, and the inverse Fourier transform is used to convert it back to the spatial domain to obtain the channel adjustment coefficient generated by adaptive learning, which is used to restore the contrast-impaired image and effectively enhance the contrast-interfered image.
[0045] The present invention proposes an interference scene target detection system based on channel enhancement, an interference scene target detection device based on channel enhancement, a computer storage medium and a computer product, which have all the advantages of the above-mentioned interference scene target detection method based on channel enhancement. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 This is the result of the traditional target detection network Faster R-CNN performing target detection on images affected by contrast and low light at night;
[0048] Figure 2 A flow chart of the interference scene target detection method based on channel enhancement provided in Example 1;
[0049] Figure 3 This is a network architecture diagram of the fast Fourier transform-based channel enhancement network FCENet in Example 1;
[0050] Figure 4 This is a comparison diagram of the image in Example 1 before and after being enhanced by FCENet after being subjected to contrast interference of degree S5;
[0051] Figure 5 This is a qualitative detection result diagram of the FCE-FRCNN method and the traditional Faster R-CNN method on the Pascal-Contrast dataset in Example 1. DETAILED DESCRIPTION
[0052] In order to make the technical problems, technical solutions and beneficial effects solved by the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application; obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments; based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0053] Example 1
[0054] As attached Figure 2 As shown, this embodiment 1 provides an interference scene target detection method based on channel enhancement, comprising the following steps:
[0055] Step 1: Input the image to be detected into x org To the Fast-fourier Channel-enhance Network (FCENet) based on fast Fourier transform, channel enhancement processing is performed, and the enhanced image x is output e ; Among them, the fast Fourier transform-based channel enhancement network FCENet is used to optimize the contrast degradation problem in the image to be detected, so as to eliminate contrast interference while maximally retaining potential features that are useful for subsequent target detection tasks.
[0056] In this embodiment 1, the fast Fourier transform-based channel enhancement network FCENet includes a global average pooling operation module, a fast Fourier transform module, a spatial domain information extraction module and an information fusion module, as shown in the attached Figure 3 As shown; the global average pooling operation module is used to perform a global average pooling operation on the image to be detected to obtain a global average pooling result; the fast Fourier transform module is used to perform a two-dimensional fast Fourier transform on the global average pooling result to obtain the frequency information after the two-dimensional Fourier transform; according to the frequency information after the two-dimensional Fourier transform, the amplitude spectrum information and the phase spectrum information are obtained; the spatial domain information extraction module is used to perform a convolution operation on the global average pooling result to obtain a first tensor and a second tensor; the information fusion module is used to perform a Hadamard product operation on the amplitude spectrum information and the first tensor to obtain enhanced amplitude spectrum information; perform a Hadamard product operation on the phase spectrum information and the second tensor to obtain enhanced phase spectrum information; fuse the enhanced amplitude spectrum information and the enhanced phase spectrum information to obtain a fused tensor; multiply the fused tensor with the image to be detected to obtain an enhanced image.
[0057] Specifically, the image to be detected is input into the channel enhancement network FCENet based on fast Fourier transform, channel enhancement processing is performed, and the enhanced image is output. The steps are as follows:
[0058] Step 11: Perform a global average pooling operation x on the image to be detected org , get the global average pooling result x g ; Wherein, the global average pooling result x g , specifically:
[0059] x g =GAP(x org )
[0060] Among them, x g is the global average pooling result, x org is the image to be detected,
[0061] It should be noted that in order to reduce the computational complexity of the model and ensure reliable and stable image restoration and enhancement, a global average pooling operation is performed on the image to be detected so that the shape of the input image to be detected is compressed to achieve global information fusion, which significantly reduces the parameter data of the convolutional layer in the channel enhancement network FCENet based on fast Fourier transform. The channel enhancement network FCENet based on fast Fourier transform significantly improves the computational efficiency of Zheng Teng without sacrificing performance, which is particularly critical for performing image restoration and enhancement tasks in an environment with limited computing resources. Among them, parameter optimization not only reduces the storage burden of the model, but also reduces the computing requirements, effectively expanding the application scope of the channel enhancement network FCENet based on fast Fourier transform.
[0062] Step 12: average pooling result x g Perform a two-dimensional fast Fourier transform, and use a preset data function to obtain the frequency information after the two-dimensional Fourier transform; obtain the amplitude spectrum information x according to the frequency information after the two-dimensional Fourier transform a and phase spectrum information x p .
[0063] The amplitude spectrum information x a , specifically:
[0064]
[0065] Among them, x a is the amplitude spectrum information; x g is the global average pooling result; FFT 2D (x g ) is the global average pooling result xg Frequency information after two-dimensional fast Fourier transform.
[0066] The phase spectrum information x p , specifically:
[0067] x p =arctan(FFT 2D (x g ))
[0068] Among them, x p is the phase spectrum information.
[0069] Step 13: average pooling result x g Perform convolution operation to get the first tensor and the second tensor Specifically, the global average pooling result x g Input to the up-convolution branch in the spatial domain information extraction module to perform a convolution operation with a convolution kernel size of 1×1 to obtain the first tensor The upper convolution branch includes convolutional layers conv1 and ReLU; the global pooling result x g Input to the down-convolution branch in the spatial domain information extraction module to perform a convolution operation with a convolution kernel size of 1×1 to obtain the second tensor The lower convolution branch includes convolutional layers conv2 and ReLU.
[0070] The first tensor Specifically:
[0071]
[0072] in, is the first tensor.
[0073] The second tensor Specifically:
[0074]
[0075] in, is the second tensor.
[0076] Step 14: convert the amplitude spectrum information x a With the first tensor Perform Hadamard product operation to obtain enhanced amplitude spectrum information The phase spectrum information x p With the second tensor Perform Hadamard product operation to obtain enhanced phase spectrum information The enhanced amplitude spectrum information With the enhanced phase spectrum information Perform fusion to obtain the fused tensor By broadcasting the multiplication operation, the fused tensor With the image to be detected x org Multiply to get the enhanced image x e .
[0077] The enhanced amplitude spectrum information Specifically:
[0078]
[0079] in, is the enhanced amplitude spectrum information; is the first tensor; x a is the amplitude spectrum information
[0080] The enhanced phase spectrum information Specifically:
[0081]
[0082] in, is the enhanced phase spectrum information; is the second tensor; x p is the phase spectrum information.
[0083] In this embodiment 1, the advantage of FCENet lies first in its lightweight design. The number of parameters of the entire network is contributed only by its two 1×1 convolutional layers, each of which has 3 channels of input and output. This greatly reduces the number of network parameters and improves computational efficiency. Through the global average pooling operation, the shape of the input tensor is compressed, further reducing the number of parameters and computational complexity of the subsequent convolution part, laying the foundation for the lightweight of FCENet.
[0084] It should be noted that in the image enhancement stage, the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the input image; FCENet aims to maximize the retention of potential features useful for subsequent target detection tasks while eliminating contrast interference; FCENet achieves this goal by fine-tuning the image in the frequency domain, and uses Fourier transform and convolution operations to improve the overall contrast of the image without sacrificing important detail information; secondly, FCENet achieves the removal of contrast interference while retaining potential features for subsequent target detection.
[0085] Step 2: The enhanced image x eInput to the Faster R-CNN model for target detection, and output the target detection result.
[0086] It should be noted that in the target detection stage, the Faster R-CNN model is used to process the image enhanced by FCENet. Faster R-CNN is an efficient target detection framework that combines the region proposal network (RPN) and the fast region convolutional neural network (Fast R-CNN). Among them, the specific structure and processing flow of the Faster R-CNN model are the same as those of the traditional target detection network Faster R-CNN, which will not be repeated here.
[0087] In this embodiment 1, a joint enhancement and restoration network framework FCE-FRCNN is constructed by integrating FCENet and Faster R-CNN; the joint enhancement and restoration network framework FCE-FRCNN constrains the training process through a standard detection loss function to ensure that the network can accurately restore target information from the enhanced image and improve the learning and representation capabilities of deep image features, thereby achieving the premise that only clear images are used to train the network, while still improving the robustness of the network to contrast-interfered images; in this embodiment 1, no additional interference data training samples are introduced, which effectively overcomes the limitation that interference scene data samples are usually required when optimizing deep learning-based target detection methods in the past.
[0088] The test and result analysis are described in detail as follows:
[0089] (1) Experimental setup: The Cityscapes dataset and Pascal dataset, which are widely used in the field of object detection, are selected as datasets for network training. In order to verify the effectiveness of FCENet in solving contrast and low-light interference at night, the corresponding Pascal-Contrast dataset is first generated. Secondly, based on the Pascal validation set, a simulated night dataset DARK is created. In addition, a real night scene dataset ExDark (exclusively dark) is selected for more comprehensive experiments.
[0090] (2) Quantitative test
[0091] In order to verify the effectiveness of the interference scene target detection method based on channel enhancement described in this Example 1 in dealing with contrast and low-light interference at night, this Example 1 compares the performance of FCE-FRCNN and other traditional methods on Cityscapes-Contrast, Pascal-Contrast, DARK and ExDark datasets.
[0092] As shown in Table 1 below, the AP of different comparison methods on the Pascal-Contrast dataset is given in Table 1. 50 , nAP, mPC and rPC scores; the nAP column shows the performance of each method on the standard Pascal dataset; methods with higher mPC and rPC scores show better robustness; mPC uses AP 50 Indicator measurement;
[0093] As can be seen from Table 1, except for the combination of YOLOv8s and FCENet, which did not show any performance improvement, the other comparison methods all achieved performance improvement under different degrees of contrast interference; especially Faster R-CNN, under the network framework FCE-FRCNN proposed in Example 1, the average performance of the five interference levels exceeded the original model by 9 percentage points; under the highest contrast interference S5, the AP of the network framework FCE-FRCNN was 50 The score is almost twice that of the baseline network Faster R-CNN, showing significant superiority; in addition, in the face of S5 level contrast interference, the YOLOv3 and YOLOv4 baseline networks have an AP 50 The scores increased by 3 percentage points and nearly 5 percentage points respectively. When RetinaNet was combined with FCENet as the baseline network, starting from the contrast interference of S3 level, AP 50 The indicators were improved by 5 percentage points, 13 percentage points and 16 percentage points respectively; when Efficient DetD0 was combined with FCENet as the baseline network, the average performance at five levels of interference was improved by nearly 7 percentage points; it should be noted that after FCENet was combined with YOLOv7, the performance on the standard Pascal dataset was improved by nearly 7 percentage points; therefore, the above results show that FCENet can effectively improve the performance of single-stage and two-stage target detection networks, proving its good versatility.
[0094] Table 1 Comparison of various indicators of the comparison methods on the Pascal-Contrast dataset
[0095]
[0096]
[0097] To further demonstrate the superior performance of FCENet, the following Table 2 shows the AP of FCE-FRCNN, Faster R-CNN baseline network, and other low-light object detection methods on the ExDark dataset. 50Performance comparison; The above results are obtained by directly training on the ExDark dataset and verifying on its test set; in addition, the performance of each model under the "Low" lighting condition in the ExDark dataset is also specially emphasized in Table 2 below; among them, the "average AP of all categories" reflects the average AP (mAP) of the model under the 10 different lighting conditions contained in the ExDark test set; from the data in Table 2 below, it can be observed that after Faster R-CNN is combined with FCENet, the detection performance under the "Low" lighting type has achieved a significant improvement of nearly 2 percentage points; among all the compared methods, the network framework FCE-FRCNN proposed in this Example 1 has achieved the best results in comprehensive performance, which further confirms the effectiveness and advancement of FCENet in the field of low-light target detection.
[0098] Table 2 AP of the compared methods on the ExDark dataset 50 index
[0099]
[0100] To further explore the performance of the model under low-light conditions at night, several comparative methods were experimented on the self-made simulated night low-light DARK dataset in Example 1. The performance of FCE-FRCNN and other comparative methods on this dataset is given in Table 3 below. The results show that, except for YOLOv7, all other comparative methods have enhanced robustness under medium to high night interference after combining with FCENet. In particular, when Faster R-CNN is used as the baseline network, under S4 and S5 night interference, the AP of FCE-FRCNN is 50 The scores are nearly 6 and 16 percentage points higher than the baseline respectively; when YOLOv3 and RetinaNet are used as baseline networks, FCENet also shows excellent performance in the face of medium to high levels of night interference; it should be noted that although the performance of YOLOv7 combined with FCENet in night interference scenarios is not significantly improved, its performance on the standard Pascal dataset is improved by nearly 7 percentage points, which shows its potential under normal lighting conditions.
[0101] Table 3 Comparison of various indicators of the comparison methods on the DARK dataset (mPC uses AP 50 Indicator measurement)
[0102]
[0103] The above quantitative analysis results show that FCENet, whether combined with a single-stage or two-stage target detection method, can effectively improve the robustness of the baseline method against contrast interference; in addition, the qualitative results of FCENet on the ExDark and DARK datasets further verify its effectiveness in enhancing the generalization ability of the model.
[0104] (3) Qualitative test
[0105] In the qualitative experiment, FCENet is first used to restore and enhance images disturbed by contrast; then some samples with obvious missing visual cues are qualitatively analyzed and presented.
[0106] 1) Visualization of images affected by contrast before and after enhancement
[0107] As attached Figure 4 As shown, attached Figure 4 The comparison of the image before and after FCENet enhancement after being subjected to contrast interference of degree S5 is given in the figure; Figure 4 The first column is the original image before adding interference, the second column is the image after adding interference, and the third column is the image enhanced by FCENet.
[0108] From the attached Figure 4 As can be seen from the figure, FCENet effectively restores the color, contrast, and brightness of the image and makes the target in the image clearly visible.
[0109] 2) Detection results of samples with obvious lack of visual cues
[0110] As attached Figure 5 As shown, attached Figure 5 The qualitative detection results of the FCE-FRCNN method and the traditional Faster R-CNN method on the Pascal-Contrast dataset are given in the figure; among them, the detection results of FCE-FRCNN and the comparison method Faster R-CNN on the five categories of person, car, bus, airplane and cat in the Pascal-Contrast dataset are specifically given.
[0111] From the attached Figure 5It can be seen that in the five selected categories, the traditional Faster R-CNN has certain problems such as missed detection and false detection; for example, in the first row, Faster R-CNN misdetects a person as a cat, in the second and third rows, two categories are detected for one target, and in the fifth row, Faster R-CNN does not detect a cat; in contrast, the interference scene target detection method based on channel enhancement described in this embodiment 1 can detect the target more accurately.
[0112] The interference scene target detection method based on channel enhancement described in this embodiment 1 uses a channel enhancement network FCENet based on fast Fourier transform to perform channel enhancement processing on the image to be detected. FCENet is different from the traditional image restoration method. It integrates information in the spatial domain and the frequency domain. Specifically, the image is converted from the spatial domain to the frequency domain through Fourier transform, and the amplitude spectrum and phase spectrum information are extracted. Then, the amplitude spectrum and the phase spectrum are enhanced using the spatial domain information after convolution fusion. The space-frequency domain information fusion method fully utilizes the characteristics of time domain convolution and frequency domain convolution, so that the spatial domain information and the frequency domain information are effectively combined. Finally, FCENet obtains the adjustment coefficient through adaptive learning, adjusts each channel of the damaged image, and realizes image restoration and enhancement. Unlike traditional data enhancement technology, the enhancement of FCENet is adaptively adjusted for specific tasks, with higher flexibility and effect.
[0113] It should be noted that the enhancement in this embodiment 1 is different from the traditional image enhancement technology; specifically, the traditional data enhancement technology is mostly performed through some predefined image processing operations, for example: generating enhanced images through rotation, scaling, cropping, inversion, etc., the purpose of which is to provide augmented data for network training, rather than to restore and enhance the image information itself.
[0114] In this embodiment 1, FCENet is a lightweight, adaptively learnable image enhancement and restoration network. By combining fast Fourier transform and adaptive learning, FCENet can restore and enhance damaged images, and is particularly suitable for processing contrast-interfered images and improving image clarity and contrast.
[0115] Example 2
[0116] This embodiment 2 provides an interference scene target detection system based on channel enhancement, including an image enhancement module and a target detection module; the image enhancement module is used to input the image to be detected into a channel enhancement network FCENet based on fast Fourier transform, perform channel enhancement processing, and output an enhanced image; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected; the target detection module is used to input the enhanced image into a Faster R-CNN model, perform target detection, and output a target detection result.
[0117] In this embodiment 2, the channel enhancement network FCENet based on fast Fourier transform includes a global average pooling operation module, a fast Fourier transform module, a spatial domain information extraction module and an information fusion module; the global average pooling operation module is used to perform a global average pooling operation on the image to be detected to obtain a global average pooling result; the fast Fourier transform module is used to perform a two-dimensional fast Fourier transform on the global average pooling result to obtain frequency information after the two-dimensional Fourier transform; according to the frequency information after the two-dimensional Fourier transform, the amplitude spectrum information and the phase spectrum information are obtained; the spatial domain information extraction module is used to perform a convolution operation on the global average pooling result to obtain a first tensor and a second tensor; the information fusion module is used to perform a Hadamard product operation on the amplitude spectrum information and the first tensor to obtain enhanced amplitude spectrum information; perform a Hadamard product operation on the phase spectrum information and the second tensor to obtain enhanced phase spectrum information; fuse the enhanced amplitude spectrum information with the enhanced phase spectrum information to obtain a fused tensor; multiply the fused tensor with the image to be detected to obtain an enhanced image.
[0118] Example 3
[0119] This embodiment 3 provides an interference scene target detection device based on channel enhancement, including: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as an interference scene target detection program based on channel enhancement.
[0120] When the processor executes the computer program, the steps in the above-mentioned interference scene target detection method based on channel enhancement are implemented, for example: the image to be detected is input into the channel enhancement network FCENet based on fast Fourier transform, channel enhancement processing is performed, and an enhanced image is output; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected; the enhanced image is input into the Faster R-CNN model, target detection is performed, and the target detection result is output.
[0121] Alternatively, when the processor executes the computer program, the functions of each module in the interference scene target detection system based on channel enhancement are realized, for example: an image enhancement module, which is used to input the image to be detected into a channel enhancement network FCENet based on fast Fourier transform, perform channel enhancement processing, and output an enhanced image; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected; a target detection module, which is used to input the enhanced image into a Faster R-CNN model, perform target detection, and output a target detection result.
[0122] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the interference scene target detection method and device based on channel enhancement.
[0123] For example, the computer program can be divided into an image enhancement module and a target detection module, and the specific functions of each module are as follows: the image enhancement module is used to input the image to be detected into a channel enhancement network FCENet based on fast Fourier transform, perform channel enhancement processing, and output an enhanced image; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected; the target detection module is used to input the enhanced image into a Faster R-CNN model, perform target detection, and output a target detection result.
[0124] The interference scene target detection device based on channel enhancement can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The interference scene target detection device based on channel enhancement can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the interference scene target detection device based on channel enhancement is only an example of a terminal device and does not constitute a limitation on the terminal device. It can include more or fewer components, or combine certain components, or different components. For example, the interference scene target detection device based on channel enhancement can also include input and output devices, network access devices, and buses.
[0125] The interference scene target detection device based on channel enhancement can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The interference scene target detection device based on channel enhancement may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that this embodiment 3 is only an example of an interference scene target detection device based on channel enhancement, and does not constitute a limitation on the interference scene target detection device based on channel enhancement, and may include more or less components than shown in the figure, or combine certain components, or different components, for example, the interference scene target detection device based on channel enhancement may also include input and output devices, network access devices, buses, etc.
[0126] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the interference scene target detection device based on channel enhancement, and uses various interfaces and lines to connect various parts of the entire interference scene target detection device based on channel enhancement.
[0127] The memory can be used to store the computer program and / or module, and the processor implements various functions of the interference scene target detection device based on channel enhancement by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory.
[0128] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc.
[0129] In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, an internal memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0130] Example 4
[0131] This embodiment 4 provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the interference scene target detection method based on channel enhancement are implemented.
[0132] If the module / unit integrated in the interference scene target detection system based on channel enhancement is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0133] Based on such understanding, this embodiment 4 implements all or part of the processes in the interference scene target detection method based on channel enhancement described in the above embodiment 1, and can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of the interference scene target detection method based on channel enhancement can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or preset intermediate form, etc.
[0134] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0135] It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electrical carrier signals and telecommunication signals.
[0136] Example 5
[0137] This embodiment 5 provides a computer product, which includes a computer program, which is stored in a computer-readable storage medium; the processor of the interference scene target detection device based on channel enhancement reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the interference scene target detection device based on channel enhancement can execute the interference scene target detection method based on channel enhancement described in Embodiment 1, which will not be repeated here.
[0138] It should be noted that a person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0139] The interference scene target detection method based on channel enhancement described in the present invention is intended to effectively cope with the challenges of target detection in contrast-impaired and low-light interference scenes at night; wherein, the channel enhancement network FCENet based on fast Fourier transform is a lightweight and adaptively learnable image restoration network, which mainly fuses spatial domain and frequency domain information by combining fast Fourier transform with convolution operation; and, FCENet generates channel adjustment coefficients obtained by adaptive learning, adjusts the channels of contrast-interfered images, effectively enhances the contrast and texture details of the image, and thus restores the original information of the image.
[0140] In the present invention, in order to further improve the robustness of the target detection method facing contrast and low-light interference at night, a strategy of cascading training the FCENet and Faster R-CNN models is adopted to construct a joint image enhancement and target detection framework; the framework uses the loss function used in the normal target detection method to constrain the network training process, which not only completes the image enhancement task, but also improves the target detection network's ability to learn and represent deep image features, so that the target detection network still shows excellent robustness for images with impaired contrast and low-light interference at night when only clear images are used for training, which can be a major breakthrough in optimizing the robustness of traditional target detection methods.
[0141] In the present invention, a joint enhancement and detection network framework FCE-FRCNN is built by combining FCENet with the normal target detection method Faster R-CNN. The results show that FCENet significantly improves the robustness of Faster R-CNN in the face of contrast interference and the generalization of Faster R-CNN in the face of low-light interference at night, and further experiments on a variety of baseline methods verify the versatility of FCENet; while maintaining the model detection efficiency, the model's ability to learn and represent deep image features is enhanced, thereby reducing the model's dependence on the target interference domain data set; ultimately, the accuracy and robustness of the target detection method in a variety of interference scenarios are improved.
[0142] The target detection method disclosed in the present invention uses Fourier transform to analyze the cause of the degradation of target detection performance caused by contrast changes and low-light conditions at night, and proposes a channel enhancement network FCENet based on fast Fourier transform; the channel enhancement network FCENet based on fast Fourier transform uses a global average pooling operation to reduce the size of the input image, laying a foundation for the lightweight of the network; the fast Fourier transform is used to transfer the global fusion features from the spatial domain to the frequency domain, and the convolution operation is used to enhance the global features of the spatial domain; the enhanced spatial domain features are fused with the frequency domain information to obtain new frequency domain information, and the inverse Fourier transform is used to convert it back to the spatial domain to obtain the channel adjustment coefficient generated by adaptive learning, which is used to restore the image with damaged contrast; finally, FCENet effectively enhances the image with disturbed contrast, and constructs a joint enhancement and detection FCE-FRCNN network framework by combining with Faster R-CNN, thereby improving the ability of the detection network to learn and represent deep features of the image, and achieving the premise that only clear images are used for network training, significantly enhancing its robustness under contrast changes and low-light conditions.
[0143] The above embodiment is only one of the implementation methods that can realize the technical solution of the present invention. The scope of protection claimed by the present invention is not limited only to this embodiment, but also includes changes, replacements and other implementation methods that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention.
Claims
1. A method for detecting interference scene targets based on channel enhancement, characterized in that: include: The image to be detected is input into a channel enhancement network FCENet based on fast Fourier transform, channel enhancement processing is performed, and an enhanced image is output; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected; The enhanced image is input into the Faster R-CNN model to perform target detection, and the target detection result is output.
2. The interference scene target detection method based on channel enhancement according to claim 1 is characterized in that: The fast Fourier transform-based channel enhancement network FCENet includes a global average pooling operation module, a fast Fourier transform module, a spatial domain information extraction module and an information fusion module; The global average pooling operation module is used to perform a global average pooling operation on the image to be detected to obtain a global average pooling result; The fast Fourier transform module is used to perform a two-dimensional fast Fourier transform on the global average pooling result to obtain frequency information after the two-dimensional Fourier transform; and obtain amplitude spectrum information and phase spectrum information according to the frequency information after the two-dimensional Fourier transform; The spatial domain information extraction module is used to perform a convolution operation on the global average pooling result to obtain a first tensor and a second tensor; The information fusion module is used to perform a Hadamard product operation on the amplitude spectrum information and the first tensor to obtain enhanced amplitude spectrum information; Performing a Hadamard product operation on the phase spectrum information and the second tensor to obtain enhanced phase spectrum information; The enhanced amplitude spectrum information and the enhanced phase spectrum information are fused to obtain a fused tensor; and the fused tensor is multiplied by the image to be detected to obtain an enhanced image.
3. The interference scene target detection method based on channel enhancement according to claim 2 is characterized in that: The amplitude spectrum information is specifically: Among them, xa is the amplitude spectrum information; xg is the global average pooling result; FFT 2D (x g ) is the frequency information after performing two-dimensional fast Fourier transform on the global average pooling result xg; The phase spectrum information is specifically: x p =arctan(FFT 2D (x g )) Among them, x p is the phase spectrum information.
4. The interference scene target detection method based on channel enhancement according to claim 2 is characterized in that: The process of performing a convolution operation on the global average pooling result to obtain the first tensor and the second tensor is as follows: Inputting the global average pooling result into the upper convolution branch and the lower convolution branch of the spatial domain information extraction module respectively, performing convolution operation, and obtaining a first tensor and a second tensor; The first tensor is specifically: in, is the first tensor; The second tensor is specifically: in, is the second tensor.
5. The method for detecting interference scene targets based on channel enhancement according to claim 2, characterized in that: The enhanced amplitude spectrum information is specifically: in, is the enhanced amplitude spectrum information; is the first tensor; x a is the amplitude spectrum information; The enhanced phase spectrum information is specifically: in, is the enhanced phase spectrum information; is the second tensor; x p is the phase spectrum information.
6. The method for detecting interference scene targets based on channel enhancement according to claim 2, characterized in that: The process of multiplying the fused tensor with the image to be detected to obtain an enhanced image is specifically: multiplying the fused tensor with the image to be detected through a broadcast multiplication operation to obtain an enhanced image.
7. A target detection system for interference scenes based on channel enhancement, characterized in that: include: An image enhancement module is used to input the image to be detected into a channel enhancement network FCENet based on fast Fourier transform, perform channel enhancement processing, and output an enhanced image; wherein the channel enhancement network FCENet based on fast Fourier transform is used to optimize the contrast degradation problem in the image to be detected; The target detection module is used to input the enhanced image into the Faster R-CNN model to perform target detection and output the target detection result.
8. A device for detecting interference scenes based on channel enhancement, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the interference scene target detection method based on channel enhancement as described in any one of claims 1 to 6 is executed.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the interference scene target detection method based on channel enhancement as described in any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the method for detecting interference scene targets based on channel enhancement according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Hotel scene picture target detection method based on Faster R-CNN-FFS model
CN113469272A
Convolutional neural network-based real-time low-illumination image target detection method
CN117115616A
Pedestrian image enhancement method and device based on multi-kernel feature fusion convolutional neural network, and medium
CN119206794A